
At a glance
Role
Principal UX Designer
Product
Technical Equivalence, an internal R&D decision workflow
Team
UX, product and business analysis, Raw Material Management experts, engineering, AI delivery, Responsible AI and specialist reviewers across R&D
What I owned
Problem framing, workflow architecture, interaction design, prototyping, component specifications, stakeholder alignment and implementation partnership
What shipped
Evidence-linked AI comparisons, factor decisions, cross-group review requests, a consolidated assessment and controlled reprocessing
14 factors
4
5+ groups
Shipped
I think the easiest way to misunderstand this project is to call it document comparison. The documents matter. The difficult part was the decision.
When the business needs an alternative raw material, R&D has to decide whether it is technically equivalent to the material already in use. That means working through supplier documents, internal data and specialist knowledge across several groups. A small omission can change the outcome.
The original request was to use AI to compare the documents. Fair enough. But an answer from the model was only useful if a reviewer could understand it, check it and decide what to do next.
How can AI help experts move faster without hiding uncertainty or acting like it gets the final say?
That became the real design problem.
The part I had to own
I led the experience from early workflow definition through a development-ready prototype and implementation support. The issue was not arranging fourteen screens. It was turning a complicated review process into something the team could build and specialists could trust.
On one line, the workflow looked simple:
Supplier evidence → AI comparison → Raw Material Management review → specialist reviews → final assessment
Then the edge cases started.
What should the AI extract? What does a reviewer need before making a decision? When should another group be pulled in? What happens when two documents disagree? And how do fourteen separate decisions become one assessment without losing the reasoning underneath?
Fourteen comparisons were only half the product
The assessment covered fourteen factors, including chemical composition, specifications, sourcing, processing temperature, heavy metals, impurity profile, safety data, regulatory requirements and manufacturing process.
Each factor could use a different document, require a different specialist and lead to a different kind of judgment. A generic AI summary was not going to carry that.
I built the factor pattern around four practical questions:
What did the system find for the current and proposed materials?
What is different, missing or potentially important?
Where did the information come from?
What does the reviewer decide next?

The comparison keeps source data, the model’s interpretation and the human decision as separate layers.
The AI needed to be useful without looking certain
A clean answer can be dangerous when it hides the messy parts.
I treated the AI output as an interpretation, not a verdict. Extracted values, similarities, discrepancies and missing information are shown separately from the reviewer’s decision. The source links stay visible too.
That separation matters. If the evidence, AI summary and human judgment look like one answer, the user is kind of beholden to us. We have made the model look more certain than it is.
The decision controls come after the evidence. Reviewers can agree, disagree, ask another group or say that more information is needed. The human action is clear and attributable.
The communication problem was part of the product
The AI comparison solved one part. People still had to ask other people what the differences meant.
Raw Material Management might find an issue but need Regulatory, Safety, Analytical, Process Development or another specialist to judge its significance. Before this work, the evidence, request and response could end up split across messages and meetings.
I put review requests inside the assessment. A reviewer can select the group, add context and keep the response attached to the factor that started the conversation.
The design does not replace expert conversation. It gives everyone the same evidence and decision state to talk about.

The request, evidence, owner and response stay connected to the factor.
UAT found the part we had simplified too much
The first model assumed one source document per factor. In practice, suppliers can provide several. They may agree, add context or conflict.
Merging everything into one polished answer sounded efficient. It also created rules nobody had defined. Which document wins? What happens when the dates or values disagree? The system could easily hide the thing a reviewer most needed to see.
I designed a more controlled approach. One primary document drives the comparison. Supporting documents stay visible. Matching evidence can strengthen the source trail. Conflicts are called out instead of quietly reconciled.
It was less magical. It was also easier to explain, audit and ship.

One primary document drives the comparison while supporting and conflicting evidence remains visible.
Reprocessing could not erase a decision
Supplier evidence changes. A replacement document may arrive, or the team may learn that the original source was incomplete.
Simply rerunning the model could overwrite the context behind an earlier decision. So the workflow separates the new AI result from the existing human decision. The reviewer chooses whether to keep or revise it.
Updated machine output can prompt another look. It cannot quietly rewrite someone’s judgment.
Responsible AI had to show up where the decision happened
The Responsible AI review required us to tell people they were interacting with AI, that the output could be wrong and that they were responsible for checking it.
I put that message next to the AI interpretation and source links. A disclaimer buried in another screen would have checked a box. It would not have helped the reviewer at the moment it mattered.
The rest of the design supports the same idea:
AI content is labeled.
Source documents stay available.
Missing and conflicting information stays visible.
Human decisions are separate and attributable.
Reprocessing does not replace an earlier decision automatically.
What I shipped
The workflow shipped into the existing R&D platform with:
A consistent review pattern across fourteen factors
Side-by-side evidence for current and proposed materials
AI-generated similarities, differences and missing information
Direct links back to source documents
Factor decisions and rationale
Specialist review requests tied to the evidence
A consolidated summary and final decision
Controlled reprocessing when the evidence changes

Fourteen factor decisions roll into one assessment without flattening the evidence underneath.
What changed
Before this work, the documents, interpretation, specialist conversations and final decisions could live in different places. The shipped workflow keeps them attached to the same assessment.
That means the evidence stays close to the AI interpretation. The interpretation stays separate from the human judgment. Review requests stay attached to the question that caused them. And the final assessment still has the factor-level reasoning underneath it.
I do not have clean post-launch analytics I can publish. That is frustrating, but it is not a reason to make up a percentage. The result I can support is the shipped workflow and the way it makes evidence, ownership and decisions visible.
What this changed in my practice
I think “human in the loop” is easy to say and not very useful on its own. The harder question is: what does that person need to make a real decision?
In this product, the answer was evidence, context, authority and a clear place to disagree.
I also learned that accountability has a shape. It shows up in the distance between a claim and its source, the order of information, the wording on a button and what happens to an earlier decision when the AI runs again.
Do not ask someone to verify AI output unless the interface gives them the evidence, context and authority to do it.
What I would do differently
I would make the measurement plan part of kickoff. I would want assessment time, waiting time between groups, source-link use, AI corrections, re-review and specialist escalations tracked from launch.
I would also bring more specialist roles into recurring usability sessions earlier. The shared pattern helped, but Regulatory, Safety and Raw Material Management do not read evidence in exactly the same way.
And I would push harder on learnability and accessibility in the inherited enterprise UI. Dense tables and specialist abbreviations may be familiar. That does not make them easy.
Why this work stayed with me
The result was not a chatbot attached to an old workflow. It was a decision system built around the line between automation and expertise.
I am proud of it because the complexity is real. It is in the documents, the handoffs, the exceptions and the consequences of being wrong. The design makes that work clearer without pretending it is simple.
Workflow
Supplier evidence → AI comparison → Raw Material Management review → specialist reviews → final assessment
Have questions about the workflow, the prototype, or the design decisions? Happy to walk through it live.
Much of my work at The Estée Lauder Companies involved confidential internal platforms and supplier information. Every screen in this case study has been fictionalized. The workflow and design decisions represent the shipped experience; the displayed names, values and documents are not production data.