
At a glance
Role
Senior UX Designer · Project Lead
Product
Technical Equivalence, an internal R&D decision workflow
Team
UX, product and business analysis, Raw Material Management experts, engineering, AI delivery, Responsible AI and specialist reviewers across R&D
What I owned
Problem framing, workflow architecture, interaction design, prototyping, component specifications, stakeholder alignment and implementation partnership
What shipped
Evidence-linked AI comparisons, factor decisions, cross-group review requests, a consolidated assessment and controlled reprocessing
Confidentiality
NDA · Screens and data fictionalized
14 factors
4
5+ groups
Shipped
I think the easiest way to misunderstand this project is to call it document comparison. The documents matter. The difficult part was the decision.
When the business needs an alternative raw material, R&D has to decide whether it is technically equivalent to the material already in use. That means working through supplier documents, internal data and specialist knowledge across several groups. A small omission can change the outcome.
The original request was to use AI to compare the documents. Fair enough. But an answer from the model was only useful if a reviewer could understand it, check it and decide what to do next.
How can AI help experts move faster without hiding uncertainty or acting like it gets the final say?
That became the real design problem.
The part I had to own
I led the experience from early workflow definition through a development-ready prototype and implementation support. The issue was not arranging fourteen screens. It was turning a complicated review process into something the team could build and specialists could trust.
On one line, the workflow looked simple:
The workflow looks linear. The evidence, handoffs and decisions inside it were not.
Then the edge cases started.
What should the AI extract? What does a reviewer need before making a decision? When should another group be pulled in? What happens when two documents disagree? And how do fourteen separate decisions become one assessment without losing the reasoning underneath?
Fourteen comparisons were only half the product
The assessment covered fourteen factors, including chemical composition, specifications, sourcing, processing temperature, heavy metals, impurity profile, safety data, regulatory requirements and manufacturing process.
Each factor could use a different document, require a different specialist and lead to a different kind of judgment. A generic AI summary was not going to carry that.
I built the factor pattern around four practical questions:
What did the system find for the current and proposed materials?
What is different, missing or potentially important?
Where did the information come from?
What does the reviewer decide next?
The comparison keeps source data, the model’s interpretation and the human decision as separate layers.
The AI needed to be useful without looking certain
A clean answer can be dangerous when it hides the messy parts.
I treated the AI output as an interpretation, not a verdict. Extracted values, similarities, discrepancies and missing information are shown separately from the reviewer’s decision. The source links stay visible too.
That separation matters. If the evidence, AI summary and human judgment look like one answer, the user is kind of beholden to us. We have made the model look more certain than it is.
The decision controls come after the evidence. Reviewers can agree, disagree, ask another group or say that more information is needed. The human action is clear and attributable.
The communication problem was part of the product
The AI comparison solved one part. People still had to ask other people what the differences meant.
Raw Material Management might find an issue but need Regulatory, Safety, Analytical, Process Development or another specialist to judge its significance. Before this work, the evidence, request and response could end up split across messages and meetings.
I put review requests inside the assessment. A reviewer can select the group, add context and keep the response attached to the factor that started the conversation.
The design does not replace expert conversation. It gives everyone the same evidence and decision state to talk about.
Technical Equivalence did not have one human in the loop. It had several. Specialist groups made decisions within their areas of expertise, while Raw Material Management originated the assessment, retained authority to override a factor decision and owned the final go-ahead. The design had to make both expertise and authority visible.
The request, evidence, owner and response stay connected to the factor.
UAT found the assumption we had not challenged
From the start, the technical implementation compared the current material with a single source document. The AI was allowed to choose that source. At the time, that looked like a technical detail. It was actually a product decision.
We were new to working with an external AI delivery partner, and the collaboration was a good one. But this was one of the learning moments. The team, including me, had not challenged who should choose the source before the comparison began.
In UAT, reviewers made the need clearer: they had to see the comparison for every relevant supplier document in the table. Those documents may agree, add context or conflict. If the AI chooses one outright, the result can look clean while hiding the other comparisons a reviewer needs.
For iteration two, I proposed a tabbed comparison. The current material stays fixed on the left, while each tab opens a one-to-one comparison with a single supplier document. A badge on the tab shows how many deviations are present, so reviewers can move through every comparison without asking the AI to merge the documents or decide which source wins.
On the surface, this was a relatively small front-end change. Underneath, it was much larger. The backend had been built around one selected source and one comparison. Supporting this direction meant generating, storing and retrieving a separate comparison for every relevant document.
The table still requires horizontal scrolling. That was already familiar in Cornerstone, and it preserved the one-to-one evidence model instead of compressing several documents into a polished answer that could hide a conflict.
This direction was designed after UAT and planned for iteration two. It was not part of the shipped MVP.
The iteration-two direction keeps each document comparison separate. Tab badges flag deviations, and reviewers move through the evidence without letting AI choose the source for them.
Reprocessing could not erase a decision
Supplier evidence changes. A replacement document may arrive, or the team may learn that the original source was incomplete.
Simply rerunning the model could overwrite the context behind an earlier decision. So the workflow separates the new AI result from the existing human decision. The reviewer chooses whether to keep or revise it.
Updated machine output can prompt another look. It cannot quietly rewrite someone’s judgment.
Responsible AI had to show up where the decision happened
The Responsible AI review required us to tell people they were interacting with AI, that the output could be wrong and that they were responsible for checking it.
I put that message next to the AI interpretation and source links. A disclaimer buried in another screen would have checked a box. It would not have helped the reviewer at the moment it mattered.
The rest of the design supports the same idea:
AI content is labeled.
Source documents stay available.
Missing and conflicting information stays visible.
Human decisions are separate and attributable.
Reprocessing does not replace an earlier decision automatically.
What I shipped
The workflow shipped into the existing R&D platform with:
A consistent review pattern across fourteen factors
Side-by-side evidence for current and proposed materials
AI-generated similarities, differences and missing information
Direct links back to source documents
Factor decisions and rationale
Specialist review requests tied to the evidence
A consolidated summary and final decision
Controlled reprocessing when the evidence changes
Fourteen factor decisions roll into one assessment without flattening the evidence underneath.
What changed
Before this work, the documents, interpretation, specialist conversations and final decisions could live in different places. The shipped workflow keeps them attached to the same assessment.
That means the evidence stays close to the AI interpretation. The interpretation stays separate from the human judgment. Review requests stay attached to the question that caused them. And the final assessment still has the factor-level reasoning underneath it.
I do not have publishable post-launch analytics. What I can support is a shipped workflow that reflected how reviewers actually worked: document comparisons remained visible, specialist judgments stayed attached to the evidence and Raw Material Management retained final accountability.
What this changed in my practice
I think “human in the loop” is useful shorthand, but it can put the technology at the center. It makes the person sound like a safeguard attached to the AI—useful for now, perhaps removable later.
That is not how I think about this work. The R&D experts already owned the decision. AI had a valuable supporting role: extracting evidence, comparing supplier documents and surfacing differences across fourteen factors.
The design problem was fitting AI into a human system of expertise, authority and accountability—not asking people to fit themselves around the technology.
I also learned that accountability has a shape. It shows up in the distance between a claim and its source, the order of information, the wording on a button and what happens to an earlier decision when the AI runs again.
Do not ask someone to verify AI output unless the interface gives them the evidence, context and authority to do it.
Human first. AI in support.
What I would do differently
I would make the measurement plan part of kickoff. I would want assessment time, waiting time between groups, source-link use, AI corrections, re-review and specialist escalations tracked from launch.
I would also map every AI-controlled choice before UAT, not only the output reviewers could see. Choosing the source was already part of the decision. We treated it like an implementation detail until testing made the consequence visible.
I would bring more specialist roles into recurring usability sessions earlier. The shared pattern helped, but Regulatory, Safety and Raw Material Management do not read evidence in exactly the same way.
And I would push harder on learnability and accessibility in the inherited enterprise UI. Dense tables and specialist abbreviations may be familiar. That does not make them easy.
Why this work stayed with me
The result was not a chatbot attached to an old workflow. It was a decision system built around the line between automation and expertise.
I am proud of it because the complexity is real. It is in the documents, the handoffs, the exceptions and the consequences of being wrong. The design makes that work clearer without pretending it is simple.
Workflow
Supplier evidence → AI comparison → Raw Material Management review → specialist reviews → final assessment
Have questions about the workflow, the prototype, or the design decisions? Happy to walk through it live.
Confidentiality Much of my work at The Estée Lauder Companies involved confidential internal platforms and supplier information. Every screen in this case study has been fictionalized. The workflow and design decisions represent the shipped experience; the displayed names, values and documents are not production data.




