CASE STUDY

AI-powered mortgage document review workflow

Redesigned the human-in-the-loop workflow for an AI extraction engine — shifting from full-document auditing to an exception-based model that safely automated 67% of extracted fields.

Due to NDA, the designs are adaptations made to mimic the original designs.

ROLE

Lead Designer

TEAM

2 PM · ML Eng · Ops · SMEs

PLATFORM

Web · Enterprise

TIMELINE

3 Quarters

00

Executive

summary 

PROBLEM

Nine months after launch, Darwin’s mortgage document AI engine was a technical success but a commercial failure. While model accuracy steadily improved, operational costs remained unchanged—reviewers were still manually verifying 100% of the data to avoid regulatory liability.

SOLUTION

I reframed the project from a technical challenge ("How do we make the AI more accurate?") into a systemic governance and behavioral design challenge ("How do we give humans the psychological and legal permission to safely skip review work?").

IMPACT

By designing an action-based Confidence × Risk Framework paired with a new operational accountability policy, we shifted the system from a full-document audit into an exception-based workflow.

01 - THE STARTING POINT

The AI did the work. People checked everything anyway.

Mortgage servicing documents feed directly into regulatory reporting, legal processes, and loan operations. Before Darwin, analysts manually transferred that information into downstream systems manually, field by field.

Darwin automated extraction. The AI pulled values straight from documents and handed them to human reviewers for sign-off.

  • Loan numberTyped by hand
  • Payoff amountTyped by hand
  • Borrower nameTyped by hand
  • Unpaid principal balanceTyped by hand
  • Middle initialTyped by hand
  • Property addressTyped by hand

Analysts transferred every value, field by field, into downstream systems.

INHERITED DESIGN

The original flow. Extraction was automated; verification was not — and verification quietly became the new bottleneck.

02 - PROBLEM

Better accuracy wasn't moving the needle

At launch, Darwin's extraction accuracy was mixed and so the reviewers had good reason to verify every field manually. As the models improved, accuracy increased significantly, but reviewer behavior barely changed.


In a high-risk domain like mortgage processing, trust is earned through consistent reliability, not accuracy metrics alone. Reviewers continued checking every field because the cost of missing a single error outweighed the perceived benefit of trusting the AI.

0%25%50%75%100%TRUST GAPLaunchV2V3V4V5V5 — OUTCOMEExpected effort: 15%Actual effort: 97%82-point Trust GapActual Review EffortExpected Review EffortModel Accuracy

Engineering improved the model from 68% to 96% accuracy, yet reviewers were still reviewing 97% of fields. Something other than accuracy was preventing adoption.

03 - RESEARCH

Reviewers weren't evaluating accuracy. They were managing accountability.

Shadowing the work

12

shadowing sessions

8

reviewers observed

70+

field types audited

What reviewers actually said

Reviewer · session 4

“I can't tell which ones to trust.”


Reviewer · session 7

“If I skip one and it's wrong, that's on me.”

Reviewer · session 9

“It's faster to just check them all than to guess which one's going to bite me.”

KEY FINDINGS

No way to tell reliable from unreliable. 40+ field types, all rendered identically — whether accuracy was 97% or 61%.

The interface looked confident even when wrong. Errors on ~1 in 6 documents, but every field looked complete. Finding mistakes felt like luck.

Trust was fragile with nothing to rebuild it. 7 of 8 reviewers wouldn't skip any field. One bad catch set the standard for everything after.

The reframe

BEFORE

"How do we increase trust in the AI?"


IMPROVED

"How do we safely remove unnecessary

review work?"

04 - Design directions

Three concepts. Two killed. One breakthrough.

Due to NDA, the designs are adaptations made to mimic the original designs.

OPTIONS CONSIDERED

OPTION 1

Confidence colour-coding

Green/amber/red fields. Reviewers saw green and still verified. Added info, removed no responsibility.

KILLED

OPTION 2

Progressive disclosure

High-confidence sections collapsed by default.


Can lead to avoidance and errors. Visibility was important

KILLED

OPTION 3

Action-based

review states

Fields classified as "review required" got utmost priority always in edit mode.

Fields with high confidence were still available for a quick glance at all times, can be edited by clicking on them,

Leaving no room for ambiguity.

SHIPPED

05 - THE FRAMEWORK BEHIND IT

The combination of combination x risk

confidence × risk FRAMEWORK

The classification decision for each field combined two axes: extraction reliability (rolling accuracy) and business consequence (downstream regulatory, legal, or operational risk).


High confidence + low risk = automated. Any other combination = review required.

Low risk
High risk
High conf.
Low conf.

Hover to interact

Governance that made automation trustworthy

Compliance-agreed sampling

A random sample of automated extractions is reviewed each week under a regime agreed with Compliance

MANUAL EDIT OPTIONS

Every automated extraction still stores version and can be edited, meeting regulatory traceability requirements end-to-end.

Earned automation

Rolling audit process continuously evaluates automated fields. Any field type exceeding error thresholds automatically returns to manual review.

06 - Outcome & impact

Review became exception-based

43%

Reduction in manual

review effort

67%

Fields fully automated

6 hrs

Saved per reviewer

each week

4 days

Faster degradation

detection

0

Increase in escaped

errors

BEFORE AND AFTER

BEFORE

Every field reviewed, every document

AFTER

67% of fields fully automated

BEFORE

AI used as a recommendation engine

AFTER

Human attention as a managed resource

BEFORE

Degradation caught via analyst complaints

AFTER

Degradation detected 4 days earlier

BEFORE

Success = model accuracy

AFTER

Success = review effort safely reduced

07 - REFLECTIONS

Overall Impact

Policy over pixels

The UI took days. The governance framework such as compliance regimes, accountability models, earned automation thresholds — took months. Screens are only trustworthy if the system behind them is.

Human-in-the-loop

The goal can never be to remove humans from the process. It should be to make sure their attention landed only where it actually mattered.

The end

Condensed this case study for a quick scan, feel free to reach out for a detailed walkthrough

  • hello

    hola

    salut

    prego

    namaste

    ni hao

    olá

    ciao

    s̄wạs̄dī

    hallo

Yashvardhan Bhardwaj

Senior User Experience Designer

Designed with ❤️, Logic, and AI.

Copyright ©2026. All rights reserved.

Designed with ❤️, Logic, and AI.

Copyright ©2026. All rights reserved.

Yashvardhan Bhardwaj

Senior User Experience Designer

  • hello

    hola

    salut

    prego

    namaste

    ni hao

    olá

    ciao

    s̄wạs̄dī

    hallo

These case studies goes deep and are built for bigger screens.

For the complete walkthrough, detailed interactions, and full visuals, view them on desktop.

Protected Content

Please enter the password to access this page

Protected Content

Please enter the password to access this page