OnCase attention view with evidence and system blind spots visible

Cases are grouped by why they need attention. The evidence panel shows what is known, what is inferred, and what OnCase still cannot confirm.

Cases are grouped by why they need attention. The evidence panel shows what is known, what is inferred, and what OnCase still cannot confirm.

At a glance

Role

Product Designer · Research, product strategy, interaction and visual design

Product

OnCase, a self-initiated attention layer for community-based case managers

Team

Solo designer · AI-assisted build

What I owned

Problem framing, evidence modeling, research protocol, information architecture, interaction design, visual design and supervised-AI boundaries

What I delivered

A published working prototype, an auditable evidence model and a corrected interview protocol. Primary research with practicing case managers remains the next step.

5

protocol stress tests

protocol stress tests

2

weeks to prototype

weeks to prototype

60

Evidence entries

Evidence entries

103

Automated checks

Automated checks

The work behind the record

Case-management systems are built to preserve the official record: notes, compliance, reporting, and the audit trail. But that is not the same as helping a worker decide what needs attention today.

The worker still has to reconstruct who has gone quiet, what is waiting on a landlord or county office, what they promised to follow up on, and what may be getting close to urgent. In the evidence I reviewed, that second layer of work lived across spreadsheets, notebooks, Outlook reminders, OneNote, inboxes, and memory.

The design question became: could one connected attention layer make that reasoning easier to see without creating another system to maintain?

OnCase groups work by reason rather than date, names who owes the next move, separates completed work from documentation, and keeps missing evidence visible. It supports the worker’s judgment; it does not replace the official record.

Project boundary. This is a self-initiated prototype built from public evidence and fictional data. It has not been validated with case managers. The evidence catalog shaped the problem, and the synthetic sessions tested the research protocol. Neither is presented as primary research.


What OnCase changes for the worker

Reasons before dates

Work is grouped by why it needs attention, not reduced to a generic urgency score. This preserves context, although the grouping still needs validation with caseworkers.

Match is not capacity

An eligible resource can still be unavailable. Results therefore show the source, last-checked date, and a capacity caveat instead of implying that a match is a recommendation.

Silence remains unknown

No provider reply stays ‘No response.’ The system does not convert an absence of evidence into a decline.

Sent is not delivered

A referral can be sent and accepted without the client receiving the service. OnCase keeps the outcome open until delivery or an unresolved outcome is documented.

AI prepares; the worker decides

The agent can gather, compare, draft, and interpret. A worker controls outreach, consequential status changes, and the official record.


How I approached the problem

I built OnCase independently to explore how a case-management attention layer could help workers recover priorities, coordinate across records, and close referrals without hiding uncertainty. With no access to active caseworkers or production data, I treated the work as a bounded prototype: public evidence shaped the problem, synthetic sessions tested the research protocol, and the interface remained explicit about what still requires primary validation.

In about two weeks, I built a 60-entry evidence catalog from secondary sources — Reddit threads, government publications, systematic reviews, vendor docs — with every entry labeled by evidence class. I piloted my own research discussion guide against five synthetic, adversarial scoping sessions before spending a single real recruit on it. I wrote down an anchor-setting decision — which setting to design for, and what would change my mind — before I knew whether I’d guessed right. And I designed and shipped a working, accessible, Material 3 prototype where every decision on screen traces back to a source.

This case study is about the operating discipline behind the screens. The prototype matters because it turns that discipline into visible product behavior: evidence stays bounded, uncertainty remains legible, and automation stops where human authority begins.

The Method: An Evidence Catalog You Can Audit

Before I designed a single screen, I built a source catalog. Sixty entries, each one labeled with what kind of evidence it actually is, because “I read this on Reddit” and “this came from a peer-reviewed study” cannot carry the same weight in a design decision, and pretending otherwise is how research gets untrustworthy without anyone noticing.

Direct exploratory: Responses to my own posts and follow-up questions. Good for generating hypotheses and language, not representative of anything.

Organic community: Existing discussions, written independent of this project. Good for recurring behaviors and edge cases; identities unverified.

Authoritative secondary: Government publications, systematic reviews, peer-reviewed studies. Stronger support for workload and coordination patterns; context may not transfer.

Vendor: Current product pages and documentation. Maps what a tool claims to do, never proof it’s usable.

Inference: A conclusion I drew by connecting two or more sources. Flagged as unproven until real research checks it.

I wrote guardrails for myself before the pressure to overclaim showed up, not after: don’t describe Reddit responses as interviews, don’t assume every case manager is secretly drowning, don’t treat healthcare evidence as automatically true for nonprofit human services, don’t use caseload count alone as a workload measure.

Four settings looked like plausible places to anchor the design — supportive housing, elder and aging services, IDD services, community health navigation. I scored what the evidence actually supported, not what I found most interesting, and picked supportive housing. I said so in writing, before I knew if I’d guessed right: “I anchored design in the setting with the deepest organic evidence and best access while recruitment continued, and stated in advance what would change my mind.” That sentence is the whole method in miniature.


The Pilot: Synthetic Scoping Sessions

I’d written a 20-minute discussion guide for real scoping conversations with case managers, but recruiting takes time. Rather than let a guide sit untested, or spend an actual scarce recruit finding out a question landed badly, I ran the full guide against five synthetic participants built from the evidence catalog: two supportive-housing workers at different tenure, one elder and aging worker, one IDD worker, one community health navigator.

Three rules ran through every session. Isolation: each participant could read only its own brief — never the hypotheses, never the other four sessions. Anti-sycophancy: every brief instructed the participant to say “I don’t know” freely, correct the interviewer, and carry at least one trait built specifically to push back against my hypotheses. Evidence bounds: each brief cited the exact catalog entries it was built from, and where evidence was thin, the brief stayed sparse on purpose.

What five rehearsal sessions earned: the center of gravity for near-misses looked more like dependencies and missing signals than memory failures. A new falsifiable hypothesis emerged around reconciliation cadence tracking contact tempo rather than a fixed daily rhythm. Waits got named by who owes the next move — the landlord, the county, the vendor, the family — never by a generic status label. That shaped how OnCase’s waiting board is organized.

The flaw I caught and logged, not smoothed over. The isolation only ran one direction. In review I found two moments where I’d referenced a detail a participant hadn’t actually said aloud. Those two exchanges carry no weight, and I wrote the fix into the real interview protocol: blind the interviewer to everything but the same screener facts a recruiter would have. The pilot caught its own methodology gap before it could cost a real participant’s time.


The Craft: Building OnCase

The prototype began with the two opportunities the evidence supported most: a reason-grouped attention view and a waiting-state model, anchored on a fictional veteran supportive-housing case manager with 22 tenants. The header keeps the product boundary explicit: the official record stays in the case-management system underneath. OnCase doesn’t replace it. It explains it.

Grouped by reason, not by date. Five groups, each with a header stating why it exists: no recent signal, waiting gone quiet, promised, due soon, done but not yet documented. Sorting by due date alone was the thing every hand-built spreadsheet in my research already did. I wasn’t interested in rebuilding that.

Done and documented, split apart. Two chips on every item: a checkmark for done in real life, a second mark only once it’s in the record. Marking something done in OnCase never silently touches the record. A “draft a note from this” action turns the capture into the paperwork, so the two moments can collapse into one.

OnCase attention view with Denise K.'s evidence and system blind spots visible

The official system keeps the record. OnCase brings forward the reasoning a worker needs to decide what to do next.

Waits named by counterparty. Waiting-board lanes for landlord, county, property manager, family — each item states who owes the next move. A visibly distinct badge marks waits that nudging can’t shorten, so the system never implies persistence will fix a supply problem.

OnCase waiting board organized by counterparty

Waits organized by who owes the next move, not by a generic status. The badge marks a scarce-supply wait where more follow-up can’t manufacture the resource.

Provenance and stated blind spots. A “why you’re seeing this” drawer on every item, every claim tagged fact, inference, or suggestion, with its source and date. A separate section states outright what OnCase can’t confirm. I built this because the thing I heard most consistently in my research wasn’t a request for more automation — it was distrust of systems that sound certain when they shouldn’t be.

Corrections are first-class, not silent. Any non-fact claim can be flagged wrong, with a reason, logged. Anything OnCase extracts from a worker’s own notes can be dismissed as easily as it can be accepted. Nothing gets to become a fact just because an algorithm noticed it.

Supervised delegation, not autonomous case management. Ray’s legal-aid referral is the end-to-end agent example. OnCase searches a fictional approved directory, explains why each program matches and how recently its information was checked, prepares editable emails, and waits. Only the worker’s explicit approval creates the simulated send. Incoming replies remain proposed status changes until the worker confirms or corrects them, and silence remains no response.

Supervised agent resource review showing source freshness and capacity uncertainty

The agent prepares a bounded set of matches from the fictional approved directory. The worker still chooses recipients, and every result keeps source freshness and capacity uncertainty visible.


Reply review showing source messages, confidence, proposed outcomes, and no response

The consequential test is correction, not agreement. The low-confidence interpretation is deliberately wrong, the original message stays beside it, and no status changes until the worker confirms the meaning.


See how I defined the agent’s authority, human approval points, failure behavior, and evaluation plan.


What changed through critique

Design critique response, not participant research.

Provider outcomes

Before: The proposed status labels collapsed acceptance, screening, and unavailability into ambiguous choices.

After: The choices became ‘Provider accepted referral,’ ‘Intake or screening required,’ and ‘Provider unavailable.’

Unresolved referrals

Before: A consequential outcome form was compressed into the bottom of a narrow drawer, creating nested scrolling and limited writing space.

After: The action now opens a focused dialog with a required reason, an optional alternative path, explicit actions, keyboard focus, and Escape behavior.


How I would measure a real pilot

These are proposed measures for a real pilot, not results from this prototype.

Worker

  • Time to understand why an item surfaced

  • Referral completion time

  • Interpretation correction rate

  • Confidence in the reason shown

Operations

  • Aged no-response referrals

  • Percentage of referrals with known outcomes

  • Accepted-to-delivered drop-off

Client and service

  • Time to intake or service delivery

  • Confirmed service receipt

  • Unresolved referrals with a visible next path

Design Principles

State the confidence, always. Every claim on screen carries its evidence class or its source. Nothing borrows more certainty than it earned.

Keep uncertainty visible. Missing data, stale signals, and conflicting sources stay on screen instead of getting smoothed into false confidence.

Support judgment, never replace it. The system can summarize and suggest. It doesn’t silently create facts, close cases, or auto-apply a recommendation.

Corrections should be as easy as agreement. If dismissing an extraction takes more effort than accepting it, people rubber-stamp instead of thinking.

Never let synthetic and real evidence touch. A rehearsal is allowed to sharpen a question. It’s never allowed to answer one.

Say what would change your mind, before you know the answer. A decision made in the open, with its own falsification condition attached, is worth more than a decision defended after the fact.


What This Project Demonstrates

Research operations judgment under real constraint. A zero-dollar recruiting budget didn’t produce a skipped research phase. It produced an evidence-classing system, a documented anchor decision, and a synthetic pilot with guardrails against circular confirmation.

Intellectual honesty as a working practice, not a slogan. The protocol flaw in the synthetic pilot got caught, logged, and fixed before it reached a real participant. Nothing in this project claims more validation than it has earned.

Systems-level interaction design. Provenance, stated blind spots, corrections-as-first-class-state, and planted errors for trust calibration are explainability patterns built for a domain where a worker’s trust in the system has to be earned claim by claim, not assumed.

Visual and accessibility craft. A working Material 3 system built from an exported token set, not defaults, checked computationally for contrast, backed by an automated test harness rather than a one-time manual pass.

Comfort with ambiguity on a compressed timeline. Two weeks, no team, no budget, and a still-honest final artifact — because the discipline scaled down instead of getting cut when the resources did.

Known limits

OnCase remains a prototype. The ‘Crisis hits’ control is a moderator affordance that compresses elapsed time for a usability session; a production system would use real elapsed gaps and event signals. An expiring referral can also be visible inside the referral workflow without yet appearing in Needs Attention. That cross-workflow integration remains unfinished.


The thing I’m proudest of in this project isn’t a screen. It’s a sentence I wrote down before I knew if it was true: that I anchored the design on the best available evidence and said in advance what would prove me wrong. That discipline transfers directly to real work, where research budgets are finite and the difference between judgment and a guess is often whether the reasoning was recorded before the outcome was known.

The thing I’d do differently next time is run the synthetic pilot’s blinding fix from day one instead of catching it in review. I got lucky that the flaw was small enough to log and correct without touching any real data.


What I’d test next

  • Whether a real supportive-housing case manager’s first three moves on the attention view actually match the group headers, without ever reading them out loud — the single test this whole prototype was built to fail or pass.

  • Whether the reconciliation-cadence hypothesis from the synthetic pilot holds: does the “rebuild plan” rhythm really track contact cadence, or was that an artifact of how I built the personas.

  • Whether the counterparty-named waiting board still makes sense in a setting further from supportive housing.

  • Whether provenance and stated blind spots actually build trust in a real session, or whether they just add friction nobody asked for — the honest risk of any explainability feature is that it’s easy to design for a hiring panel and hard to validate with someone who just wants to get through their caseload.

Have questions about the workflow, the prototype, or the design decisions? Happy to walk through it live.

© 2026 Michael Cullinan Principal UX Designer

© 2026 Michael Cullinan Principal UX Designer

© 2026 Michael Cullinan Principal UX Designer