
$0
60
5
103
Start Here
I was between roles when I started this: no budget for recruiting, no company backing, no client data to draw on. I wanted a portfolio piece that would hold up in front of a health, clinical, or nonprofit hiring team — not just look good to a general audience.
In about two weeks, I built a 60-entry evidence catalog from secondary sources — Reddit threads, government publications, systematic reviews, vendor docs — with every entry labeled by evidence class. I piloted my own research discussion guide against five synthetic, adversarial scoping sessions before spending a single real recruit on it. I wrote down an anchor-setting decision — which setting to design for, and what would change my mind — before I knew whether I’d guessed right. And I designed and shipped a working, accessible, Material 3 prototype where every decision on screen traces back to a source.
This case study is about the operating discipline behind the screens. The prototype matters because it turns that discipline into visible product behavior: evidence stays bounded, uncertainty remains legible, and automation stops where human authority begins.
The Constraint
I do not want to design from a hunch, and I can’t afford not to. I started this project unemployed, which is its own kind of design brief: no budget to recruit case managers, no employer to give a cold outreach email any weight, no access to a real case-management system or real client data. Most solo projects in that position skip research or run a handful of informal conversations and call it research. Neither holds up if the room includes someone who has actually run a study.
So I treated the constraint as the brief itself. If I couldn’t buy primary research, I could still build something defensible: evidence with its confidence level stated up front, a synthetic rehearsal that never gets mistaken for a finding, and a prototype that shows its reasoning instead of hiding it behind a clean screen.
The Problem
Case-management platforms carry the official record: the compliance burden, the reporting burden, the audit trail. What they don’t reliably answer is a different, more immediate set of questions a frontline worker asks every day: Who needs my attention right now? Why does this case need it? What am I actually waiting on, and from whom? What did I promise, and to whom? What’s quietly about to become urgent?
Workers answer that second set of questions by hand, constantly, in tools nobody designed for the job: color-coded Excel trackers reinvented independently across separate threads, a five-by-eight paper calendar paired with a separate to-do notebook, OneNote pages per client, Outlook reminders built by hand, a literal checkbox for “was this entered into the record yet.”
The pattern across all of it: the official system holds the record, and the worker often has to reconstruct the reasoning. Existing products advertise tasks, reminders, dashboards, AI documentation, and referral tools — so I could not claim the market lacked these capabilities. The sharper question behind OnCase was whether a connected attention layer could make timely action and its reasons easier to see, without becoming another place to maintain.
The Method: An Evidence Catalog You Can Audit
Before I designed a single screen, I built a source catalog. Sixty entries, each one labeled with what kind of evidence it actually is, because “I read this on Reddit” and “this came from a peer-reviewed study” cannot carry the same weight in a design decision, and pretending otherwise is how research gets untrustworthy without anyone noticing.
Direct exploratory: Responses to my own posts and follow-up questions. Good for generating hypotheses and language, not representative of anything.
Organic community: Existing discussions, written independent of this project. Good for recurring behaviors and edge cases; identities unverified.
Authoritative secondary: Government publications, systematic reviews, peer-reviewed studies. Stronger support for workload and coordination patterns; context may not transfer.
Vendor: Current product pages and documentation. Maps what a tool claims to do, never proof it’s usable.
Inference: A conclusion I drew by connecting two or more sources. Flagged as unproven until real research checks it.
I wrote guardrails for myself before the pressure to overclaim showed up, not after: don’t describe Reddit responses as interviews, don’t assume every case manager is secretly drowning, don’t treat healthcare evidence as automatically true for nonprofit human services, don’t use caseload count alone as a workload measure.
Four settings looked like plausible places to anchor the design — supportive housing, elder and aging services, IDD services, community health navigation. I scored what the evidence actually supported, not what I found most interesting, and picked supportive housing. I said so in writing, before I knew if I’d guessed right: “I anchored design in the setting with the deepest organic evidence and best access while recruitment continued, and stated in advance what would change my mind.” That sentence is the whole method in miniature.
The Pilot: Synthetic Scoping Sessions
I’d written a 20-minute discussion guide for real scoping conversations with case managers, but recruiting takes time. Rather than let a guide sit untested, or spend an actual scarce recruit finding out a question landed badly, I ran the full guide against five synthetic participants built from the evidence catalog: two supportive-housing workers at different tenure, one elder and aging worker, one IDD worker, one community health navigator.
Three rules ran through every session. Isolation: each participant could read only its own brief — never the hypotheses, never the other four sessions. Anti-sycophancy: every brief instructed the participant to say “I don’t know” freely, correct the interviewer, and carry at least one trait built specifically to push back against my hypotheses. Evidence bounds: each brief cited the exact catalog entries it was built from, and where evidence was thin, the brief stayed sparse on purpose.
What five rehearsal sessions earned: the center of gravity for near-misses looked more like dependencies and missing signals than memory failures. A new falsifiable hypothesis emerged around reconciliation cadence tracking contact tempo rather than a fixed daily rhythm. Waits got named by who owes the next move — the landlord, the county, the vendor, the family — never by a generic status label. That shaped how OnCase’s waiting board is organized.
The flaw I caught and logged, not smoothed over. The isolation only ran one direction. In review I found two moments where I’d referenced a detail a participant hadn’t actually said aloud. Those two exchanges carry no weight, and I wrote the fix into the real interview protocol: blind the interviewer to everything but the same screener facts a recruiter would have. The pilot caught its own methodology gap before it could cost a real participant’s time.
The Craft: Building OnCase
The prototype began with the two opportunities the evidence supported most: a reason-grouped attention view and a waiting-state model, anchored on a fictional veteran supportive-housing case manager with 22 tenants. The header keeps the product boundary explicit: the official record stays in the case-management system underneath. OnCase doesn’t replace it. It explains it.
Grouped by reason, not by date. Five groups, each with a header stating why it exists: no recent signal, waiting gone quiet, promised, due soon, done but not yet documented. Sorting by due date alone was the thing every hand-built spreadsheet in my research already did. I wasn’t interested in rebuilding that.
Done and documented, split apart. Two chips on every item: a checkmark for done in real life, a second mark only once it’s in the record. Marking something done in OnCase never silently touches the record. A “draft a note from this” action turns the capture into the paperwork, so the two moments can collapse into one.

Cases grouped by why they need attention. Every item states its source, confidence, and what OnCase can’t confirm.
Waits named by counterparty. Waiting-board lanes for landlord, county, property manager, family — each item states who owes the next move. A visibly distinct badge marks waits that nudging can’t shorten, so the system never implies persistence will fix a supply problem.

Waits organized by who owes the next move, not by a generic status. The badge marks a scarce-supply wait where more follow-up can’t manufacture the resource.
Provenance and stated blind spots. A “why you’re seeing this” drawer on every item, every claim tagged fact, inference, or suggestion, with its source and date. A separate section states outright what OnCase can’t confirm. I built this because the thing I heard most consistently in my research wasn’t a request for more automation — it was distrust of systems that sound certain when they shouldn’t be.
Corrections are first-class, not silent. Any non-fact claim can be flagged wrong, with a reason, logged. Anything OnCase extracts from a worker’s own notes can be dismissed as easily as it can be accepted. Nothing gets to become a fact just because an algorithm noticed it.
Supervised delegation, not autonomous case management. Ray’s legal-aid referral is the end-to-end agent example. OnCase searches a fictional approved directory, explains why each program matches and how recently its information was checked, prepares editable emails, and waits. Only the worker’s explicit approval creates the simulated send. Incoming replies remain proposed status changes until the worker confirms or corrects them, and silence remains no response.

The agent prepares a bounded set of matches from the fictional approved directory. The worker still chooses recipients, and every result keeps source freshness and capacity uncertainty visible.

The consequential test is correction, not agreement. The low-confidence interpretation is deliberately wrong, the original message stays beside it, and no status changes until the worker confirms the meaning.
See how I defined the agent’s authority, human approval points, failure behavior, and evaluation plan.
The thing I’m proudest of in this project isn’t a screen. It’s a sentence I wrote down before I knew if it was true: that I anchored the design on the best available evidence and said in advance what would prove me wrong. That’s a small discipline, and it’s the one that actually transfers to a real job, where I won’t have unlimited research budget there either, and the difference between good judgment and a guess is usually whether you wrote down your reasoning before you knew the outcome.
The thing I’d do differently next time is run the synthetic pilot’s blinding fix from day one instead of catching it in review. I got lucky that the flaw was small enough to log and correct without touching any real data.
What I’d test next
Whether a real supportive-housing case manager’s first three moves on the attention view actually match the group headers, without ever reading them out loud — the single test this whole prototype was built to fail or pass.
Whether the reconciliation-cadence hypothesis from the synthetic pilot holds: does the “rebuild plan” rhythm really track contact cadence, or was that an artifact of how I built the personas.
Whether the counterparty-named waiting board still makes sense in a setting further from supportive housing.
Whether provenance and stated blind spots actually build trust in a real session, or whether they just add friction nobody asked for — the honest risk of any explainability feature is that it’s easy to design for a hiring panel and hard to validate with someone who just wants to get through their caseload.
Have questions about the workflow, the prototype, or the design decisions? Happy to walk through it live.