Research ran with the client and with the people who do this work — reviewers today, and the shift leads and policy owners who would inherit the product. Alongside it I audited the QA tool they already used, walked the existing paths for review, refund approval and policy change, and compared products in AI observability, evaluation and human-in-the-loop approval.
Two findings rewrote the brief. The route for one hard conversation ran QA tool → vendor panel → knowledge base → Confluence → Slack and took an hour and a half against a six-minute norm, so the evidence had to arrive inside the verdict card or the product had no reason to exist. And confident prose turned out to be active camouflage: reviewers cannot spot an unsupported claim by reading, and asked for it to be marked — while ruling out the engineering answer, since a confidence score gets the screen dismissed as noise rather than read.
Then the role model — Reviewer, Shift Lead, Policy Owner — the information architecture, the design system, and 18 of 35 scoped artboards assembled into a clickable prototype. Testing it with users added decisions the brief never had: a cluster needs its members visible and removable or reviewers stop trusting clusters at all, a wrong verdict needs a short window to take back before it poisons the reporting, and a reviewer needs to see what happened to their correction, or they quietly stop classifying causes at all.