Healthcare · Clinical summarization
Lumen Health ships AI changes with evidence
How Lumen Health built a 400-example clinical golden set and gated every model change on clinician-calibrated evaluations.
Results
- Judge agreement with clinician labels
- 91%
- Clinician-labeled golden examples
- 400
- More releases per month
- 4x
- Down from 2 weeks for clinical review
- 2 hrs
The challenge
Lumen Health summarizes clinical notes for care coordinators. A wrong summary is a patient safety issue, so every model change needed clinician sign-off. That review took two weeks and did not scale past one release a month.
What they did
Clinicians labeled 400 summaries across the failure modes they cared about: omitted medications, wrong dosage, invented findings. Cloudmind judges were calibrated against those labels until agreement passed 91%. Now the judge runs on every change and clinicians only review the disagreements.
“I stopped asking engineers whether the summaries got better. I just look at the scorecard on the release. That changed my job more than any other tool this year.”
Stop guessing whether your AI feature got better
Free for 50,000 traces a month. No credit card, no sales call, five minutes to your first trace.
Start free