Sign InOpen Brain
AI EngineerVideoSource Linked

Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo

Plausible outputs can hide consequential omissions that generic LLM judges miss. Production evals need real failure discovery and retrieved expert judgments, not a frozen rubric alone.

AI Engineer · Aug 22, 2026
Open Source Open MarkdownOpen JSON
Source Summary

In the cited production study, **about 1 in 20 notes** contained an error capable of significant harm, nearly 1 in 5 had an important omission, and more than 1 in 10 hallucinated. A strong judge still passed notes containing serious errors.

Practical Implication

For any agent where confident mistakes carry real cost, inspect production outputs, organize recurring failure modes, capture expert corrections and reasoning, then retrieve relevant past judgments into each evaluation. Start with experts commenting freely on real cases.

Agent-Ready Context
In the cited production study, **about 1 in 20 notes** contained an error capable of significant harm, nearly 1 in 5 had an important omission, and more than 1 in 10 hallucinated. A strong judge still passed notes containing serious errors.

For any agent where confident mistakes carry real cost, inspect production outputs, organize recurring failure modes, capture expert corrections and reasoning, then retrieve relevant past judgments into each evaluation. Start with experts commenting freely on real cases.

This approach depends on sustained expert review and representative production data. It improves the judge’s domain context, but the talk does not establish that retrieval catches every novel failure or removes the need for human oversight.
Connected Context · Feed7 Judgment

This turns expert-led error analysis from general advice into a production-grounded evaluation loop: collect real failures and expert reasoning, then retrieve relevant prior judgments for each new case. The clinical error rates and judge misses sharply limit confidence in model grading alone. Retrieval improves domain context but does not cover novel failures or eliminate ongoing human oversight.

Trading Desks to Clinical Trials: Parallels in Applied Vertical AI — Ayush Bhardwaj, Allos AIIt confirms with production clinical notes that generic models and self-grading cannot establish domain usefulness, then adds a concrete mechanism for carrying expert corrections into later evaluations.Demystifying evals for AI agentsBoth start eval design from real failures, but this Signal extends the small-task roadmap into a continuously updated memory of expert judgments retrieved per case.Verifiable Environments for AI in Biology — Kenny Workman, LatchBioBoth show that domain experts remain necessary when automated graders miss serious errors or reject valid work; this Signal focuses on production clinical outputs rather than constructed scientific tasks.Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World SystemsThe clinical study shows acceptable-looking outputs can conceal harmful content errors, paralleling the finding that correct answers can conceal invalid computation; both require inspection beyond surface success.
Context Map
benchmarkdata#agent-evals#agent-reliability#retrieval
Uncertainty
This approach depends on sustained expert review and representative production data. It improves the judge’s domain context, but the talk does not establish that retrieval catches every novel failure or removes the need for human oversight.