Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo
Plausible outputs can hide consequential omissions that generic LLM judges miss. Production evals need real failure discovery and retrieved expert judgments, not a frozen rubric alone.
In the cited production study, **about 1 in 20 notes** contained an error capable of significant harm, nearly 1 in 5 had an important omission, and more than 1 in 10 hallucinated. A strong judge still passed notes containing serious errors.
For any agent where confident mistakes carry real cost, inspect production outputs, organize recurring failure modes, capture expert corrections and reasoning, then retrieve relevant past judgments into each evaluation. Start with experts commenting freely on real cases.
In the cited production study, **about 1 in 20 notes** contained an error capable of significant harm, nearly 1 in 5 had an important omission, and more than 1 in 10 hallucinated. A strong judge still passed notes containing serious errors. For any agent where confident mistakes carry real cost, inspect production outputs, organize recurring failure modes, capture expert corrections and reasoning, then retrieve relevant past judgments into each evaluation. Start with experts commenting freely on real cases. This approach depends on sustained expert review and representative production data. It improves the judge’s domain context, but the talk does not establish that retrieval catches every novel failure or removes the need for human oversight.
This turns expert-led error analysis from general advice into a production-grounded evaluation loop: collect real failures and expert reasoning, then retrieve relevant prior judgments for each new case. The clinical error rates and judge misses sharply limit confidence in model grading alone. Retrieval improves domain context but does not cover novel failures or eliminate ongoing human oversight.