# Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo

Source: [AI Engineer](https://www.youtube.com/watch?v=yqF6XhzbWBk)  
Feed7 permalink: https://feed7.dev/p/inside-847-production-clinical-ai-notes-sebastian-fox-composo-0lrc8td  
Published: 2026-08-22T17:00:32.000Z  
Trust: Source Linked (source_linked)

## Why Included

Plausible outputs can hide consequential omissions that generic LLM judges miss. Production evals need real failure discovery and retrieved expert judgments, not a frozen rubric alone.

## Source Summary

In the cited production study, **about 1 in 20 notes** contained an error capable of significant harm, nearly 1 in 5 had an important omission, and more than 1 in 10 hallucinated. A strong judge still passed notes containing serious errors.

## Practical Implication

For any agent where confident mistakes carry real cost, inspect production outputs, organize recurring failure modes, capture expert corrections and reasoning, then retrieve relevant past judgments into each evaluation. Start with experts commenting freely on real cases.

## Agent-Ready Context

In the cited production study, **about 1 in 20 notes** contained an error capable of significant harm, nearly 1 in 5 had an important omission, and more than 1 in 10 hallucinated. A strong judge still passed notes containing serious errors.

For any agent where confident mistakes carry real cost, inspect production outputs, organize recurring failure modes, capture expert corrections and reasoning, then retrieve relevant past judgments into each evaluation. Start with experts commenting freely on real cases.

This approach depends on sustained expert review and representative production data. It improves the judge’s domain context, but the talk does not establish that retrieval catches every novel failure or removes the need for human oversight.

## Connected Context

Feed7 judgment across 545 accumulated Signals:

This supplies production evidence that strong model judges can approve clinically dangerous outputs, tightening the case against self-grading in high-cost domains. It turns expert review into reusable evaluation context: collect real failures, preserve corrections and reasoning, and retrieve relevant judgments per case. That improves domain grounding but does not automate away expert oversight or guarantee coverage of novel failures.

- [Trading Desks to Clinical Trials: Parallels in Applied Vertical AI — Ayush Bhardwaj, Allos AI](https://feed7.dev/p/trading-desks-to-clinical-trials-parallels-in-applied-vertical-ai-ayush-1hwvsg3) — The clinical error findings provide concrete production support for the claim that generic models and self-grading cannot replace expert causal judgment in vertical agents.
- [Demystifying evals for AI agents](https://feed7.dev/p/demystifying-evals-for-ai-agents-1kh2tdz) — Anthropic recommends starting with tasks drawn from real failures; this signal extends that recipe by capturing expert reasoning and retrieving related past judgments into each evaluation.
- [Verifiable Environments for AI in Biology — Kenny Workman, LatchBio](https://feed7.dev/p/verifiable-environments-for-ai-in-biology-kenny-workman-latchbio-1vs6y66) — Both show that plausible outputs and brittle graders can miss domain-critical errors, reinforcing human validation even when evaluation includes deterministic or model-based checks.
- [Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta Software](https://feed7.dev/p/rethinking-environments-for-long-horizon-work-rayan-garg-theta-software-11r7wbx) — The clinical results strengthen the warning that flexible model judges are fallible; inspecting production outputs and expert corrections complements final-state and trajectory-aware evaluation.

## Context Map

- Layer: benchmark
- Domains: data
- Topics: agent-evals, agent-reliability, retrieval

## Uncertainty

- This approach depends on sustained expert review and representative production data. It improves the judge’s domain context, but the talk does not establish that retrieval catches every novel failure or removes the need for human oversight.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
