Sign InOpen Brain
arXivPaperNeeds Review

A Living Benchmark for Information Retrieval from Electronic Health Records

BRIE generates refreshable EHR retrieval evaluations from longitudinal notes, addressing benchmark staleness and leakage while exposing omissions in multi-document clinical synthesis.

arXiv · Sep 24, 2026
Open Source Open MarkdownOpen JSON
Source Summary

BRIE automatically generates question-answer pairs from longitudinal health records, with its generator validated by **19 clinicians**. The study evaluates **nine LLMs** using **five inference strategies**.

Practical Implication

Builders of retrieval agents can borrow the living-benchmark pattern: validate the generation process, refresh cases to reduce leakage, and allow multiple reference answers where expert reasoning legitimately varies.

Agent-Ready Context
BRIE automatically generates question-answer pairs from longitudinal health records, with its generator validated by **19 clinicians**. The study evaluates **nine LLMs** using **five inference strategies**.

Builders of retrieval agents can borrow the living-benchmark pattern: validate the generation process, refresh cases to reduce leakage, and allow multiple reference answers where expert reasoning legitimately varies.

The evaluated systems frequently omitted important clinical details, especially when answers required synthesis across documents and encounters. The evidence is specific to electronic health records, where omissions carry unusually high stakes.
Connected Context · Feed7 Judgment

This turns several retrieval-evaluation principles into a renewable clinical benchmark design: validate generation with experts, refresh cases against leakage, and accept legitimate answer plurality. Its omission findings make cross-document synthesis a concrete failure mode rather than a generic retrieval concern, while the EHR setting limits direct performance generalization and raises the cost of incomplete answers.

Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question AnsweringClaim-level coverage and entailment checks offer a direct way to distinguish BRIE’s clinically important omissions from answers that merely contain accurate cited fragments.Inside 847 Production Clinical AI Notes — Sebastian Fox, ComposoProduction clinical notes independently reinforce omissions as a consequential failure mode and support BRIE’s use of clinician judgment rather than generic model grading alone.Verifiable Environments for AI in Biology — Kenny Workman, LatchBioBoth treat expert validation and multiple valid reasoning paths as necessary constraints on domain benchmarks, rather than assuming one brittle reference or grader captures correctness.Eval awareness in Claude Opus 4.6’s BrowseComp performanceBrowseComp leakage demonstrates the benchmark-integrity risk that BRIE’s refreshed case generation is designed to reduce, though it does not establish that refreshes eliminate leakage.
Context Map
benchmarkresearchdata#retrieval#agent-evals#benchmark-integrity
Uncertainty
The evaluated systems frequently omitted important clinical details, especially when answers required synthesis across documents and encounters. The evidence is specific to electronic health records, where omissions carry unusually high stakes.