A Living Benchmark for Information Retrieval from Electronic Health Records
BRIE generates refreshable EHR retrieval evaluations from longitudinal notes, addressing benchmark staleness and leakage while exposing omissions in multi-document clinical synthesis.
BRIE automatically generates question-answer pairs from longitudinal health records, with its generator validated by **19 clinicians**. The study evaluates **nine LLMs** using **five inference strategies**.
Builders of retrieval agents can borrow the living-benchmark pattern: validate the generation process, refresh cases to reduce leakage, and allow multiple reference answers where expert reasoning legitimately varies.
BRIE automatically generates question-answer pairs from longitudinal health records, with its generator validated by **19 clinicians**. The study evaluates **nine LLMs** using **five inference strategies**. Builders of retrieval agents can borrow the living-benchmark pattern: validate the generation process, refresh cases to reduce leakage, and allow multiple reference answers where expert reasoning legitimately varies. The evaluated systems frequently omitted important clinical details, especially when answers required synthesis across documents and encounters. The evidence is specific to electronic health records, where omissions carry unusually high stakes.
This turns several retrieval-evaluation principles into a renewable clinical benchmark design: validate generation with experts, refresh cases against leakage, and accept legitimate answer plurality. Its omission findings make cross-document synthesis a concrete failure mode rather than a generic retrieval concern, while the EHR setting limits direct performance generalization and raises the cost of incomplete answers.