Don’t be data poor — Anuj Iravane, Anterior
When production data cannot be retained, generate eval cases backward from sampled labels and reasoning paths, build records in layers, and let domain experts steer the pipeline.
Anterior evaluates healthcare agents against records that are sensitive, varied, and sometimes **over 300 pages**; contracts can prohibit retaining even redacted derivatives. Its pipeline samples a label and policy reasoning path, then generates a patient journey and encounter documents from coarse to fine.
Reverse the inference workflow to control scenario diversity, emulate how the source documents arise, and round-trip generated records against their starting labels. Give domain experts control at each generation stage and package their reusable guidance as skills rather than leaving dataset design solely to AI engineers.
Anterior evaluates healthcare agents against records that are sensitive, varied, and sometimes **over 300 pages**; contracts can prohibit retaining even redacted derivatives. Its pipeline samples a label and policy reasoning path, then generates a patient journey and encounter documents from coarse to fine. Reverse the inference workflow to control scenario diversity, emulate how the source documents arise, and round-trip generated records against their starting labels. Give domain experts control at each generation stage and package their reusable guidance as skills rather than leaving dataset design solely to AI engineers. Anterior says **roughly 90% of its datasets** are synthetic, while clinicians distinguished synthetic from real records **about 60% of the time** in a blind review. That is a fidelity signal, not proof that the generated distribution captures production prevalence, rare failures, or clinical validity.
This adds a concrete way to build healthcare eval data when privacy and retention rules make real records unusable: generate documents backward from controlled labels and reasoning paths, then round-trip them for consistency. It also places dataset design with clinicians through reusable skills. The blind-review result supports surface fidelity only, leaving production coverage and clinical validity unresolved.