onepot-Bench 0: towards lab-aware in silico chemistry benchmarks
onepot-Bench 0 evaluates chemistry models with three lab-oriented tests, including private experimental data to reduce contamination risk. It is a useful eval design pattern, though results and reproducibility are absent here.
**onepot-Bench 0** contains **three evaluations** for synthetic chemistry: ChemAbacus tests tool-free cheminformatics and numerical reasoning, SynthRefusal examines safety behavior, and SynthBench uses private lab data for reaction prediction and catalyst selection.
Builders evaluating research agents should test distinct failure modes rather than collapse competence, safety, and domain judgment into one score. Private, task-relevant data can also reduce the risk that benchmark examples appeared in model training corpora.
**onepot-Bench 0** contains **three evaluations** for synthetic chemistry: ChemAbacus tests tool-free cheminformatics and numerical reasoning, SynthRefusal examines safety behavior, and SynthBench uses private lab data for reaction prediction and catalyst selection. Builders evaluating research agents should test distinct failure modes rather than collapse competence, safety, and domain judgment into one score. Private, task-relevant data can also reduce the risk that benchmark examples appeared in model training corpora. The supplied material reports the benchmark design but no model scores, dataset size, protocol detail, or access terms. Because SynthBench relies on proprietary experiments, independent reproduction and contamination auditing remain open questions.
This narrows benchmark-integrity guidance for scientific agents into three separately scored concerns: computational competence, safety refusal, and lab-grounded judgment. Its private experimental data strengthens contamination resistance, but also trades away some reproducibility and auditability; without scores or protocols, it establishes an evaluation design rather than evidence that any system performs well.