Sign InOpen Brain
arXivPaperNeeds Review

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

onepot-Bench 0 evaluates chemistry models with three lab-oriented tests, including private experimental data to reduce contamination risk. It is a useful eval design pattern, though results and reproducibility are absent here.

arXiv · Aug 3, 2026
Open Source Open MarkdownOpen JSON
Source Summary

**onepot-Bench 0** contains **three evaluations** for synthetic chemistry: ChemAbacus tests tool-free cheminformatics and numerical reasoning, SynthRefusal examines safety behavior, and SynthBench uses private lab data for reaction prediction and catalyst selection.

Practical Implication

Builders evaluating research agents should test distinct failure modes rather than collapse competence, safety, and domain judgment into one score. Private, task-relevant data can also reduce the risk that benchmark examples appeared in model training corpora.

Agent-Ready Context
**onepot-Bench 0** contains **three evaluations** for synthetic chemistry: ChemAbacus tests tool-free cheminformatics and numerical reasoning, SynthRefusal examines safety behavior, and SynthBench uses private lab data for reaction prediction and catalyst selection.

Builders evaluating research agents should test distinct failure modes rather than collapse competence, safety, and domain judgment into one score. Private, task-relevant data can also reduce the risk that benchmark examples appeared in model training corpora.

The supplied material reports the benchmark design but no model scores, dataset size, protocol detail, or access terms. Because SynthBench relies on proprietary experiments, independent reproduction and contamination auditing remain open questions.
Connected Context · Feed7 Judgment

This narrows benchmark-integrity guidance for scientific agents into three separately scored concerns: computational competence, safety refusal, and lab-grounded judgment. Its private experimental data strengthens contamination resistance, but also trades away some reproducibility and auditability; without scores or protocols, it establishes an evaluation design rather than evidence that any system performs well.

Verifiable Environments for AI in Biology — Kenny Workman, LatchBioBoth argue that scientific-agent evaluation must reflect real experimental analysis and distinct valid paths; the biology case additionally warns that brittle deterministic grading may reject legitimate alternatives.When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AIonepot-Bench’s private lab tasks address the contamination concern behind distrust of leaderboard results, while its undisclosed protocol and proprietary data preserve the candidate’s concerns about inspectability and verifier coverage.DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, DatacurveBoth use original, non-public task material to reduce training contamination, but each remains limited as a general capability proxy because its domain and task coverage are narrow.
Context Map
benchmarkresearch#agent-evals#benchmark-integrity#agent-reliability
Uncertainty
The supplied material reports the benchmark design but no model scores, dataset size, protocol detail, or access terms. Because SynthBench relies on proprietary experiments, independent reproduction and contamination auditing remain open questions.