Sign InOpen Brain
arXivPaperNeeds Review

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Molecular benchmark accuracy can reflect retrieval of published values rather than prediction. Contamination audits should test digit-level recall and repeat runs across reasoning settings.

arXiv · Sep 4, 2026
Open Source Open MarkdownOpen JSON
Source Summary

An audit of **22 frontier models** across **12 molecular regression benchmarks** found verbatim value retrieval concentrated by dataset. On five datasets, **more than 50% of models** showed it; elsewhere it appeared only in isolated cases.

Practical Implication

Builders evaluating scientific agents should separate prediction from memorized lookup, perturb representations and labels, and test multiple reasoning settings. A high benchmark score alone cannot establish general molecular prediction ability.

Agent-Ready Context
An audit of **22 frontier models** across **12 molecular regression benchmarks** found verbatim value retrieval concentrated by dataset. On five datasets, **more than 50% of models** showed it; elsewhere it appeared only in isolated cases.

Builders evaluating scientific agents should separate prediction from memorized lookup, perturb representations and labels, and test multiple reasoning settings. A high benchmark score alone cannot establish general molecular prediction ability.

With the same molecules and prompts, higher reasoning was flagged for retrieval **89% more often** than the lowest setting. Even transformed SMILES strings did not always interrupt recognition, and the study does not claim memorization alone determines predictive capability.
Connected Context · Feed7 Judgment

This identifies memorized value retrieval as a dataset-dependent confound in molecular benchmarks and shows that higher reasoning settings may amplify rather than remove it. It strengthens the case for controlled exposure and leakage-resistant evaluation, while warning that representation changes alone may not separate lookup from prediction and that retrieval evidence does not by itself negate predictive ability.

WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live TournamentWorldCup Arena avoids answer leakage prospectively; this Signal shows why repairing public benchmarks with prompt or representation perturbations may be insufficient once values are recognizable.LittleLearner: Language Models Under Pedagogically Controlled Knowledge ExposureLittleLearner controls prior knowledge exposure to distinguish acquisition from reuse, while this audit detects reuse of published benchmark values in frontier models whose exposure is unknown.Beyond Scale and Generation: Understanding Language Model-based Entity MatchingBoth make evaluation configuration part of model selection: architecture and variant affect entity matching, while reasoning level affects the measured incidence of molecular-value retrieval.
Context Map
benchmarkresearchdata#benchmark-integrity#reasoning#model-selection
Uncertainty
With the same molecules and prompts, higher reasoning was flagged for retrieval **89% more often** than the lowest setting. Even transformed SMILES strings did not always interrupt recognition, and the study does not claim memorization alone determines predictive capability.