Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
Molecular benchmark accuracy can reflect retrieval of published values rather than prediction. Contamination audits should test digit-level recall and repeat runs across reasoning settings.
An audit of **22 frontier models** across **12 molecular regression benchmarks** found verbatim value retrieval concentrated by dataset. On five datasets, **more than 50% of models** showed it; elsewhere it appeared only in isolated cases.
Builders evaluating scientific agents should separate prediction from memorized lookup, perturb representations and labels, and test multiple reasoning settings. A high benchmark score alone cannot establish general molecular prediction ability.
An audit of **22 frontier models** across **12 molecular regression benchmarks** found verbatim value retrieval concentrated by dataset. On five datasets, **more than 50% of models** showed it; elsewhere it appeared only in isolated cases. Builders evaluating scientific agents should separate prediction from memorized lookup, perturb representations and labels, and test multiple reasoning settings. A high benchmark score alone cannot establish general molecular prediction ability. With the same molecules and prompts, higher reasoning was flagged for retrieval **89% more often** than the lowest setting. Even transformed SMILES strings did not always interrupt recognition, and the study does not claim memorization alone determines predictive capability.
This identifies memorized value retrieval as a dataset-dependent confound in molecular benchmarks and shows that higher reasoning settings may amplify rather than remove it. It strengthens the case for controlled exposure and leakage-resistant evaluation, while warning that representation changes alone may not separate lookup from prediction and that retrieval evidence does not by itself negate predictive ability.