Sign InOpen Brain
arXivPaperNeeds Review

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

P-Bench tests whether agents choose statistically valid methods, not merely execute code. Fisher-R1-14B improved single-trial success over DeepSeek-V4-Pro by 21% on average.

arXiv
Open Source Open MarkdownOpen JSON
Source Summary

P-Bench contains **425 open-ended tasks** across economics, biology, and medicine. Each asks an agent to select a statistical test, compute a p-value, and conclude from a hypothesis and dataset; Fisher-R1-14B is trained with synthetic tasks and verified statistical rewards.

Practical Implication

If agents analyze data, evaluate whether their assumptions justify the reported statistic. Correctly running code is insufficient when the chosen test or p-value is invalid; checks should cover method selection and inference.

Agent-Ready Context
P-Bench contains **425 open-ended tasks** across economics, biology, and medicine. Each asks an agent to select a statistical test, compute a p-value, and conclude from a hypothesis and dataset; Fisher-R1-14B is trained with synthetic tasks and verified statistical rewards.

If agents analyze data, evaluate whether their assumptions justify the reported statistic. Correctly running code is insufficient when the chosen test or p-value is invalid; checks should cover method selection and inference.

On P-Bench, Fisher-R1-14B reported a **21% average relative improvement** in single-trial success over DeepSeek-V4-Pro, reaching **up to 26%** on the hardest tasks. These are benchmark results, not evidence of reliability on every scientific workflow.
Context Map
benchmarkresearchdata#agent-evals#benchmark-integrity#agent-reliability
Uncertainty
On P-Bench, Fisher-R1-14B reported a **21% average relative improvement** in single-trial success over DeepSeek-V4-Pro, reaching **up to 26%** on the hardest tasks. These are benchmark results, not evidence of reliability on every scientific workflow.