arXivPaperNeeds Review
Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
P-Bench tests whether agents choose statistically valid methods, not merely execute code. Fisher-R1-14B improved single-trial success over DeepSeek-V4-Pro by 21% on average.
arXiv
Source Summary
P-Bench contains **425 open-ended tasks** across economics, biology, and medicine. Each asks an agent to select a statistical test, compute a p-value, and conclude from a hypothesis and dataset; Fisher-R1-14B is trained with synthetic tasks and verified statistical rewards.
Practical Implication
If agents analyze data, evaluate whether their assumptions justify the reported statistic. Correctly running code is insufficient when the chosen test or p-value is invalid; checks should cover method selection and inference.
Agent-Ready Context
P-Bench contains **425 open-ended tasks** across economics, biology, and medicine. Each asks an agent to select a statistical test, compute a p-value, and conclude from a hypothesis and dataset; Fisher-R1-14B is trained with synthetic tasks and verified statistical rewards. If agents analyze data, evaluate whether their assumptions justify the reported statistic. Correctly running code is insufficient when the chosen test or p-value is invalid; checks should cover method selection and inference. On P-Bench, Fisher-R1-14B reported a **21% average relative improvement** in single-trial success over DeepSeek-V4-Pro, reaching **up to 26%** on the hardest tasks. These are benchmark results, not evidence of reliability on every scientific workflow.
Context Map
benchmarkresearchdata#agent-evals#benchmark-integrity#agent-reliabilityUncertainty
On P-Bench, Fisher-R1-14B reported a **21% average relative improvement** in single-trial success over DeepSeek-V4-Pro, reaching **up to 26%** on the hardest tasks. These are benchmark results, not evidence of reliability on every scientific workflow.