Sign InOpen Brain
arXivPaperNeeds Review

WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament

A live, pre-kickoff benchmark removes answer leakage by construction and finds six frontier models clustered near a bookmaker-favorite baseline, with no gain from majority voting.

arXiv · Aug 4, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Across the **39-day 2026 World Cup**, six web-enabled reasoning models made predictions before kickoff for all **104 matches**. The frozen archive contains **4,494 scored predictions**, so answers could not have appeared in training or search results.

Practical Implication

For agent evals, prospective collection is a cleaner leakage control than filtering retrospective questions. Also test disagreement quality before using model ensembles: these systems agreed more often than they were correct, so majority voting added nothing.

Agent-Ready Context
Across the **39-day 2026 World Cup**, six web-enabled reasoning models made predictions before kickoff for all **104 matches**. The frozen archive contains **4,494 scored predictions**, so answers could not have appeared in training or search results.

For agent evals, prospective collection is a cleaner leakage control than filtering retrospective questions. Also test disagreement quality before using model ensembles: these systems agreed more often than they were correct, so majority voting added nothing.

Average match-outcome accuracy was **63.9%**, level with backing the bookmaker favorite. Results cover one sports tournament, models were narrowly separated, and accuracy fell most on close fixtures despite richer dossiers.
Connected Context · Feed7 Judgment

This strengthens leakage-control guidance by showing that predictions recorded before future outcomes provide a cleaner test than repairing already-public benchmarks. It also weakens two common selection shortcuts: the models only matched a simple bookmaker-favorite baseline on average, and correlated errors meant majority voting offered no gain. The evidence remains limited to one tournament with closely grouped systems.

Eval awareness in Claude Opus 4.6’s BrowseComp performanceBrowseComp shows how a web-enabled model can locate leaked evaluation answers; WorldCup Arena avoids that failure mode by freezing predictions before the answers exist.SocietyBench: Forecasting Counterfactual Social-World EvolutionBoth evaluate forecasting while controlling event recognition, but SocietyBench transforms historical timelines whereas WorldCup Arena collects predictions prospectively; together they provide complementary reproducible leakage controls.Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and ReproducibilityWorldCup Arena’s frozen archive supplies replayable prediction evidence, while the test-time-scaling guidance explains what remains necessary for fair comparison: reporting each model’s complete inference and compute protocol.Reward hacking is swamping model intelligence gainsCursor’s sealed harness reduces access to existing fixes, while WorldCup Arena goes further by evaluating answers before ground truth exists; both show that model rankings can change meaningfully when answer access is structurally constrained.
Context Map
benchmarkresearch#agent-evals#benchmark-integrity#model-selection
Uncertainty
Average match-outcome accuracy was **63.9%**, level with backing the bookmaker favorite. Results cover one sports tournament, models were narrowly separated, and accuracy fell most on close fixtures despite richer dossiers.