WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament
A live, pre-kickoff benchmark removes answer leakage by construction and finds six frontier models clustered near a bookmaker-favorite baseline, with no gain from majority voting.
Across the **39-day 2026 World Cup**, six web-enabled reasoning models made predictions before kickoff for all **104 matches**. The frozen archive contains **4,494 scored predictions**, so answers could not have appeared in training or search results.
For agent evals, prospective collection is a cleaner leakage control than filtering retrospective questions. Also test disagreement quality before using model ensembles: these systems agreed more often than they were correct, so majority voting added nothing.
Across the **39-day 2026 World Cup**, six web-enabled reasoning models made predictions before kickoff for all **104 matches**. The frozen archive contains **4,494 scored predictions**, so answers could not have appeared in training or search results. For agent evals, prospective collection is a cleaner leakage control than filtering retrospective questions. Also test disagreement quality before using model ensembles: these systems agreed more often than they were correct, so majority voting added nothing. Average match-outcome accuracy was **63.9%**, level with backing the bookmaker favorite. Results cover one sports tournament, models were narrowly separated, and accuracy fell most on close fixtures despite richer dossiers.
This strengthens leakage-control guidance by showing that predictions recorded before future outcomes provide a cleaner test than repairing already-public benchmarks. It also weakens two common selection shortcuts: the models only matched a simple bookmaker-favorite baseline on average, and correlated errors meant majority voting offered no gain. The evidence remains limited to one tournament with closely grouped systems.