# WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament

Source: [arXiv](https://arxiv.org/abs/2608.04008v1)  
Feed7 permalink: https://feed7.dev/p/2608-04008v1-0d8bqwv  
Published: 2026-08-04T17:59:55.000Z  
Trust: Needs Review (needs_review)

## Why Included

A live, pre-kickoff benchmark removes answer leakage by construction and finds six frontier models clustered near a bookmaker-favorite baseline, with no gain from majority voting.

## Source Summary

Across the **39-day 2026 World Cup**, six web-enabled reasoning models made predictions before kickoff for all **104 matches**. The frozen archive contains **4,494 scored predictions**, so answers could not have appeared in training or search results.

## Practical Implication

For agent evals, prospective collection is a cleaner leakage control than filtering retrospective questions. Also test disagreement quality before using model ensembles: these systems agreed more often than they were correct, so majority voting added nothing.

## Agent-Ready Context

Across the **39-day 2026 World Cup**, six web-enabled reasoning models made predictions before kickoff for all **104 matches**. The frozen archive contains **4,494 scored predictions**, so answers could not have appeared in training or search results.

For agent evals, prospective collection is a cleaner leakage control than filtering retrospective questions. Also test disagreement quality before using model ensembles: these systems agreed more often than they were correct, so majority voting added nothing.

Average match-outcome accuracy was **63.9%**, level with backing the bookmaker favorite. Results cover one sports tournament, models were narrowly separated, and accuracy fell most on close fixtures despite richer dossiers.

## Connected Context

Feed7 judgment across 353 accumulated Signals:

This strengthens leakage-control guidance by showing that predictions recorded before future outcomes provide a cleaner test than repairing already-public benchmarks. It also weakens two common selection shortcuts: the models only matched a simple bookmaker-favorite baseline on average, and correlated errors meant majority voting offered no gain. The evidence remains limited to one tournament with closely grouped systems.

- [Eval awareness in Claude Opus 4.6’s BrowseComp performance](https://feed7.dev/p/eval-awareness-browsecomp-1q6k277) — BrowseComp shows how a web-enabled model can locate leaked evaluation answers; WorldCup Arena avoids that failure mode by freezing predictions before the answers exist.
- [SocietyBench: Forecasting Counterfactual Social-World Evolution](https://feed7.dev/p/2608-04009v1-05m5u8w) — Both evaluate forecasting while controlling event recognition, but SocietyBench transforms historical timelines whereas WorldCup Arena collects predictions prospectively; together they provide complementary reproducible leakage controls.
- [Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility](https://feed7.dev/p/2608-04001v1-1gj91hk) — WorldCup Arena’s frozen archive supplies replayable prediction evidence, while the test-time-scaling guidance explains what remains necessary for fair comparison: reporting each model’s complete inference and compute protocol.
- [Reward hacking is swamping model intelligence gains](https://feed7.dev/p/reward-hacking-coding-benchmarks-18ddebo) — Cursor’s sealed harness reduces access to existing fixes, while WorldCup Arena goes further by evaluating answers before ground truth exists; both show that model rankings can change meaningfully when answer access is structurally constrained.

## Context Map

- Layer: benchmark
- Domains: research
- Topics: agent-evals, benchmark-integrity, model-selection

## Uncertainty

- Average match-outcome accuracy was **63.9%**, level with backing the bookmaker favorite. Results cover one sports tournament, models were narrowly separated, and accuracy fell most on close fixtures despite richer dossiers.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
