SocietyBench: Forecasting Counterfactual Social-World Evolution
SocietyBench tests forecasting in anonymized social timelines, exposing gaps that task-completion evals miss and showing agent frameworks did not improve the shared base model.
SocietyBench converts news and posts from **five platforms** into anonymized, date-shifted timelines. It scores **125 prediction points** across five events on probability calibration and temporal accuracy.
Builders evaluating research or monitoring agents should score confidence and timing separately, anonymize recognizable events, and test across several event types. The released timelines, questions, ground truth, and scoring code make that setup reproducible.
SocietyBench converts news and posts from **five platforms** into anonymized, date-shifted timelines. It scores **125 prediction points** across five events on probability calibration and temporal accuracy. Builders evaluating research or monitoring agents should score confidence and timing separately, anonymize recognizable events, and test across several event types. The released timelines, questions, ground truth, and scoring code make that setup reproducible. The strongest of six frontier models reached **75.0/100**, while three agent frameworks failed to beat their shared base model. This is a small event set, and per-event gaps of **21.4 points** warn against broad conclusions from one scenario.
This turns social forecasting into a reproducible agent evaluation that separates confidence calibration from temporal accuracy and reduces recognition through anonymized, shifted timelines. It also narrows claims about agent scaffolding: on this small, variable event set, three frameworks did not improve their common base model, so results should be reported per event rather than generalized from the aggregate.