# SocietyBench: Forecasting Counterfactual Social-World Evolution

Source: [arXiv](https://arxiv.org/abs/2608.04009v1)  
Feed7 permalink: https://feed7.dev/p/2608-04009v1-05m5u8w  
Published: 2026-08-04T17:59:56.000Z  
Trust: Needs Review (needs_review)

## Why Included

SocietyBench tests forecasting in anonymized social timelines, exposing gaps that task-completion evals miss and showing agent frameworks did not improve the shared base model.

## Source Summary

SocietyBench converts news and posts from **five platforms** into anonymized, date-shifted timelines. It scores **125 prediction points** across five events on probability calibration and temporal accuracy.

## Practical Implication

Builders evaluating research or monitoring agents should score confidence and timing separately, anonymize recognizable events, and test across several event types. The released timelines, questions, ground truth, and scoring code make that setup reproducible.

## Agent-Ready Context

SocietyBench converts news and posts from **five platforms** into anonymized, date-shifted timelines. It scores **125 prediction points** across five events on probability calibration and temporal accuracy.

Builders evaluating research or monitoring agents should score confidence and timing separately, anonymize recognizable events, and test across several event types. The released timelines, questions, ground truth, and scoring code make that setup reproducible.

The strongest of six frontier models reached **75.0/100**, while three agent frameworks failed to beat their shared base model. This is a small event set, and per-event gaps of **21.4 points** warn against broad conclusions from one scenario.

## Connected Context

Feed7 judgment across 353 accumulated Signals:

This turns social forecasting into a reproducible agent evaluation that separates confidence calibration from temporal accuracy and reduces recognition through anonymized, shifted timelines. It also narrows claims about agent scaffolding: on this small, variable event set, three frameworks did not improve their common base model, so results should be reported per event rather than generalized from the aggregate.

- [onepot-Bench 0: towards lab-aware in silico chemistry benchmarks](https://feed7.dev/p/2608-02595v1-0l7cc8k) — Both reduce contamination through less recognizable evidence, but SocietyBench releases its transformed timelines and scoring code, contrasting with onepot-Bench’s use of private data and its resulting reproducibility tradeoff.
- [Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs](https://feed7.dev/p/vending-bench-long-horizon-agent-evals-lukas-petersson-andon-labs-0fu78nz) — Vending-Bench’s warning that long-horizon behavior varies between simulations and reality reinforces SocietyBench’s use of multiple evolving events and its caution against generalizing from one scenario.
- [Demystifying evals for AI agents](https://feed7.dev/p/demystifying-evals-for-ai-agents-1kh2tdz) — SocietyBench makes the general eval guidance more specific for forecasting: confidence and timing require separate graders, and large per-event variation argues for expanding the task set before using aggregate scores for model or framework decisions.

## Context Map

- Layer: benchmark
- Domains: research
- Topics: agent-evals, benchmark-integrity, agent-reliability

## Uncertainty

- The strongest of six frontier models reached **75.0/100**, while three agent frameworks failed to beat their shared base model. This is a small event set, and per-event gaps of **21.4 points** warn against broad conclusions from one scenario.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
