Sign InOpen Brain
arXivPaperNeeds Review

Dutch Books for Language Models

A label-free Dutch-book test finds internally inconsistent probabilities from language models, especially when prompts add logical complexity or irrelevant context.

arXiv · Sep 2, 2026
Open Source Open MarkdownOpen JSON
Source Summary

The authors elicit forecasts for events derived from stock-return data, then use linear programs to find guaranteed betting profit against the model. The test needs **no outcome labels**, so it can assess unresolved events.

Practical Implication

Treat agent-generated probabilities as claims to validate, not a coherent world model. Strip irrelevant prompt details and test related forecasts together; the paper reports that such details can raise incoherence by **an order of magnitude**.

Agent-Ready Context
The authors elicit forecasts for events derived from stock-return data, then use linear programs to find guaranteed betting profit against the model. The test needs **no outcome labels**, so it can assess unresolved events.

Treat agent-generated probabilities as claims to validate, not a coherent world model. Strip irrelevant prompt details and test related forecasts together; the paper reports that such details can raise incoherence by **an order of magnitude**.

This measures internal coherence, not whether forecasts are calibrated or accurate. The supplied abstract also does not identify the evaluated models or provide per-model results.
Connected Context · Feed7 Judgment

This adds a label-free reliability test for unresolved forecasts: evaluate whether related probabilities admit a guaranteed-loss betting strategy, independently of eventual accuracy or calibration. It therefore complements task-success and domain-validity benchmarks rather than replacing them. The reported sensitivity to irrelevant details also makes prompt variation part of coherence testing.

SocietyBench: Forecasting Counterfactual Social-World EvolutionSocietyBench separates calibration from temporal accuracy; Dutch-book testing adds a third, distinct property—internal coherence among forecasts—and can assess it before outcomes resolve.Fisher-R1: Training LLM Agents for Reliable Hypothesis TestingFisher-R1 checks whether an agent selects statistically valid methods, whereas this test checks whether its probability claims are mutually coherent; together they cover different failures in quantitative reasoning.MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMsMedPRESS shows judgment can shift under escalating pressure, while this paper reports incoherence rising with irrelevant prompt details; both support evaluating reliability across prompt variations rather than a single static formulation.
Context Map
benchmarkresearchdata#agent-evals#agent-reliability
Uncertainty
This measures internal coherence, not whether forecasts are calibrated or accurate. The supplied abstract also does not identify the evaluated models or provide per-model results.