Dutch Books for Language Models
A label-free Dutch-book test finds internally inconsistent probabilities from language models, especially when prompts add logical complexity or irrelevant context.
The authors elicit forecasts for events derived from stock-return data, then use linear programs to find guaranteed betting profit against the model. The test needs **no outcome labels**, so it can assess unresolved events.
Treat agent-generated probabilities as claims to validate, not a coherent world model. Strip irrelevant prompt details and test related forecasts together; the paper reports that such details can raise incoherence by **an order of magnitude**.
The authors elicit forecasts for events derived from stock-return data, then use linear programs to find guaranteed betting profit against the model. The test needs **no outcome labels**, so it can assess unresolved events. Treat agent-generated probabilities as claims to validate, not a coherent world model. Strip irrelevant prompt details and test related forecasts together; the paper reports that such details can raise incoherence by **an order of magnitude**. This measures internal coherence, not whether forecasts are calibrated or accurate. The supplied abstract also does not identify the evaluated models or provide per-model results.
This adds a label-free reliability test for unresolved forecasts: evaluate whether related probabilities admit a guaranteed-loss betting strategy, independently of eventual accuracy or calibration. It therefore complements task-success and domain-validity benchmarks rather than replacing them. The reported sensitivity to irrelevant details also makes prompt variation part of coherence testing.