# Dutch Books for Language Models

Source: [arXiv](https://arxiv.org/abs/2609.02797v1)  
Feed7 permalink: https://feed7.dev/p/2609-02797v1-15ybio3  
Published: 2026-09-02T16:31:47.000Z  
Trust: Needs Review (needs_review)

## Why Included

A label-free Dutch-book test finds internally inconsistent probabilities from language models, especially when prompts add logical complexity or irrelevant context.

## Source Summary

The authors elicit forecasts for events derived from stock-return data, then use linear programs to find guaranteed betting profit against the model. The test needs **no outcome labels**, so it can assess unresolved events.

## Practical Implication

Treat agent-generated probabilities as claims to validate, not a coherent world model. Strip irrelevant prompt details and test related forecasts together; the paper reports that such details can raise incoherence by **an order of magnitude**.

## Agent-Ready Context

The authors elicit forecasts for events derived from stock-return data, then use linear programs to find guaranteed betting profit against the model. The test needs **no outcome labels**, so it can assess unresolved events.

Treat agent-generated probabilities as claims to validate, not a coherent world model. Strip irrelevant prompt details and test related forecasts together; the paper reports that such details can raise incoherence by **an order of magnitude**.

This measures internal coherence, not whether forecasts are calibrated or accurate. The supplied abstract also does not identify the evaluated models or provide per-model results.

## Connected Context

Feed7 judgment across 669 accumulated Signals:

This adds a label-free reliability test for unresolved forecasts: evaluate whether related probabilities admit a guaranteed-loss betting strategy, independently of eventual accuracy or calibration. It therefore complements task-success and domain-validity benchmarks rather than replacing them. The reported sensitivity to irrelevant details also makes prompt variation part of coherence testing.

- [SocietyBench: Forecasting Counterfactual Social-World Evolution](https://feed7.dev/p/2608-04009v1-05m5u8w) — SocietyBench separates calibration from temporal accuracy; Dutch-book testing adds a third, distinct property—internal coherence among forecasts—and can assess it before outcomes resolve.
- [Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing](https://feed7.dev/p/2608-07437v1-0t4v96m) — Fisher-R1 checks whether an agent selects statistically valid methods, whereas this test checks whether its probability claims are mutually coherent; together they cover different failures in quantitative reasoning.
- [MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs](https://feed7.dev/p/2608-02520v1-16negmk) — MedPRESS shows judgment can shift under escalating pressure, while this paper reports incoherence rising with irrelevant prompt details; both support evaluating reliability across prompt variations rather than a single static formulation.

## Context Map

- Layer: benchmark
- Domains: research, data
- Topics: agent-evals, agent-reliability

## Uncertainty

- This measures internal coherence, not whether forecasts are calibrated or accurate. The supplied abstract also does not identify the evaluated models or provide per-model results.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
