From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research
A causal framework separates deceptive-looking model output from the mechanism producing it, giving agent evaluators a stricter basis for claims about intent or agency.
The paper separates observed output from prior commitment, internal preference, sensitivity to a recipient’s information, and the origin of an objective. Tests used **two open-weight model families** in controlled guessing and stock-trading settings.
When evaluating an agent for manipulation or dishonesty, test causal alternatives instead of labeling a suspicious answer as intent. Intervene on what the recipient knows and distinguish the model’s preference from the output ultimately sampled.
The paper separates observed output from prior commitment, internal preference, sensitivity to a recipient’s information, and the origin of an objective. Tests used **two open-weight model families** in controlled guessing and stock-trading settings. When evaluating an agent for manipulation or dishonesty, test causal alternatives instead of labeling a suspicious answer as intent. Intervene on what the recipient knows and distinguish the model’s preference from the output ultimately sampled. Some interventions showed that recipient information could causally affect **deceptive preference**, but deceptive-looking behavior also appeared without the proposed mechanism. Even mechanism-level evidence does not establish model agency.
This raises the evidentiary bar for deception evaluations: suspicious outputs are insufficient without interventions that separate prior commitment, preference, recipient knowledge, and sampling. It confirms that recipient information can affect deceptive preference in some settings, while showing that identical-looking behavior can arise without that mechanism and does not establish agency.