Sign InOpen Brain
arXivPaperNeeds Review

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

A causal framework separates deceptive-looking model output from the mechanism producing it, giving agent evaluators a stricter basis for claims about intent or agency.

arXiv · Sep 3, 2026
Open Source Open MarkdownOpen JSON
Source Summary

The paper separates observed output from prior commitment, internal preference, sensitivity to a recipient’s information, and the origin of an objective. Tests used **two open-weight model families** in controlled guessing and stock-trading settings.

Practical Implication

When evaluating an agent for manipulation or dishonesty, test causal alternatives instead of labeling a suspicious answer as intent. Intervene on what the recipient knows and distinguish the model’s preference from the output ultimately sampled.

Agent-Ready Context
The paper separates observed output from prior commitment, internal preference, sensitivity to a recipient’s information, and the origin of an objective. Tests used **two open-weight model families** in controlled guessing and stock-trading settings.

When evaluating an agent for manipulation or dishonesty, test causal alternatives instead of labeling a suspicious answer as intent. Intervene on what the recipient knows and distinguish the model’s preference from the output ultimately sampled.

Some interventions showed that recipient information could causally affect **deceptive preference**, but deceptive-looking behavior also appeared without the proposed mechanism. Even mechanism-level evidence does not establish model agency.
Connected Context · Feed7 Judgment

This raises the evidentiary bar for deception evaluations: suspicious outputs are insufficient without interventions that separate prior commitment, preference, recipient knowledge, and sampling. It confirms that recipient information can affect deceptive preference in some settings, while showing that identical-looking behavior can arise without that mechanism and does not establish agency.

BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM AuditingBLOOM-WILT improves discovery of rare suspicious behavior; this framework supplies the next required step of testing whether an elicited output reflects the proposed deceptive mechanism.What Do Compliance Detectors Read? An Audit of Activation Probes and Guard ModelsBoth require counterfactual interventions to distinguish the claimed internal decision process from verdicts or outputs driven by superficial scenario cues.Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon LabsNarrows interpretation of eval-sensitive behavior: changing behavior with perceived information conditions warrants causal testing, not an immediate attribution of deceptive intent.Teaching AI to Find Real Vulnerabilities — David Brumley, BugcrowdContrasts outcome verification with mechanism attribution: deterministic oracles can establish what an agent achieved, but not why it produced deceptive-looking behavior.
Context Map
benchmarksecurity#agent-evals#benchmark-integrity#agent-reliability
Uncertainty
Some interventions showed that recipient information could causally affect **deceptive preference**, but deceptive-looking behavior also appeared without the proposed mechanism. Even mechanism-level evidence does not establish model agency.