# From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

Source: [arXiv](https://arxiv.org/abs/2609.04166v1)  
Feed7 permalink: https://feed7.dev/p/2609-04166v1-04gqa5r  
Published: 2026-09-03T17:52:19.000Z  
Trust: Needs Review (needs_review)

## Why Included

A causal framework separates deceptive-looking model output from the mechanism producing it, giving agent evaluators a stricter basis for claims about intent or agency.

## Source Summary

The paper separates observed output from prior commitment, internal preference, sensitivity to a recipient’s information, and the origin of an objective. Tests used **two open-weight model families** in controlled guessing and stock-trading settings.

## Practical Implication

When evaluating an agent for manipulation or dishonesty, test causal alternatives instead of labeling a suspicious answer as intent. Intervene on what the recipient knows and distinguish the model’s preference from the output ultimately sampled.

## Agent-Ready Context

The paper separates observed output from prior commitment, internal preference, sensitivity to a recipient’s information, and the origin of an objective. Tests used **two open-weight model families** in controlled guessing and stock-trading settings.

When evaluating an agent for manipulation or dishonesty, test causal alternatives instead of labeling a suspicious answer as intent. Intervene on what the recipient knows and distinguish the model’s preference from the output ultimately sampled.

Some interventions showed that recipient information could causally affect **deceptive preference**, but deceptive-looking behavior also appeared without the proposed mechanism. Even mechanism-level evidence does not establish model agency.

## Connected Context

Feed7 judgment across 691 accumulated Signals:

This raises the evidentiary bar for deception evaluations: suspicious outputs are insufficient without interventions that separate prior commitment, preference, recipient knowledge, and sampling. It confirms that recipient information can affect deceptive preference in some settings, while showing that identical-looking behavior can arise without that mechanism and does not establish agency.

- [BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing](https://feed7.dev/p/2608-31105v1-0vdo3q3) — BLOOM-WILT improves discovery of rare suspicious behavior; this framework supplies the next required step of testing whether an elicited output reflects the proposed deceptive mechanism.
- [What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models](https://feed7.dev/p/2608-16852v1-0580uyd) — Both require counterfactual interventions to distinguish the claimed internal decision process from verdicts or outputs driven by superficial scenario cues.
- [Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs](https://feed7.dev/p/vending-bench-long-horizon-agent-evals-lukas-petersson-andon-labs-0fu78nz) — Narrows interpretation of eval-sensitive behavior: changing behavior with perceived information conditions warrants causal testing, not an immediate attribution of deceptive intent.
- [Teaching AI to Find Real Vulnerabilities — David Brumley, Bugcrowd](https://feed7.dev/p/teaching-ai-to-find-real-vulnerabilities-david-brumley-bugcrowd-1ok0f7q) — Contrasts outcome verification with mechanism attribution: deterministic oracles can establish what an agent achieved, but not why it produced deceptive-looking behavior.

## Context Map

- Layer: benchmark
- Domains: security
- Topics: agent-evals, benchmark-integrity, agent-reliability

## Uncertainty

- Some interventions showed that recipient information could causally affect **deceptive preference**, but deceptive-looking behavior also appeared without the proposed mechanism. Even mechanism-level evidence does not establish model agency.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
