# Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence

Source: [arXiv](https://arxiv.org/abs/2609.05385v1)  
Feed7 permalink: https://feed7.dev/p/2609-05385v1-0bf7wcv  
Published: 2026-09-04T17:37:17.000Z  
Trust: Needs Review (needs_review)

## Why Included

LLMs' stated top decision factors only weakly tracked factors shown to affect outputs. For agent oversight, validate explanations with controlled interventions before using them for escalation.

## Source Summary

Across **eight Claude, GPT, and Gemini models**, cited factor rankings had mean Spearman correlations of **0.349–0.580** with measured necessity or sufficiency in two synthetic decision tasks.

## Practical Implication

If agent monitoring depends on explanations, test each named factor by changing it and by retaining it while removing other mutable information. Do not treat a plausible top-three rationale as a faithful account of the model's decision boundary.

## Agent-Ready Context

Across **eight Claude, GPT, and Gemini models**, cited factor rankings had mean Spearman correlations of **0.349–0.580** with measured necessity or sufficiency in two synthetic decision tasks.

If agent monitoring depends on explanations, test each named factor by changing it and by retaining it while removing other mutable information. Do not treat a plausible top-three rationale as a faithful account of the model's decision boundary.

In advisor recommendations, uncited factors outranked a cited factor in **57.6% under necessity and 58.1% under sufficiency**; prompt-monitoring rates were **25.8% and 8.9%**. The method covers individual decisions in synthetic settings, not whole agent trajectories.

## Connected Context

Feed7 judgment across 703 accumulated Signals:

This makes intervention-based evaluation concrete for ordinary model explanations: a plausible rationale is neither evidence that a cited factor was necessary nor that it was sufficient. It reinforces prior warnings about treating readable traces or suspicious outputs as causal accounts, while narrowing the result to individual decisions in synthetic tasks rather than whole-agent behavior.

- [From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research](https://feed7.dev/p/2609-04166v1-04gqa5r) — Both require counterfactual intervention before inferring an underlying mechanism from model output, whether the claim concerns a decision rationale or deception.

## Context Map

- Layer: benchmark
- Domains: security
- Topics: agent-evals, agent-reliability

## Uncertainty

- In advisor recommendations, uncited factors outranked a cited factor in **57.6% under necessity and 58.1% under sufficiency**; prompt-monitoring rates were **25.8% and 8.9%**. The method covers individual decisions in synthetic settings, not whole agent trajectories.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
