# What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

Source: [arXiv](https://arxiv.org/abs/2608.16852v1)  
Feed7 permalink: https://feed7.dev/p/2608-16852v1-0580uyd  
Published: 2026-08-17T17:37:07.000Z  
Trust: Needs Review (needs_review)

## Why Included

Tested compliance guards often ignored the governing rule and classified from scenario cues. Builders should counterfactually swap policies before trusting a detector as an audit control.

## Source Summary

Across every tested guard and activation probe, deleting, permuting, or replacing the governing rule left accuracy unchanged. Even a policy-conditioned guard could cite a clause while barely changing its verdict when the clause became permissive; the authors call this **rule blindness**.

## Practical Implication

Audit compliance detectors with crossed examples where neither the scenario nor rule alone predicts the label. The paper found **step-by-step reasoning** escaped the failure seen in fast detectors and released a counterfactual protocol for testing future claims.

## Agent-Ready Context

Across every tested guard and activation probe, deleting, permuting, or replacing the governing rule left accuracy unchanged. Even a policy-conditioned guard could cite a clause while barely changing its verdict when the clause became permissive; the authors call this **rule blindness**.

Audit compliance detectors with crossed examples where neither the scenario nor rule alone predicts the label. The paper found **step-by-step reasoning** escaped the failure seen in fast detectors and released a counterfactual protocol for testing future claims.

The proposed **Internal Compliance Score** missed its preregistered bar, matched a bag-of-words baseline in pooled generalization, and lost its ranking gain under an adaptive white-box attack. Its advantage is low-cost auditing, not demonstrated rule understanding.

## Connected Context

Feed7 judgment across 479 accumulated Signals:

This turns general warnings about weak verifiers into a specific audit requirement: vary the governing rule independently of the scenario and verify that verdicts follow the rule. It also sharply limits activation-probe and fast-guard claims, because tested accuracy may reflect scenario cues rather than compliance reasoning; step-by-step reasoning is the reported exception.

- [Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models](https://feed7.dev/p/2607-12962v1-0q4i26c) — Both use controlled substitutions to test whether a system responds to the claimed content; each finds that apparent success can persist when the supposedly causal information is replaced or mismatched.
- [When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI](https://feed7.dev/p/when-will-the-benchmaxxing-plague-end-nick-heiner-surge-ai-178gqcg) — The paper supplies a concrete instance of the weak-verifier and shortcut problem: compliance scores remain stable when the governing policy changes.
- [Cursor earns AIUC-1 certification for agent security and reliability](https://feed7.dev/p/aiuc-1-0s1nr9l) — AIUC-1 adds live adversarial testing to controls audits; this paper identifies crossed policy–scenario counterfactuals as a useful test for whether compliance guards actually enforce those controls.
- [MakazhanAlpamys/Soup](https://feed7.dev/p/soup-1rqg9hp) — Soup’s over-refusal and noise-floor checks already argue against trusting one aggregate score; rule-blind detectors add a distinct need to test whether the intended rule causally changes verdicts.

## Context Map

- Layer: benchmark
- Domains: security
- Topics: agent-evals, benchmark-integrity, agent-reliability

## Uncertainty

- The proposed **Internal Compliance Score** missed its preregistered bar, matched a bag-of-words baseline in pooled generalization, and lost its ranking gain under an adaptive white-box attack. Its advantage is low-cost auditing, not demonstrated rule understanding.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
