Sign InOpen Brain
arXivPaperNeeds Review

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

Tested compliance guards often ignored the governing rule and classified from scenario cues. Builders should counterfactually swap policies before trusting a detector as an audit control.

arXiv
Open Source Open MarkdownOpen JSON
Source Summary

Across every tested guard and activation probe, deleting, permuting, or replacing the governing rule left accuracy unchanged. Even a policy-conditioned guard could cite a clause while barely changing its verdict when the clause became permissive; the authors call this **rule blindness**.

Practical Implication

Audit compliance detectors with crossed examples where neither the scenario nor rule alone predicts the label. The paper found **step-by-step reasoning** escaped the failure seen in fast detectors and released a counterfactual protocol for testing future claims.

Agent-Ready Context
Across every tested guard and activation probe, deleting, permuting, or replacing the governing rule left accuracy unchanged. Even a policy-conditioned guard could cite a clause while barely changing its verdict when the clause became permissive; the authors call this **rule blindness**.

Audit compliance detectors with crossed examples where neither the scenario nor rule alone predicts the label. The paper found **step-by-step reasoning** escaped the failure seen in fast detectors and released a counterfactual protocol for testing future claims.

The proposed **Internal Compliance Score** missed its preregistered bar, matched a bag-of-words baseline in pooled generalization, and lost its ranking gain under an adaptive white-box attack. Its advantage is low-cost auditing, not demonstrated rule understanding.
Context Map
benchmarksecurity#agent-evals#benchmark-integrity#agent-reliability
Uncertainty
The proposed **Internal Compliance Score** missed its preregistered bar, matched a bag-of-words baseline in pooled generalization, and lost its ranking gain under an adaptive white-box attack. Its advantage is low-cost auditing, not demonstrated rule understanding.