{
  "schema_version": "1.1",
  "id": "archive:https://arxiv.org/abs/2608.16852v1",
  "slug": "2608-16852v1-0580uyd",
  "url": "https://feed7.dev/p/2608-16852v1-0580uyd",
  "title": "What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models",
  "why_included": "Tested compliance guards often ignored the governing rule and classified from scenario cues. Builders should counterfactually swap policies before trusting a detector as an audit control.",
  "summary": "Across every tested guard and activation probe, deleting, permuting, or replacing the governing rule left accuracy unchanged. Even a policy-conditioned guard could cite a clause while barely changing its verdict when the clause became permissive; the authors call this **rule blindness**.",
  "practical_implication": "Audit compliance detectors with crossed examples where neither the scenario nor rule alone predicts the label. The paper found **step-by-step reasoning** escaped the failure seen in fast detectors and released a counterfactual protocol for testing future claims.",
  "agent_context": "Across every tested guard and activation probe, deleting, permuting, or replacing the governing rule left accuracy unchanged. Even a policy-conditioned guard could cite a clause while barely changing its verdict when the clause became permissive; the authors call this **rule blindness**.\n\nAudit compliance detectors with crossed examples where neither the scenario nor rule alone predicts the label. The paper found **step-by-step reasoning** escaped the failure seen in fast detectors and released a counterfactual protocol for testing future claims.\n\nThe proposed **Internal Compliance Score** missed its preregistered bar, matched a bag-of-words baseline in pooled generalization, and lost its ranking gain under an adaptive white-box attack. Its advantage is low-cost auditing, not demonstrated rule understanding.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2608.16852v1",
    "published_at": "2026-08-17T17:37:07.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "benchmark",
  "domains": [
    "security"
  ],
  "topics": [
    "agent-evals",
    "benchmark-integrity",
    "agent-reliability"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "The proposed **Internal Compliance Score** missed its preregistered bar, matched a bag-of-words baseline in pooled generalization, and lost its ranking gain under an adaptive white-box attack. Its advantage is low-cost auditing, not demonstrated rule understanding."
  ],
  "connected_context": {
    "meaning": "This turns general warnings about weak verifiers into a specific audit requirement: vary the governing rule independently of the scenario and verify that verdicts follow the rule. It also sharply limits activation-probe and fast-guard claims, because tested accuracy may reflect scenario cues rather than compliance reasoning; step-by-step reasoning is the reported exception.",
    "corpus_size": 479,
    "generated_at": "2026-08-18T10:05:01.713Z",
    "connections": [
      {
        "title": "Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.12962v1",
        "feed7_url": "https://feed7.dev/p/2607-12962v1-0q4i26c",
        "reason": "Both use controlled substitutions to test whether a system responds to the claimed content; each finds that apparent success can persist when the supposedly causal information is replaced or mismatched."
      },
      {
        "title": "When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=-npY6XjM8CQ",
        "feed7_url": "https://feed7.dev/p/when-will-the-benchmaxxing-plague-end-nick-heiner-surge-ai-178gqcg",
        "reason": "The paper supplies a concrete instance of the weak-verifier and shortcut problem: compliance scores remain stable when the governing policy changes."
      },
      {
        "title": "Cursor earns AIUC-1 certification for agent security and reliability",
        "source_name": "Cursor",
        "source_url": "https://cursor.com/blog/aiuc-1",
        "feed7_url": "https://feed7.dev/p/aiuc-1-0s1nr9l",
        "reason": "AIUC-1 adds live adversarial testing to controls audits; this paper identifies crossed policy–scenario counterfactuals as a useful test for whether compliance guards actually enforce those controls."
      },
      {
        "title": "MakazhanAlpamys/Soup",
        "source_name": "GitHub",
        "source_url": "https://github.com/MakazhanAlpamys/Soup",
        "feed7_url": "https://feed7.dev/p/soup-1rqg9hp",
        "reason": "Soup’s over-refusal and noise-floor checks already argue against trusting one aggregate score; rule-blind detectors add a distinct need to test whether the intended rule causally changes verdicts."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-08-17T17:37:07.000Z",
  "modified_at": "2026-08-17T17:37:07.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2608-16852v1-0580uyd",
    "json": "https://feed7.dev/p/2608-16852v1-0580uyd.json",
    "markdown": "https://feed7.dev/p/2608-16852v1-0580uyd.md"
  }
}