{
  "schema_version": "1.1",
  "id": "archive:https://arxiv.org/abs/2608.20338v1",
  "slug": "2608-20338v1-1robmt3",
  "url": "https://feed7.dev/p/2608-20338v1-1robmt3",
  "title": "ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models",
  "why_included": "ConceptGuard tests whether model unlearning blocks harmful uses of a concept while preserving benign ones. Current methods show weak contextual control and sharp forgetting-versus-utility trade-offs.",
  "summary": "**ConceptGuard** replaces independent forget and retain facts with complementary uses of **dual-use concepts**. The intent-sensitive benchmark measures whether harmful applications disappear while useful applications remain; its **17-page** paper and **public dataset** were released with the submission.",
  "practical_implication": "Use paired harmful and benign contexts when evaluating selective removal. Direct recall scores alone can hide whether a model erased an entire useful concept or still applies it unsafely under a different intent.",
  "agent_context": "**ConceptGuard** replaces independent forget and retain facts with complementary uses of **dual-use concepts**. The intent-sensitive benchmark measures whether harmful applications disappear while useful applications remain; its **17-page** paper and **public dataset** were released with the submission.\n\nUse paired harmful and benign contexts when evaluating selective removal. Direct recall scores alone can hide whether a model erased an entire useful concept or still applies it unsafely under a different intent.\n\nThe authors report weak contextual separation, inconsistent concept-level control, and strong forgetting-utility trade-offs across current techniques. The material provides no evidence that the proposed benchmark itself resolves those unlearning failures.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2608.20338v1",
    "published_at": "2026-08-20T17:59:57.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "benchmark",
  "domains": [
    "security",
    "research"
  ],
  "topics": [
    "benchmark-integrity"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "The authors report weak contextual separation, inconsistent concept-level control, and strong forgetting-utility trade-offs across current techniques. The material provides no evidence that the proposed benchmark itself resolves those unlearning failures."
  ],
  "connected_context": {
    "meaning": "ConceptGuard adds intent-sensitive selectivity to unlearning evaluation: success requires suppressing harmful uses without erasing benign uses of the same concept. This complements localization and resurfacing tests by exposing a different failure mode—concept-wide deletion or unsafe transfer across contexts—and confirms that recall-only scores are insufficient. It diagnoses current methods but does not improve them.",
    "corpus_size": 525,
    "generated_at": "2026-08-22T21:15:25.320Z",
    "connections": [
      {
        "title": "LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.02513v1",
        "feed7_url": "https://feed7.dev/p/2607-02513v1-0lwaytn",
        "reason": "LACUNA tests whether targeted knowledge is actually erased from known parameters; ConceptGuard tests whether removal is selectively expressed across harmful and benign uses. Together they separate localization and persistence failures from failures of contextual control."
      },
      {
        "title": "What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.16852v1",
        "feed7_url": "https://feed7.dev/p/2608-16852v1-0580uyd",
        "reason": "Both require counterfactual context changes to reveal shortcut behavior: policy swaps test whether compliance detectors follow rules, while paired harmful and benign uses test whether unlearning follows intent rather than suppressing a whole concept."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-08-20T17:59:57.000Z",
  "modified_at": "2026-08-20T17:59:57.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2608-20338v1-1robmt3",
    "json": "https://feed7.dev/p/2608-20338v1-1robmt3.json",
    "markdown": "https://feed7.dev/p/2608-20338v1-1robmt3.md"
  }
}