# ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

Source: [arXiv](https://arxiv.org/abs/2608.20338v1)  
Feed7 permalink: https://feed7.dev/p/2608-20338v1-1robmt3  
Published: 2026-08-20T17:59:57.000Z  
Trust: Needs Review (needs_review)

## Why Included

ConceptGuard tests whether model unlearning blocks harmful uses of a concept while preserving benign ones. Current methods show weak contextual control and sharp forgetting-versus-utility trade-offs.

## Source Summary

**ConceptGuard** replaces independent forget and retain facts with complementary uses of **dual-use concepts**. The intent-sensitive benchmark measures whether harmful applications disappear while useful applications remain; its **17-page** paper and **public dataset** were released with the submission.

## Practical Implication

Use paired harmful and benign contexts when evaluating selective removal. Direct recall scores alone can hide whether a model erased an entire useful concept or still applies it unsafely under a different intent.

## Agent-Ready Context

**ConceptGuard** replaces independent forget and retain facts with complementary uses of **dual-use concepts**. The intent-sensitive benchmark measures whether harmful applications disappear while useful applications remain; its **17-page** paper and **public dataset** were released with the submission.

Use paired harmful and benign contexts when evaluating selective removal. Direct recall scores alone can hide whether a model erased an entire useful concept or still applies it unsafely under a different intent.

The authors report weak contextual separation, inconsistent concept-level control, and strong forgetting-utility trade-offs across current techniques. The material provides no evidence that the proposed benchmark itself resolves those unlearning failures.

## Connected Context

Feed7 judgment across 525 accumulated Signals:

ConceptGuard adds intent-sensitive selectivity to unlearning evaluation: success requires suppressing harmful uses without erasing benign uses of the same concept. This complements localization and resurfacing tests by exposing a different failure mode—concept-wide deletion or unsafe transfer across contexts—and confirms that recall-only scores are insufficient. It diagnoses current methods but does not improve them.

- [LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning](https://feed7.dev/p/2607-02513v1-0lwaytn) — LACUNA tests whether targeted knowledge is actually erased from known parameters; ConceptGuard tests whether removal is selectively expressed across harmful and benign uses. Together they separate localization and persistence failures from failures of contextual control.
- [What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models](https://feed7.dev/p/2608-16852v1-0580uyd) — Both require counterfactual context changes to reveal shortcut behavior: policy swaps test whether compliance detectors follow rules, while paired harmful and benign uses test whether unlearning follows intent rather than suppressing a whole concept.

## Context Map

- Layer: benchmark
- Domains: security, research
- Topics: benchmark-integrity

## Uncertainty

- The authors report weak contextual separation, inconsistent concept-level control, and strong forgetting-utility trade-offs across current techniques. The material provides no evidence that the proposed benchmark itself resolves those unlearning failures.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
