Sign InOpen Brain
arXivPaperNeeds Review

ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

ConceptGuard tests whether model unlearning blocks harmful uses of a concept while preserving benign ones. Current methods show weak contextual control and sharp forgetting-versus-utility trade-offs.

arXiv
Open Source Open MarkdownOpen JSON
Source Summary

**ConceptGuard** replaces independent forget and retain facts with complementary uses of **dual-use concepts**. The intent-sensitive benchmark measures whether harmful applications disappear while useful applications remain; its **17-page** paper and **public dataset** were released with the submission.

Practical Implication

Use paired harmful and benign contexts when evaluating selective removal. Direct recall scores alone can hide whether a model erased an entire useful concept or still applies it unsafely under a different intent.

Agent-Ready Context
**ConceptGuard** replaces independent forget and retain facts with complementary uses of **dual-use concepts**. The intent-sensitive benchmark measures whether harmful applications disappear while useful applications remain; its **17-page** paper and **public dataset** were released with the submission.

Use paired harmful and benign contexts when evaluating selective removal. Direct recall scores alone can hide whether a model erased an entire useful concept or still applies it unsafely under a different intent.

The authors report weak contextual separation, inconsistent concept-level control, and strong forgetting-utility trade-offs across current techniques. The material provides no evidence that the proposed benchmark itself resolves those unlearning failures.
Context Map
benchmarksecurityresearch#benchmark-integrity
Uncertainty
The authors report weak contextual separation, inconsistent concept-level control, and strong forgetting-utility trade-offs across current techniques. The material provides no evidence that the proposed benchmark itself resolves those unlearning failures.