Sign InOpen Brain
arXivPaperNeeds Review

RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution

RedEvoAgent turns prior jailbreak trajectories into a compact attack skill, then keeps only validated improvements. It offers a more interpretable way to probe tool-using agents for unsafe actions.

arXiv
Open Source Open MarkdownOpen JSON
Source Summary

RedEvoAgent is a **black-box red-teaming agent** that distills attack trajectories into a concise skill. Tool-effectiveness profiling and **Deciding-Tool Attribution** determine how that skill changes, while a validation ratchet rejects regressions.

Practical Implication

Builders should test agent harnesses for harmful tool calls and persistent state changes, not just unsafe text. Compact, inspectable attack skills may also be easier to audit and cheaper to reuse than full trajectory retrieval.

Agent-Ready Context
RedEvoAgent is a **black-box red-teaming agent** that distills attack trajectories into a concise skill. Tool-effectiveness profiling and **Deciding-Tool Attribution** determine how that skill changes, while a validation ratchet rejects regressions.

Builders should test agent harnesses for harmful tool calls and persistent state changes, not just unsafe text. Compact, inspectable attack skills may also be easier to audit and cheaper to reuse than full trajectory retrieval.

The abstract reports gains across multiple benchmarks, models, and execution harnesses but provides no figures here. Transfer claims and operational cost still need inspection in the full paper.
Context Map
benchmarksecurity#agent-evals#agent-reliability#skills
Uncertainty
The abstract reports gains across multiple benchmarks, models, and execution harnesses but provides no figures here. Transfer claims and operational cost still need inspection in the full paper.