Sign InOpen Brain
AI EngineerVideoSource Linked

Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face

This security eval tests whether agents can discover and exploit logic flaws across live chained services, using hidden zero-days and deterministic grading instead of source-code pattern matching.

AI Engineer · Jul 24, 2026
Open Source Open MarkdownOpen JSON
Source Summary

The benchmark builds live, chained application environments around **real zero-days** found by human researchers. Agents start with limited access, receive **no internet or source code**, and must infer the system’s state while pursuing an exploit.

Practical Implication

For security-agent evals, test black-box reasoning across authentication, permissions, and service boundaries. Record partial progress with **deterministic graders** so model or harness changes can be compared even when full exploitation is rare.

Agent-Ready Context
The benchmark builds live, chained application environments around **real zero-days** found by human researchers. Agents start with limited access, receive **no internet or source code**, and must infer the system’s state while pursuing an exploit.

For security-agent evals, test black-box reasoning across authentication, permissions, and service boundaries. Record partial progress with **deterministic graders** so model or harness changes can be compared even when full exploitation is rare.

The presented benchmark remained extremely difficult, with **one solve at k=1**. Its access-control scenarios demonstrate a demanding evaluation method, but not yet the broader claim that specialized open models can replace existing defensive systems.
Context Map
benchmarksecurity#agent-evals#reasoning#open-models
Uncertainty
The presented benchmark remained extremely difficult, with **one solve at k=1**. Its access-control scenarios demonstrate a demanding evaluation method, but not yet the broader claim that specialized open models can replace existing defensive systems.