Sign InOpen Brain
AI EngineerVideoSource Linked

How long can your skills be before your agent forgets what you told it? — Laurie Voss, Arize AI

Instruction capacity rose sharply, but long skill files still need evals. Current models can track thousands of constraints, with large differences by model, wording, and order.

AI Engineer · Sep 9, 2026
Open Source Open MarkdownOpen JSON
Source Summary

A rerun of IFScale reproduced the old failure point around **200–300 rules**. Current frontier models reached roughly **2,000–5,000 rules** before major decline; GPT-5.5 reportedly held **99% accuracy at 5,000 rules** on this proxy task.

Practical Implication

Revisit prompt-length assumptions before splitting a style guide or policy into many agents. Capacity is less likely to be the first constraint now, so test whether one coherent skill is cheaper and easier to maintain than a handoff-heavy design.

Agent-Ready Context
A rerun of IFScale reproduced the old failure point around **200–300 rules**. Current frontier models reached roughly **2,000–5,000 rules** before major decline; GPT-5.5 reportedly held **99% accuracy at 5,000 rules** on this proxy task.

Revisit prompt-length assumptions before splitting a style guide or policy into many agents. Capacity is less likely to be the first constraint now, so test whether one coherent skill is cheaper and easier to maintain than a handoff-heavy design.

IFScale checks whether required words appear in a synthetic report, not whether an agent reasons well. Reported limits ranged from **750 to 9,000-plus rules**, and performance can shift with instruction wording or order, making output verification essential.
Connected Context · Feed7 Judgment

This raises the estimated instruction-capacity ceiling and weakens capacity alone as a reason to split coherent guidance across agents. It does not show that long skills improve real agent work: IFScale measures synthetic rule adherence, and ordering and wording still matter. Against the prior candidates, the practical conclusion is to compare one maintained skill with handoff-heavy designs using repeated, architecture-relevant regression tests.

Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMindThe higher rule ceiling makes longer skills plausible, while this candidate supplies the necessary test discipline: compare triggering and outcomes with and without the skill across repeated runs.The Regression Tax: Decomposing Why Skills Help and Hurt LLM AgentsIt limits the capacity result’s interpretation: an agent may retain thousands of instructions yet still regress on tasks it previously solved, so adherence and net utility must be measured separately.Your Agent Evolved. Your Evals Didn't. — Ameya Bhatawdekar, BraintrustChoosing one large skill instead of multiple agents changes the harness architecture, reinforcing the candidate’s requirement that eval coverage follow planning, handoffs, tools, and other system structure.SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering AgentsSWE-Gate reinforces that following observable requirements is not equivalent to complete engineering correctness, matching the warning that IFScale’s word-inclusion proxy cannot establish real task performance.
Context Map
benchmarkcoding#skills#agent-evals#agent-reliability
Uncertainty
IFScale checks whether required words appear in a synthetic report, not whether an agent reasons well. Reported limits ranged from **750 to 9,000-plus rules**, and performance can shift with instruction wording or order, making output verification essential.