# How long can your skills be before your agent forgets what you told it? — Laurie Voss, Arize AI

Source: [AI Engineer](https://www.youtube.com/watch?v=XzJD1bvXKjs)  
Feed7 permalink: https://feed7.dev/p/how-long-can-your-skills-be-before-your-agent-forgets-what-you-told-it-l-0i0jeww  
Published: 2026-09-09T15:00:06.000Z  
Trust: Source Linked (source_linked)

## Why Included

Instruction capacity rose sharply, but long skill files still need evals. Current models can track thousands of constraints, with large differences by model, wording, and order.

## Source Summary

A rerun of IFScale reproduced the old failure point around **200–300 rules**. Current frontier models reached roughly **2,000–5,000 rules** before major decline; GPT-5.5 reportedly held **99% accuracy at 5,000 rules** on this proxy task.

## Practical Implication

Revisit prompt-length assumptions before splitting a style guide or policy into many agents. Capacity is less likely to be the first constraint now, so test whether one coherent skill is cheaper and easier to maintain than a handoff-heavy design.

## Agent-Ready Context

A rerun of IFScale reproduced the old failure point around **200–300 rules**. Current frontier models reached roughly **2,000–5,000 rules** before major decline; GPT-5.5 reportedly held **99% accuracy at 5,000 rules** on this proxy task.

Revisit prompt-length assumptions before splitting a style guide or policy into many agents. Capacity is less likely to be the first constraint now, so test whether one coherent skill is cheaper and easier to maintain than a handoff-heavy design.

IFScale checks whether required words appear in a synthetic report, not whether an agent reasons well. Reported limits ranged from **750 to 9,000-plus rules**, and performance can shift with instruction wording or order, making output verification essential.

## Connected Context

Feed7 judgment across 732 accumulated Signals:

This raises the estimated instruction-capacity ceiling and weakens capacity alone as a reason to split coherent guidance across agents. It does not show that long skills improve real agent work: IFScale measures synthetic rule adherence, and ordering and wording still matter. Against the prior candidates, the practical conclusion is to compare one maintained skill with handoff-heavy designs using repeated, architecture-relevant regression tests.

- [Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMind](https://feed7.dev/p/don-t-ship-skills-without-evals-philipp-schmid-google-deepmind-0fuh3ko) — The higher rule ceiling makes longer skills plausible, while this candidate supplies the necessary test discipline: compare triggering and outcomes with and without the skill across repeated runs.
- [The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents](https://feed7.dev/p/2607-22520v1-0mz9wnf) — It limits the capacity result’s interpretation: an agent may retain thousands of instructions yet still regress on tasks it previously solved, so adherence and net utility must be measured separately.
- [Your Agent Evolved. Your Evals Didn't. — Ameya Bhatawdekar, Braintrust](https://feed7.dev/p/your-agent-evolved-your-evals-didn-t-ameya-bhatawdekar-braintrust-1loaqv2) — Choosing one large skill instead of multiple agents changes the harness architecture, reinforcing the candidate’s requirement that eval coverage follow planning, handoffs, tools, and other system structure.
- [SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents](https://feed7.dev/p/2609-04167v1-1vvofgw) — SWE-Gate reinforces that following observable requirements is not equivalent to complete engineering correctness, matching the warning that IFScale’s word-inclusion proxy cannot establish real task performance.

## Context Map

- Layer: benchmark
- Domains: coding
- Topics: skills, agent-evals, agent-reliability

## Uncertainty

- IFScale checks whether required words appear in a synthetic report, not whether an agent reasons well. Reported limits ranged from **750 to 9,000-plus rules**, and performance can shift with instruction wording or order, making output verification essential.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
