{
  "schema_version": "1.0",
  "id": "s13:https://arxiv.org/abs/2607.22520v1",
  "slug": "2607-22520v1-0mz9wnf",
  "url": "https://feed7.dev/p/2607-22520v1-0mz9wnf",
  "title": "The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents",
  "why_included": "Procedural skills can make an agent fail tasks it previously solved. Evaluate gains and regressions separately, and design skills to preserve input grounding and output verification.",
  "summary": "Across **nearly 6,000 runs**, two office-automation benchmarks, and three harness stacks, adding skills caused meaningful regressions. The strongest skills led mainly by breaking fewer previously solved tasks, not by creating more new wins.",
  "practical_implication": "Measure each skill against a no-skill baseline and split net improvement into gains and regressions. Prioritize grounding and verification support instead of adding more procedure, since those stages dominated persistent failures.",
  "agent_context": "Across **nearly 6,000 runs**, two office-automation benchmarks, and three harness stacks, adding skills caused meaningful regressions. The strongest skills led mainly by breaking fewer previously solved tasks, not by creating more new wins.\n\nMeasure each skill against a no-skill baseline and split net improvement into gains and regressions. Prioritize grounding and verification support instead of adding more procedure, since those stages dominated persistent failures.\n\nThe authors identify **three regression modes**: description osmosis, grounding displacement, and verification displacement. The evidence comes from office automation, so the prevalence and size of these effects in repository-scale coding agents remain open.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2607.22520v1",
    "published_at": "2026-07-24T17:50:03.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "agent",
  "domains": [],
  "topics": [
    "skills",
    "agent-reliability",
    "agent-evals"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "The authors identify **three regression modes**: description osmosis, grounding displacement, and verification displacement. The evidence comes from office automation, so the prevalence and size of these effects in repository-scale coding agents remain open."
  ],
  "lifecycle": "Current",
  "published_at": "2026-07-24T17:50:03.000Z",
  "modified_at": "2026-07-24T17:50:03.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2607-22520v1-0mz9wnf",
    "json": "https://feed7.dev/p/2607-22520v1-0mz9wnf.json",
    "markdown": "https://feed7.dev/p/2607-22520v1-0mz9wnf.md"
  }
}