{
  "schema_version": "1.1",
  "id": "archive:https://arxiv.org/abs/2608.13547v1",
  "slug": "2608-13547v1-130h6xd",
  "url": "https://feed7.dev/p/2608-13547v1-130h6xd",
  "title": "QuoteBench: How Matched Scores Can Hide Command-Path Failures",
  "why_included": "QuoteBench shows that shell-command scores can conceal failures introduced by serialization and reparsing. Agent evals should identify the execution path, not attribute every result to the model.",
  "summary": "QuoteBench tests **56 one-shot tasks** from 14 incident-derived families across eight configurations. Adding one unescaped parser cut replayed-command success by **55.4–73.2 percentage points**.",
  "practical_implication": "When the boundary was disclosed, six configurations recovered **30.4–60.7 points**. Builders should test generated commands through the exact production transport and validate final state, especially where wrappers interpolate or reparse shell text.",
  "agent_context": "QuoteBench tests **56 one-shot tasks** from 14 incident-derived families across eight configurations. Adding one unescaped parser cut replayed-command success by **55.4–73.2 percentage points**.\n\nWhen the boundary was disclosed, six configurations recovered **30.4–60.7 points**. Builders should test generated commands through the exact production transport and validate final state, especially where wrappers interpolate or reparse shell text.\n\nAdaptation was uneven: two configurations recovered nothing or declined slightly. GPT-5.6-sol’s **−3.6-point matched gap** concealed −64.3 points of transport damage plus 60.7 points of compensation, so aggregate scores can misstate both model and harness quality.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2608.13547v1",
    "published_at": "2026-08-13T17:57:20.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "benchmark",
  "domains": [
    "coding",
    "security"
  ],
  "topics": [
    "agent-evals",
    "benchmark-integrity",
    "harness-engineering"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "Adaptation was uneven: two configurations recovered nothing or declined slightly. GPT-5.6-sol’s **−3.6-point matched gap** concealed −64.3 points of transport damage plus 60.7 points of compensation, so aggregate scores can misstate both model and harness quality."
  ],
  "connected_context": {
    "meaning": "This provides quantified evidence that command transport is part of the evaluated system: a model can appear unchanged in aggregate while suffering severe parser-induced failures and partially compensating after disclosure. It strengthens calls for cross-harness testing and final-state validation, and further narrows leaderboard interpretation because matched scores can hide offsetting model adaptation and harness damage.",
    "corpus_size": 462,
    "generated_at": "2026-08-16T10:04:24.195Z",
    "connections": [
      {
        "title": "State of Data — Sean Cai, Independent / State of Data",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=ZyIoTOAbRfs",
        "feed7_url": "https://feed7.dev/p/state-of-data-sean-cai-independent-state-of-data-0v9fy69",
        "reason": "QuoteBench quantifies the earlier claim that scores shift with scaffolding and shows why evaluation should inspect tool trajectories and resulting state rather than only aggregate task success."
      },
      {
        "title": "The Bitter Lesson of Tool Calling",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.06370v1",
        "feed7_url": "https://feed7.dev/p/2608-06370v1-1ar4rfv",
        "reason": "Both establish interface representation as a consequential harness variable; QuoteBench adds that reparsing and escaping along the production command path can dominate apparent model performance."
      },
      {
        "title": "Teaching AI to Find Real Vulnerabilities — David Brumley, Bugcrowd",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=ZFxh7sqbUZo",
        "feed7_url": "https://feed7.dev/p/teaching-ai-to-find-real-vulnerabilities-david-brumley-bugcrowd-1ok0f7q",
        "reason": "Bugcrowd’s deterministic exploit oracles reinforce QuoteBench’s requirement to validate concrete final effects instead of trusting generated commands or self-reported completion."
      },
      {
        "title": "Quantifying infrastructure noise in agentic coding evals",
        "source_name": "Anthropic",
        "source_url": "https://www.anthropic.com/engineering/infrastructure-noise",
        "feed7_url": "https://feed7.dev/p/infrastructure-noise-1jyyyw1",
        "reason": "Anthropic shows resource configuration can move scores, while QuoteBench demonstrates a different hidden systems effect: transport damage can be masked by model compensation even when the aggregate gap is small."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-08-13T17:57:20.000Z",
  "modified_at": "2026-08-13T17:57:20.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2608-13547v1-130h6xd",
    "json": "https://feed7.dev/p/2608-13547v1-130h6xd.json",
    "markdown": "https://feed7.dev/p/2608-13547v1-130h6xd.md"
  }
}