# QuoteBench: How Matched Scores Can Hide Command-Path Failures

Source: [arXiv](https://arxiv.org/abs/2608.13547v1)  
Feed7 permalink: https://feed7.dev/p/quotebench-how-matched-scores-can-hide-command-path-fail-31ceb9467c  
Published: 2026-08-13T00:00:00.000Z  
Trust: Needs Review (needs_review)

## Why Included

Test commands through the exact production transport and verify final state, since one parser cut task completion by 55.4–73.2 points.

## Source Summary

QuoteBench shows that shell-command scores can conceal failures introduced by serialization and reparsing. Agent evals should identify the execution path, not attribute every result to the model.

## Practical Implication

When the boundary was disclosed, six configurations recovered 30.4–60.7 points. Builders should test generated commands through the exact production transport and validate final state, especially where wrappers interpolate or reparse shell text.

## Agent-Ready Context

QuoteBench tests **56 one-shot tasks** from 14 incident-derived families across eight configurations. Adding one unescaped parser cut replayed-command success by **55.4–73.2 percentage points**.

When the boundary was disclosed, six configurations recovered **30.4–60.7 points**. Builders should test generated commands through the exact production transport and validate final state, especially where wrappers interpolate or reparse shell text.

Adaptation was uneven: two configurations recovered nothing or declined slightly. GPT-5.6-sol’s **−3.6-point matched gap** concealed −64.3 points of transport damage plus 60.7 points of compensation, so aggregate scores can misstate both model and harness quality.

## Context Map

- Layer: benchmark
- Domains: coding, security
- Topics: agent-evals, benchmark-integrity, harness-engineering

## Uncertainty

- Automatically selected from source material; feed7 has not independently tested the claim.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
