Sign InOpen Brain
Back
arXivPaperNeeds ReviewbenchmarkNew

QuoteBench: How Matched Scores Can Hide Command-Path Failures

Test commands through the exact production transport and verify final state, since one parser cut task completion by 55.4–73.2 points.

arXivAug 13, 20262 min
Open SourceOpen MarkdownOpen JSON
Source Summary

QuoteBench shows that shell-command scores can conceal failures introduced by serialization and reparsing. Agent evals should identify the execution path, not attribute every result to the model.

Practical Implication

When the boundary was disclosed, six configurations recovered 30.4–60.7 points. Builders should test generated commands through the exact production transport and validate final state, especially where wrappers interpolate or reparse shell text.

Agent-Ready Context
QuoteBench tests **56 one-shot tasks** from 14 incident-derived families across eight configurations. Adding one unescaped parser cut replayed-command success by **55.4–73.2 percentage points**.

When the boundary was disclosed, six configurations recovered **30.4–60.7 points**. Builders should test generated commands through the exact production transport and validate final state, especially where wrappers interpolate or reparse shell text.

Adaptation was uneven: two configurations recovered nothing or declined slightly. GPT-5.6-sol’s **−3.6-point matched gap** concealed −64.3 points of transport damage plus 60.7 points of compensation, so aggregate scores can misstate both model and harness quality.
Context Map
benchmarkcodingsecurity#agent-evals#benchmark-integrity#harness-engineeringGeneric AgentPrepare Coding Session
Uncertainty
Automatically selected from source material; feed7 has not independently tested the claim.
Rate This Item
Personal Note

No note yet. Notes are included in exported bundles.

Related — Every Edge Explained

No approved edges yet for this post.