Sign InOpen Brain
Back
arXivPaperNeeds ReviewbenchmarkNew

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

Green tests missed review constraints in 221 of 644 passing repairs, so encode repository expectations as separate executable checks.

arXivSep 3, 20262 min
Open SourceOpen MarkdownOpen JSON
Source Summary

SWE-Gate shows why green tests are an incomplete agent-eval signal: 221 of 644 functionally passing repairs still violated constraints derived from code review.

Practical Implication

Treat green tests as one gate, not final acceptance, for coding-agent patches. Encode review expectations as executable checks where possible, and evaluate issue resolution separately from compliance with repository-specific requirements.

Agent-Ready Context
SWE-Gate contains **303 repair instances** from **75 Python repositories**, with separate tests for functionality and constraints derived from real pull-request reviews. Across four model backends, 644 repairs passed functional tests, but **221** still violated review constraints.

Treat green tests as one gate, not final acceptance, for coding-agent patches. Encode review expectations as executable checks where possible, and evaluate issue resolution separately from compliance with repository-specific requirements.

The benchmark uses synthesized repair instances and a common agent scaffold, so its failure rates may not transfer directly to your repos. It nevertheless exposes a concrete blind spot in functional-only coding-agent evaluations.
Context Map
benchmarkcoding#agent-evals#benchmark-integrity#agent-reliabilityGeneric AgentPrepare Coding Session
Uncertainty
Automatically selected from source material; feed7 has not independently tested the claim.
Rate This Item
Personal Note

No note yet. Notes are included in exported bundles.

Related — Every Edge Explained

No approved edges yet for this post.