Sign InOpen Brain
arXivPaperNeeds Review

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

SWE-Gate shows why green tests are an incomplete agent-eval signal: 221 of 644 functionally passing repairs still violated constraints derived from code review.

arXiv · Sep 3, 2026
Open Source Open MarkdownOpen JSON
Source Summary

SWE-Gate contains **303 repair instances** from **75 Python repositories**, with separate tests for functionality and constraints derived from real pull-request reviews. Across four model backends, 644 repairs passed functional tests, but **221** still violated review constraints.

Practical Implication

Treat green tests as one gate, not final acceptance, for coding-agent patches. Encode review expectations as executable checks where possible, and evaluate issue resolution separately from compliance with repository-specific requirements.

Agent-Ready Context
SWE-Gate contains **303 repair instances** from **75 Python repositories**, with separate tests for functionality and constraints derived from real pull-request reviews. Across four model backends, 644 repairs passed functional tests, but **221** still violated review constraints.

Treat green tests as one gate, not final acceptance, for coding-agent patches. Encode review expectations as executable checks where possible, and evaluate issue resolution separately from compliance with repository-specific requirements.

The benchmark uses synthesized repair instances and a common agent scaffold, so its failure rates may not transfer directly to your repos. It nevertheless exposes a concrete blind spot in functional-only coding-agent evaluations.
Connected Context · Feed7 Judgment

This converts a general warning about coding-agent benchmarks into a specific acceptance gap: functional success can coexist with violations of repository review constraints. It argues for separately executable functionality and compliance gates, while synthesized tasks and a shared scaffold limit direct extrapolation to production repositories.

DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, DatacurveComplements DeepSWE’s contamination and long-horizon focus by measuring a different hidden failure mode: patches that work but disregard review-derived constraints.Rethinking Environments for Long-Horizon Work — Rayan Garg, Theta SoftwareImplements the call for final-state inspection by splitting acceptance into functional state and repository-specific constraint compliance.Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code ModelsTogether they weaken simplistic repair evidence: retries may not use error content, and passing functional tests may still leave the requested constraints unresolved.When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AIProvides concrete support for the claim that verifier design can inflate coding-agent success when tests omit requirements that human reviewers enforce.
Context Map
benchmarkcoding#agent-evals#benchmark-integrity#agent-reliability
Uncertainty
The benchmark uses synthesized repair instances and a common agent scaffold, so its failure rates may not transfer directly to your repos. It nevertheless exposes a concrete blind spot in functional-only coding-agent evaluations.