# DeepSWE: A Contamination-Resistant Coding Benchmark — James Shi, Datacurve

Source: [AI Engineer](https://www.youtube.com/watch?v=Yk87oUPVaxU)  
Feed7 permalink: https://feed7.dev/p/deepswe-a-contamination-resistant-coding-benchmark-james-shi-datacurve-08p61c0  
Published: 2026-07-26T18:10:56.000Z  
Trust: Source Linked (source_linked)

## Why Included

DeepSWE uses original long-horizon tasks to reduce contamination and expose coding-agent behaviors hidden by saturated PR-mined suites. Its current task mix still underrepresents some everyday work.

## Source Summary

DeepSWE contains **113 original tasks** spanning **91 repositories** and five languages, with a median of one task per repository. Its short prompts still yield solutions averaging five times the lines of code of SWE-Bench Pro.

## Practical Implication

Use it when comparing coding agents on sustained repository work, and inspect behavioral traces alongside scores. **DeepSWE v1.1** separates verifier and agent runtimes and removes extra Git references to limit reward hacking.

## Agent-Ready Context

DeepSWE contains **113 original tasks** spanning **91 repositories** and five languages, with a median of one task per repository. Its short prompts still yield solutions averaging five times the lines of code of SWE-Bench Pro.

Use it when comparing coding agents on sustained repository work, and inspect behavioral traces alongside scores. **DeepSWE v1.1** separates verifier and agent runtimes and removes extra Git references to limit reward hacking.

The suite currently underrepresents bug localization and refactoring. Its authors also want broader repository coverage and hybrid verification, so it is not yet a complete proxy for routine engineering work.

## Context Map

- Layer: benchmark
- Domains: coding
- Topics: agent-evals, benchmark-integrity, agent-reliability

## Uncertainty

- The suite currently underrepresents bug localization and refactoring. Its authors also want broader repository coverage and hybrid verification, so it is not yet a complete proxy for routine engineering work.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
