Skip to main content

> ML_DATASET // SWE-BENCH-DATASET_v1.0

SWE-bench: Real-World GitHub Software Engineering Issues

Princeton NLP (Jimenez et al.) · Software Engineering & Autonomous Agents · 2,294 real-world GitHub issues and pull request test patches across 12 popular open-source Python repos

Software Engineering & Autonomous AgentsMIT2,294 real-world GitHub issues and pull request test patches across 12 popular open-source Python reposopen

Dataset Profile & Characteristics

Label Type:Gold patch commit diff and unit test verification suite (PASS_TO_PASS and FAIL_TO_PASS)
Languages:en, code
License Tier:permissive-open-source
Modalities:text, code

Intended Use

  • Evaluating autonomous AI coding agents on end-to-end bug resolution via unit test execution

Prohibited / Discouraged Use

  • Unsandboxed execution without Docker container isolation

Bias, Leakage & Privacy Risk Analysis

Privacy / Sensitive Data Risks:

Public open-source GitHub issues; developer usernames in commit logs.

Known Bias:

Python open-source library engineering practices (Django, SymPy, scikit-learn, etc.).

Known Benchmark Leakage:

Pre-training cutoff dates of older models include training repo history.

Compatible Tools & Libraries