> ML_DATASET // SWE-BENCH-DATASET_v1.0
SWE-bench: Real-World GitHub Software Engineering Issues
Princeton NLP (Jimenez et al.) · Software Engineering & Autonomous Agents · 2,294 real-world GitHub issues and pull request test patches across 12 popular open-source Python repos
Software Engineering & Autonomous AgentsMIT2,294 real-world GitHub issues and pull request test patches across 12 popular open-source Python reposopen
Dataset Profile & Characteristics
Label Type:Gold patch commit diff and unit test verification suite (PASS_TO_PASS and FAIL_TO_PASS)
Languages:en, code
License Tier:permissive-open-source
Modalities:text, code
Intended Use
- Evaluating autonomous AI coding agents on end-to-end bug resolution via unit test execution
Prohibited / Discouraged Use
- Unsandboxed execution without Docker container isolation
Bias, Leakage & Privacy Risk Analysis
Privacy / Sensitive Data Risks:
Public open-source GitHub issues; developer usernames in commit logs.
Known Bias:
Python open-source library engineering practices (Django, SymPy, scikit-learn, etc.).
Known Benchmark Leakage:
Pre-training cutoff dates of older models include training repo history.
