> ML_LITERATURE // JIMENEZ-2024-SWE-BENCH-CAN-LANGUAGE-MODELS-RESOLVE-REAL-WORLD-GITHUB-ISSUES_v1.0
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, Karthik Narasimhan · International Conference on Learning Representations (ICLR) (2024)
benchmark2024industry-standardartifactsAvailable
Principal Contribution
Created SWE-bench: 2,294 real-world software engineering issues from top open-source Python repositories evaluated via execution of full test suites.
Operational Relevance
Directly guides deployment choices and architecture selection for task-code-generation, task-autonomous-agents.
Assumptions
- Standard empirical regularity and statistical stability hold across evaluation domains
Limitations
- Performance characteristics depend on domain distribution and compute allocation parameters
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
