Skip to main content

> ML_LITERATURE // JIMENEZ-2024-SWE-BENCH-CAN-LANGUAGE-MODELS-RESOLVE-REAL-WORLD-GITHUB-ISSUES_v1.0

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, Karthik Narasimhan · International Conference on Learning Representations (ICLR) (2024)

benchmark2024industry-standardartifactsAvailable

Principal Contribution

Created SWE-bench: 2,294 real-world software engineering issues from top open-source Python repositories evaluated via execution of full test suites.

Operational Relevance

Directly guides deployment choices and architecture selection for task-code-generation, task-autonomous-agents.

Assumptions

  • Standard empirical regularity and statistical stability hold across evaluation domains

Limitations

  • Performance characteristics depend on domain distribution and compute allocation parameters

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: