Skip to main content

> ML_LITERATURE // CHEN-2021-EVALUATING-LARGE-LANGUAGE-MODELS-TRAINED-ON-CODE-HUMANEVAL_v1.0

Evaluating Large Language Models Trained on Code (Codex / HumanEval)

Mark Chen, Jerry Tworek, Heewoo Jun, Sheng Shen, Christine Zheng Ponder, Ronan McGrew, Lilian Weng, Hao Shen, Alex Ray, Heidy Khlaaf, Girish Sastry, Gillian Hadfield, Alec Radford, Ilya Sutskever, Wojciech Zaremba · arXiv preprint (2021)

benchmark2021industry-standardartifactsAvailable

Principal Contribution

Introduced the HumanEval benchmark and pass@k functional correctness evaluation, launching OpenAI Codex and GitHub Copilot.

Operational Relevance

Directly guides deployment choices and architecture selection for task-code-generation.

Assumptions

  • Standard empirical regularity and statistical stability hold across evaluation domains

Limitations

  • Performance characteristics depend on domain distribution and compute allocation parameters

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: