Skip to main content

> ML_DATASET // HUMANEVAL-DATASET_v1.0

HumanEval: Handcrafted Python Coding Benchmark

OpenAI (Chen et al. 2021) · Code Generation & Functional Correctness · 164 handwritten Python programming problems with docstrings, solutions, and unit test cases

Code Generation & Functional CorrectnessMIT164 handwritten Python programming problems with docstrings, solutions, and unit test casesopen

Dataset Profile & Characteristics

Label Type:Docstring specification and test assertions evaluated via pass@k execution
Languages:en, code
License Tier:permissive-open-source
Modalities:text, code

Intended Use

  • Standard functional correctness evaluation of code generation LLMs (pass@1)

Prohibited / Discouraged Use

  • Single definitive measure of full-repository software engineering ability

Bias, Leakage & Privacy Risk Analysis

Privacy / Sensitive Data Risks:

Handcrafted algorithm puzzles; zero personal data.

Known Bias:

Small scope (164 problems); Python standard library arithmetic and string manipulation focus.

Known Benchmark Leakage:

Severe memorization leakage in GitHub-scraped pretraining corpora.

Compatible Tools & Libraries