> ML_DATASET // HUMANEVAL-DATASET_v1.0
HumanEval: Handcrafted Python Coding Benchmark
OpenAI (Chen et al. 2021) · Code Generation & Functional Correctness · 164 handwritten Python programming problems with docstrings, solutions, and unit test cases
Code Generation & Functional CorrectnessMIT164 handwritten Python programming problems with docstrings, solutions, and unit test casesopen
Dataset Profile & Characteristics
Label Type:Docstring specification and test assertions evaluated via pass@k execution
Languages:en, code
License Tier:permissive-open-source
Modalities:text, code
Intended Use
- Standard functional correctness evaluation of code generation LLMs (pass@1)
Prohibited / Discouraged Use
- Single definitive measure of full-repository software engineering ability
Bias, Leakage & Privacy Risk Analysis
Privacy / Sensitive Data Risks:
Handcrafted algorithm puzzles; zero personal data.
Known Bias:
Small scope (164 problems); Python standard library arithmetic and string manipulation focus.
Known Benchmark Leakage:
Severe memorization leakage in GitHub-scraped pretraining corpora.
