Skip to main content

> ML_LITERATURE // HENDRYCKS-2020-MEASURING-MASSIVE-MULTITASK-LANGUAGE-UNDERSTANDING-MMLU_v1.0

Measuring Massive Multitask Language Understanding (MMLU)

Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, Jacob Steinhardt · International Conference on Learning Representations (ICLR) (2020)

benchmark2020industry-standardartifactsAvailable

Principal Contribution

Created the premier multidisciplinary knowledge benchmark covering 57 subjects across STEM, humanities, social sciences, and professional exams (law, medicine).

Operational Relevance

Directly guides deployment choices and architecture selection for task-text-generation, task-knowledge-evaluation.

Assumptions

  • Standard empirical regularity and statistical stability hold across evaluation domains

Limitations

  • Performance characteristics depend on domain distribution and compute allocation parameters

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: