Skip to main content

> ML_LITERATURE // WANG-2019-SUPERGLUE-STICKIER-BENCHMARK-GENERAL-LANGUAGE-UNDERSTANDING_v1.0

SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, Samuel R. Bowman · Advances in Neural Information Processing Systems (NeurIPS) (2019)

benchmark2019industry-standardartifactsAvailable

Principal Contribution

Designed a substantially harder benchmark replacing GLUE with challenging reasoning tasks (MultiRC, BoolQ, WiC, WSC) after BERT saturated original GLUE.

Operational Relevance

Directly guides deployment choices and architecture selection for task-natural-language-inference, task-question-answering.

Assumptions

  • Standard empirical regularity and statistical stability hold across evaluation domains

Limitations

  • Performance characteristics depend on domain distribution and compute allocation parameters

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: