Skip to main content

> ML_LITERATURE // WANG-2018-GLUE-MULTI-TASK-BENCHMARK-ANALYSIS-PLATFORM-NLU_v1.0

GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, Samuel R. Bowman · BlackboxNLP Workshop at EMNLP (2018)

benchmark2018industry-standardartifactsAvailable

Principal Contribution

Aggregated 9 diverse NLU tasks (MNLI, QQP, SST-2, CoLA, STS-B, etc.) with a single unified evaluation leaderboard for general language representations.

Operational Relevance

Directly guides deployment choices and architecture selection for task-text-classification, task-natural-language-inference.

Assumptions

  • Standard empirical regularity and statistical stability hold across evaluation domains

Limitations

  • Performance characteristics depend on domain distribution and compute allocation parameters

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: