> ML_LITERATURE // WANG-2018-GLUE-MULTI-TASK-BENCHMARK-ANALYSIS-PLATFORM-NLU_v1.0
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, Samuel R. Bowman · BlackboxNLP Workshop at EMNLP (2018)
benchmark2018industry-standardartifactsAvailable
Principal Contribution
Aggregated 9 diverse NLU tasks (MNLI, QQP, SST-2, CoLA, STS-B, etc.) with a single unified evaluation leaderboard for general language representations.
Operational Relevance
Directly guides deployment choices and architecture selection for task-text-classification, task-natural-language-inference.
Assumptions
- Standard empirical regularity and statistical stability hold across evaluation domains
Limitations
- Performance characteristics depend on domain distribution and compute allocation parameters
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
