> ML_BENCHMARKS_ATLAS_v1.0
AI Benchmarks & Evaluation Protocols
53 canonical AI evaluation protocols spanning GLUE, SuperGLUE, MMLU, GSM8K, HumanEval, ARC, and LMSYS Chatbot Arena with saturation and contamination audits.
MMLU (Massive Multitask Language Understanding)
5-shot Accuracy (%) · 5-shot multiple-choice question answering across 57 academic subjects (humanities, STEM, social sciences, business).
HumanEval (Python Code Generation Benchmark)
pass@1, pass@10, pass@100 · 164 handcrafted Python programming problems evaluating docstring-to-code generation against unit test test-cases.
SWE-bench (Software Engineering GitHub Issue Benchmark)
Resolved Rate (% of issues passing full test regressions) · Agent receives an entire GitHub repository and natural language issue description, generates git patch diff executed inside Docker container.
GSM8K (Grade School Math Multi-step Reasoning)
Strict Match Accuracy (%) · 8-shot chain-of-thought generation followed by numerical regex answer extraction on 1,319 test problems.
MATH (Hendrycks High School Competition Benchmark)
Exact Match Accuracy on Final Answer (%) · 5-shot chain-of-thought generation over 5,000 challenging competition math problems with SymPy symbolic equivalence checking.
LMSYS Chatbot Arena (Crowdsourced Human Preference Elo)
Bradley-Terry Arena Elo Rating · Blind side-by-side human A/B testing on arbitrary user prompts with randomized model pairings and statistical bootstrapping.
TruthfulQA (Model Falsehood & Hallucination Benchmark)
% Truthful * % Informative (MC1, MC2, Fine-tuned GPT Judge) · 817 questions in zero-shot setting evaluated across multiple-choice selection and free-form generation judged by fine-tuned truthfulness classifier.
ARC Challenge (AI2 Reasoning Challenge 2590 Question Subset)
25-shot Accuracy (%) · Multiple-choice question answering on questions that retrieval algorithms and co-occurrence baselines fail to answer correctly.
ImageNet-1K Top-1 Accuracy (ILSVRC 2012 Validation)
Top-1 Accuracy (%), Top-5 Accuracy (%) · 50,000 validation images evaluated on center crop or multi-crop without test-time augmentation bells and whistles.
MS COCO Object Detection (AP @ IoU=0.50:0.95)
mean Average Precision (mAP 0.50:0.95), AP50, AP75, AP_small, AP_medium, AP_large · Evaluated on 5,000 COCO val2017 images or blind test-dev server averaging AP across 10 IoU thresholds from 0.50 to 0.95 with step 0.05.
MS COCO Instance Segmentation (Mask AP)
Mask mAP (0.50:0.95), Mask AP50, Mask AP75 · Pixel-level polygon mask overlap evaluated against ground truth segmentation masks across 80 object classes.
Cityscapes Semantic Segmentation (mIoU)
mean Intersection-over-Union (mIoU %), iIoU · 500 validation images evaluated on 19 urban driving classes at full 2048x1024 resolution.
LibriSpeech ASR Word Error Rate (WER)
Word Error Rate (WER % on test-clean and test-other) · Standard greedy or beam search decoding evaluated on test-clean and noisy test-other splits without language model rescoring.
LLM Reasoning & Architecture Standardized Protocol 1
Accuracy (%) · Standardized production evaluation protocol 1 executing multi-trial cross-validated metrics on canonical test partitions.
LLM Reasoning & Architecture Standardized Protocol 2
Accuracy (%) · Standardized production evaluation protocol 2 executing multi-trial cross-validated metrics on canonical test partitions.
LLM Reasoning & Architecture Standardized Protocol 3
Accuracy (%) · Standardized production evaluation protocol 3 executing multi-trial cross-validated metrics on canonical test partitions.
LLM Reasoning & Architecture Standardized Protocol 4
Accuracy (%) · Standardized production evaluation protocol 4 executing multi-trial cross-validated metrics on canonical test partitions.
LLM Reasoning & Architecture Standardized Protocol 5
Accuracy (%) · Standardized production evaluation protocol 5 executing multi-trial cross-validated metrics on canonical test partitions.
