> ML_LITERATURE_ATLAS_v1.0
Research Literature Atlas
253 qualified literature records from foundational statistical learning to frontier reasoning LLMs: verified DOIs, arXiv IDs, and original bilingual syntheses.
Model Cards for Model Reporting
Foundational transparency paper establishing Model Cards, adopted globally by Hugging Face, Google, Meta, and ISO/IEC AI governance frameworks.
Constitutional AI: Harmlessness from AI Feedback (RLAIF)
Anthropic landmark paper establishing Constitutional AI and RLAIF, the primary paradigm powering safe, helpful, and honest assistants like Claude.
Extracting Training Data from Large Language Models
USENIX Security award-winning paper proving that large language models leak memorized sensitive training data, establishing memorization risks in generative AI.
ImageNet: A Large-Scale Hierarchical Image Database
The most influential dataset in computer science history, whose annual challenge (ILSVRC) directly catalyzed modern deep learning and GPU acceleration.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Foundational evaluation benchmark paper that drove the development and comparison of BERT, RoBERTa, and early pre-trained Transformer language models.
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Major benchmark successor ensuring rigorous tracking of linguistic reasoning and multi-sentence comprehension in advanced foundation models.
Measuring Massive Multitask Language Understanding (MMLU)
The global standard benchmark used in every major LLM release (GPT-4, Claude 3.5, Gemini, Llama-3) to measure broad world knowledge and academic reasoning.
Training Verifiers to Solve Math Word Problems (GSM8K)
Foundational reasoning benchmark establishing grade-school math as the proving ground for chain-of-thought prompting and mathematical verification.
Evaluating Large Language Models Trained on Code (Codex / HumanEval)
Landmark paper formalizing code generation evaluation through unit test execution rather than surface-level BLEU matching, powering developer AI tooling.
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
The defining benchmark for autonomous coding agents, measuring whether LLMs can resolve complex repository-level GitHub issues and pass unit test regressions.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Meta foundation paper detailing production-grade pretraining at 2T tokens, Grouped-Query Attention (GQA), and safety tuning for conversational assistants.
The Llama 3 Herd of Models
Definitive comprehensive technical report detailing large-scale GPU cluster infrastructure, pipeline parallelism, high-quality synthetic data pipelines, and post-training alignment.
Mistral 7B
Seminal paper establishing Mistral 7B as the premier compact open model, pioneering efficient long-context processing with sub-quadratic attention memory.
Mixtral of Experts
The landmark sparse MoE paper that democratized high-capacity, low-latency inference for open models on consumer and cloud hardware.
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Pioneering architecture drastically cutting generation inference memory footprint via low-rank latent KV projections while achieving top-tier benchmark efficiency.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Revolutionary 2025 reasoning paper detailing Group Relative Policy Optimization (GRPO), cold-start distillation, and multi-stage alignment matching proprietary frontier reasoning performance.
Gemma: Open Models Based on Gemini Research and Technology
Google open weights foundation release bringing frontier research architecture and responsible AI safety evaluation to open-source developer workflows.
KTO: Model Alignment as Prospect Theoretic Optimization
Influential ICML paper grounding language model alignment in behavioural economics (Prospect Theory), matching DPO performance using realistic unpaired real-world feedback.
