> ML_LITERATURE_ATLAS_v1.0
Research Literature Atlas
253 qualified literature records from foundational statistical learning to frontier reasoning LLMs: verified DOIs, arXiv IDs, and original bilingual syntheses.
GLU Variants Improve Transformer
Foundational architecture paper establishing SwiGLU activation functions, standard in modern architectures including LLaMA, PaLM, and Mistral.
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
The landmark selective state space architecture paper that challenged transformer supremacy, providing quadratic-free long context processing with constant memory states.
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Monumental ICML paper introducing Mamba-2, achieving 8x faster training throughput than Mamba-1 by mapping state updates directly onto GPU Tensor Cores via duality theory.
RWKV: Reinventing RNNs for the Transformer Era
Foundational open sequence modeling paper introducing RWKV, scaling recurrent architectures to 14B parameters with 100% linear attention complexity.
Retentive Network: A Successor to Transformer for Large Language Models (RetNet)
Influential Microsoft Research paper proposing the "impossible triangle" solution: achieving parallel training, O(1) inference cost, and linear long-context modeling simultaneously.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Landmark JMLR systems paper presenting Switch Transformers, establishing techniques for stable, low-communication large-scale sparse mixture-of-experts training.
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
The landmark Google systems paper introducing GShard, scaling multilingual neural machine translation to 600B parameters across 2,048 TPU v3 accelerators.
Code Llama: Open Foundation Models for Code
Foundational Meta coding model report establishing state-of-the-art open code generation, Python specialization, and instruction-tuned code assistants.
StarCoder 2 and The Stack v2: The Next Generation
BigCode Consortium open research paper establishing ethical, transparent code pretraining datasets and models for enterprise-grade developer tooling.
Scaling Laws for Neural Language Models
The landmark OpenAI paper formalizing neural scaling laws, guiding billions of dollars in compute allocation for training the modern era of foundation models.
Deep Reinforcement Learning from Human Preferences
The foundational OpenAI / DeepMind paper originating modern RLHF, laying the mathematical groundwork for aligning deep neural networks to human intent.
Improving Language Understanding by Generative Pre-Training (GPT-1)
The original Generative Pre-trained Transformer (GPT) paper that established the generative pre-training + fine-tuning blueprint for natural language processing.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Landmark ACL paper creating BART, establishing the preeminent architecture for text summarization, abstractive generation, and machine comprehension.
PaLM: Scaling Language Modeling with Pathways
Landmark Google systems paper presenting PaLM (540B), detailing multi-pod distributed training, SwiGLU activation, and breakthroughs in reasoning and code.
Training data-efficient image transformers & distillation through attention (DeiT)
Landmark ICML computer vision paper democratizing Vision Transformers by eliminating the requirement for proprietary hundred-million-image training sets.
Robust Speech Recognition via Large-Scale Weak Supervision (Whisper)
The landmark OpenAI paper presenting Whisper, transforming speech recognition from fragile acoustic-phonetic models into a robust foundation model for multilingual transcription.
Segment Anything (SAM)
The landmark ICCV paper creating the Segment Anything Model (SAM), establishing the first general-purpose visual foundation model for image segmentation.
Flow Matching for Generative Modeling
Influential generative modeling paper introducing Flow Matching, the foundational mathematical formulation powering modern fast-sampling models including Stable Diffusion 3 and Flux.1.
