> ML_LITERATURE_ATLAS_v1.0
Research Literature Atlas
253 qualified literature records from foundational statistical learning to frontier reasoning LLMs: verified DOIs, arXiv IDs, and original bilingual syntheses.
Improving Language Understanding by Generative Pre-Training (GPT-1)
Pioneering OpenAI report establishing that left-to-right autoregressive language modeling creates powerful general representations.
Language Models are Unsupervised Multitask Learners (GPT-2)
Seminal OpenAI report introducing GPT-2 and WebText, demonstrating that language models naturally become general multi-task learners at scale.
Language Models are Few-Shot Learners (GPT-3)
Monumental NeurIPS paper demonstrating that 175-billion parameter GPT-3 exhibits remarkable in-context few-shot performance across hundreds of diverse tasks.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Rigorous 140-page empirical benchmark study introducing T5 and the Colossal Clean Crawled Corpus (C4), standardizing transfer learning across NLP.
LLaMA: Open and Efficient Foundation Language Models
Historic Meta paper releasing the LLaMA foundation models, proving that 7B-65B models trained on trillions of tokens achieve GPT-3 competitive parity on consumer hardware.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Groundbreaking IO-aware systems paper accelerating attention by 2x-4x and cutting memory footprint by 10x without any mathematical approximation.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Landmark SOSP systems paper solving the GPU memory bottleneck in LLM serving through virtual memory paging of key-value caches.
Training language models to follow instructions with human feedback
Seminal NeurIPS paper showing that a 1.3B aligned InstructGPT model is preferred by humans over an unaligned 175B GPT-3, launching conversational AI.
Training Compute-Optimal Large Language Models (Chinchilla)
DeepMind landmark study proving that existing large models were severely undertrained and that equal scaling of parameters and tokens is compute-optimal.
Very Deep Convolutional Networks for Large-Scale Image Recognition (VGG)
Influential ICLR paper establishing the VGG architecture, proving that simple deep stacks of 3x3 convolutions outperform complex heterogeneous filter sizes.
Going Deeper with Convolutions (GoogLeNet / Inception)
CVPR 2015 winning architecture GoogLeNet, demonstrating that multi-scale factorized convolutions achieve high accuracy with 12x fewer parameters than AlexNet.
MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Foundational mobile computer vision paper introducing MobileNets, popularizing depthwise separable convolutions for edge and smartphone deployment.
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
Award-winning ICML paper showing that compound scaling of depth, width, and resolution yields EfficientNet models that surpass existing CNNs by 8x.
You Only Look Once: Unified, Real-Time Object Detection (YOLO)
Seminal CVPR paper creating the real-time object detection paradigm, running at 45 frames per second on standard GPUs and transforming computer vision.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT)
Revolutionary ICLR paper proving that pure Transformer architectures without convolutional layers achieve state-of-the-art vision results at scale.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Marr Prize-winning paper introducing Swin Transformer, achieving linear computational complexity in vision transformers via shifted window hierarchical processing.
Learning Transferable Visual Models From Natural Language Supervision (CLIP)
Breakthrough OpenAI paper introducing CLIP, matching ResNet-50 ImageNet accuracy with zero training samples by predicting which caption matches which image.
Auto-Encoding Variational Bayes (VAE)
Foundational ICLR paper introducing Variational Autoencoders, establishing the reparameterization trick to allow backpropagation through stochastic latent nodes.
