> ML_LITERATURE_ATLAS_v1.0
Research Literature Atlas
253 qualified literature records from foundational statistical learning to frontier reasoning LLMs: verified DOIs, arXiv IDs, and original bilingual syntheses.
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
The defining Vision-Language-Action (VLA) paper demonstrating emergent semantic reasoning and zero-shot physical affordance understanding in real robots.
Completely Derandomized Self-Adaptation in Evolution Strategies (CMA-ES)
The landmark evolutionary computation paper establishing CMA-ES as the gold standard algorithm for difficult, rugged, non-linear black-box optimization.
Xception: Deep Learning with Depthwise Separable Convolutions
Landmark CVPR paper by François Chollet creating Xception, establishing depthwise separable convolutions as the standard building block for efficient mobile vision.
MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
The landmark Google paper introducing MobileNetV1, enabling real-time computer vision inference on smartphones, drones, and edge embedded devices.
MobileNetV2: Inverted Residuals and Linear Bottlenecks
Monumental CVPR paper establishing MobileNetV2, the ubiquitous industry standard for mobile and on-device computer vision backbones.
A Simple Framework for Contrastive Learning of Visual Representations (SimCLR)
The landmark Google ICML paper establishing SimCLR, transforming self-supervised visual representation learning through contrastive InfoNCE optimization.
Momentum Contrast for Unsupervised Visual Representation Learning (MoCo)
Landmark Facebook CVPR paper introducing MoCo, enabling large negative sample dictionaries without requiring thousands of GPU memory allocations.
Emerging Properties in Self-Supervised Vision Transformers (DINO)
Landmark ICCV paper creating DINO, demonstrating that self-supervised ViT attention maps discover semantic object boundaries without any human annotation.
DINOv2: Learning Robust Visual Features without Supervision
Monumental Meta paper establishing DINOv2, the foundational vision feature backbone for monocular depth estimation, dense segmentation, and image retrieval.
Masked Autoencoders Are Scalable Vision Learners (MAE)
Landmark CVPR paper creating Masked Autoencoders (MAE), establishing the BERT-style masked autoencoding paradigm as a dominant foundation for scalable computer vision.
Highly accurate protein structure prediction with AlphaFold (AlphaFold 2)
The historic Nature cover paper announcing AlphaFold 2, resolving the 50-year-old protein folding grand challenge in biology and earning the 2024 Nobel Prize in Chemistry.
Zero-Shot Text-to-Image Generation (DALL-E)
The historic ICML paper creating DALL-E, proving that multimodal autoregressive modeling enables creative zero-shot text-to-image synthesis.
Hierarchical Text-Conditional Image Generation with CLIP Latents (DALL-E 2 / unCLIP)
The landmark unCLIP / DALL-E 2 paper from OpenAI, establishing two-stage CLIP latent diffusion for photorealistic text-to-image generation and semantic image variations.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding (Imagen)
Landmark Google NeurIPS paper establishing Imagen and the DrawBench benchmark, proving the decisive role of large language models in text-to-image synthesis.
Video generation models as world simulators (Sora)
The landmark OpenAI technical report presenting Sora, establishing spacetime patch diffusion transformers as scalable world simulators for continuous video generation.
SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size
Foundational edge ML paper establishing SqueezeNet, proving that architectural compression enables high-accuracy deep learning on low-power microcontrollers and FPGAs.
ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
Landmark CVPR mobile vision paper establishing ShuffleNet, delivering state-of-the-art speed-accuracy trade-offs for embedded robotics and mobile phones.
Designing Network Design Spaces (RegNet)
Winner of CVPR 2020 Best Paper, introducing the RegNet design space methodology for scalable, hardware-friendly convolutional architectures.
