> ML_LITERATURE_ATLAS_v1.0
Research Literature Atlas
253 qualified literature records from foundational statistical learning to frontier reasoning LLMs: verified DOIs, arXiv IDs, and original bilingual syntheses.
Communication-Efficient Learning of Deep Networks from Decentralized Data (FedAvg)
The landmark Google paper founding Federated Learning, the primary enterprise and edge architecture for training on decentralized private data (smartphones, hospitals).
Practical Secure Aggregation for Privacy-Preserving Machine Learning
Foundational cryptographic systems paper enabling zero-knowledge parameter aggregation in production federated learning systems across hundreds of millions of edge devices.
Some methods of speeding up the convergence of iteration methods (Heavy-Ball Momentum)
Foundational Soviet mathematical paper originating momentum in optimization, indispensable for training deep neural networks and gradient boosted trees.
A method for solving the convex programming problem with convergence rate O(1/k^2)
Monumental optimization paper proving that standard gradient descent is sub-optimal and constructing the optimal first-order accelerated method.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization (AdaGrad)
The landmark JMLR paper introducing adaptive learning rates, laying the algorithmic groundwork for RMSProp and Adam in deep learning.
Lecture 6.5 - RMSProp: Divide the gradient by a running average of its recent magnitude
The famous unpublished Coursera lecture by Geoffrey Hinton that introduced RMSProp, widely used for training recurrent neural networks and early deep reinforcement learning.
Very Deep Convolutional Networks for Large-Scale Image Recognition (VGG)
Landmark Oxford ICLR paper establishing VGG-16 and VGG-19, proving that depth and small convolution kernels are the fundamental drivers of visual representation quality.
Going Deeper with Convolutions (GoogLeNet / Inception)
Winner of ILSVRC 2014, introducing GoogLeNet and proving that architectural width and multi-scale filtering achieve superior efficiency over naive stacking.
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
Landmark Google ICML paper establishing the EfficientNet family (B0 to B7), achieving state-of-the-art accuracy with up to 8.4x fewer parameters.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Best Paper Award winner at ICCV 2021, establishing the Swin Transformer as the primary general-purpose vision backbone for dense prediction tasks.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
The landmark NeurIPS paper that unified region proposal and object detection into a single deep network, forming the foundation of modern two-stage detectors.
Mask R-CNN
Winner of ICCV 2017 Marr Prize, establishing Mask R-CNN as the global reference architecture for instance segmentation and human pose estimation.
End-to-End Object Detection with Transformers (DETR)
Landmark ECCV paper establishing DETR, replacing handcrafted pipelines in object detection with pure encoder-decoder Transformer architectures.
U-Net: Convolutional Networks for Biomedical Image Segmentation
Monumental medical imaging paper introducing U-Net, the foundational backbone for biomedical image analysis and later the core architectural engine of diffusion models.
Flamingo: a Visual Language Model for Few-Shot Learning
Landmark DeepMind NeurIPS paper establishing the modern architecture for visual language models (VLMs), enabling few-shot open-ended visual dialogue and reasoning.
Visual Instruction Tuning (LLaVA)
The landmark open-source visual instruction tuning paper introducing LLaVA, establishing the dominant blueprint for multimodal conversational open-weights assistants.
WaveNet: A Generative Model for Raw Audio
The historic DeepMind paper creating WaveNet, transforming synthetic speech from unnatural concatenative units into human-sounding speech.
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
The premier open-source neural vocoder architecture establishing HiFi-GAN, universal across production text-to-speech systems for fast mel-spectrogram-to-audio conversion.
