> ML_LITERATURE_ATLAS_v1.0
Research Literature Atlas
253 qualified literature records from foundational statistical learning to frontier reasoning LLMs: verified DOIs, arXiv IDs, and original bilingual syntheses.
Scalable Diffusion Models with Transformers (DiT)
The landmark vision paper introducing Diffusion Transformers (DiT), proving that generative diffusion follows transformer compute-scaling laws and powering Sora, SD3, and Flux.
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
The landmark ECCV paper founding the neural radiance field (NeRF) revolution in 3D computer vision and computer graphics.
3D Gaussian Splatting for Real-Time Radiance Field Rendering
Monumental SIGGRAPH paper establishing 3D Gaussian Splatting, transforming novel view synthesis from minutes-per-frame to real-time interactive rendering.
Classifier-Free Diffusion Guidance (CFG)
Foundational generative algorithm establishing Classifier-Free Guidance (CFG), universal across modern text-to-image and text-to-video systems (DALL-E 2/3, Stable Diffusion, Midjourney).
SGLang: Efficient Execution of Structured Language Model Programs
Foundational serving paper introducing SGLang and RadixAttention, enabling radix-tree caching of KV tensors across branched generation and chained prompting.
llama.cpp: Inference of LLaMA model in pure C/C++
The landmark open-source systems achievement that democratized consumer-hardware and Apple Silicon LLM inference worldwide, powering Ollama and on-device AI apps.
Accelerating the Machine Learning Lifecycle with MLflow
The landmark Databricks paper introducing MLflow, the most widely deployed open-source MLOps platform for tracking experiments, artifacts, and model deployments.
"Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI
Award-winning empirical study from Google demonstrating that poor data collection and documentation cause 90%+ of AI failures in clinical, agricultural, and safety systems.
"Why Should I Trust You?": Explaining the Predictions of Any Classifier (LIME)
The landmark KDD paper founding modern model-agnostic local explainability, enabling auditing of computer vision, tabular, and text predictions for debugging and compliance.
A Unified Approach to Interpreting Model Predictions (SHAP)
Monumental NeurIPS paper establishing SHAP values as the gold standard for feature attribution in credit scoring, clinical diagnostics, and regulated machine learning.
Axiomatic Attribution for Deep Networks (Integrated Gradients)
Landmark ICML paper defining axiomatic attribution for deep neural networks, widely used across Google Cloud Explainable AI and PyTorch Captum.
Deep Learning with Differential Privacy (DP-SGD)
The landmark CCS security paper proving that deep neural networks can be trained with mathematically rigorous (epsilon, delta)-differential privacy guarantees against membership inference.
Quantifying Memorization Across Neural Language Models
Major ICLR study evaluating memorization across EleutherAI Pythia models, providing formal scaling laws for copyright, privacy, and deduplication in AI training corpora.
Towards Deep Learning Models Resistant to Adversarial Attacks (PGD Training)
The landmark MIT paper establishing PGD adversarial training, the cornerstone defense against imperceptible input perturbations and evasion attacks.
Equality of Opportunity in Supervised Learning
Landmark NeurIPS algorithmic fairness paper providing mathematically sound post-processing techniques to eliminate disparate impact without harming overall model accuracy.
Datasheets for Datasets
Landmark paper establishing "Datasheets for Datasets", now adopted across industry, academic conferences, and global regulatory standards including the EU AI Act.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜
High-impact FAccT paper warning of the risks of uncurated internet pre-training, sparking widespread industry reform around data provenance and environmental accounting.
