> ML_LITERATURE_ATLAS_v1.0
Research Literature Atlas
253 qualified literature records from foundational statistical learning to frontier reasoning LLMs: verified DOIs, arXiv IDs, and original bilingual syntheses.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
The primary benchmark for evaluating hallucination and epistemic truthfulness in language models, proving that larger models are often less truthful unless explicitly aligned.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
The most influential empirical benchmark of the generative AI era (LMSYS Chatbot Arena), establishing human blind preference battles as the definitive benchmark for frontier LLMs.
Q-learning
The seminal paper that introduced Q-learning, the foundational model-free off-policy reinforcement learning algorithm in computer science.
Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning (REINFORCE)
Foundational paper creating policy gradient methods, forming the theoretical bedrock of modern deep RL (PPO, TRPO, DPO, RLHF).
Human-level control through deep reinforcement learning (Nature DQN)
Historic Nature cover paper that inaugurated the field of Deep Reinforcement Learning, bridging deep convolutional networks and dynamic programming.
Deep Reinforcement Learning with Double Q-learning (Double DQN)
Landmark AAAI paper resolving value overestimation in deep Q-networks, improving stability and score performance across complex environments.
Trust Region Policy Optimization (TRPO)
The landmark ICML paper establishing Trust Region Policy Optimization (TRPO), solving destructive step-size collapse in continuous control robotics and locomotion.
Addressing Function Approximation Error in Actor-Critic Methods (TD3)
Foundational continuous control paper resolving fundamental overestimation bias in actor-critic architectures, establishing TD3 as a standard benchmark in robotic simulation.
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm (AlphaZero)
Groundbreaking Science paper demonstrating that a single general-purpose reinforcement learning algorithm can achieve superhuman mastery across multiple complex strategic domains tabula rasa.
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
Landmark RSS robotics paper establishing Diffusion Policy, the leading paradigm for training dextrous robot manipulation from human demonstrations.
Prioritized Experience Replay (PER)
Foundational ICLR paper introducing Prioritized Experience Replay, standard across DQN and off-policy actor-critic architectures.
Dueling Network Architectures for Deep Reinforcement Learning (Dueling DQN)
Winner of ICML 2016 Best Paper, introducing Dueling DQN to learn which states are valuable without having to learn the effect of each action for each state.
Asynchronous Methods for Deep Reinforcement Learning (A3C)
Landmark DeepMind ICML paper establishing A3C and A2C, enabling high-performance deep reinforcement learning directly on multi-core standard CPUs.
Continuous control with deep reinforcement learning (DDPG)
The landmark ICLR paper establishing DDPG, enabling deep reinforcement learning to solve 20+ continuous physical control tasks in physics simulators.
Hindsight Experience Replay (HER)
Landmark OpenAI paper solving sparse-reward robotic manipulation, enabling robots to master robotic arm pushing, sliding, and pick-and-place without reward shaping.
Rainbow: Combining Improvements in Deep Reinforcement Learning
Influential DeepMind AAAI paper establishing Rainbow, providing comprehensive ablation studies on what drives sample efficiency and score frontiers in discrete deep RL.
Outracing champion Gran Turismo drivers with deep reinforcement learning (GT Sophy)
Historic Nature cover paper presenting Gran Turismo Sophy, solving real-time continuous vehicle dynamics, tire friction physics, and high-speed tactical etiquette.
RT-1: Robotics Transformer for Real-World Control at Scale
Landmark robotics foundation paper demonstrating that large-scale multitask Transformer models generalize robustly to new tasks, environments, and objects.
