> ML_LITERATURE_ATLAS_v1.0
Research Literature Atlas
253 qualified literature records from foundational statistical learning to frontier reasoning LLMs: verified DOIs, arXiv IDs, and original bilingual syntheses.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Foundational reinforcement learning paper that saved massive GPU memory in LLM reasoning alignment, serving as the core mathematical engine later powering DeepSeek-R1.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Landmark Google Brain paper discovering Chain-of-Thought (CoT) prompting, transforming natural language generation into explicit computational deliberation.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Landmark reasoning ensemble technique boosting math and reasoning benchmark accuracy by up to 15% through stochastic path sampling and consensus aggregation.
ReAct: Synergizing Reasoning and Acting in Language Models
The foundational autonomous agent paper establishing ReAct loops, allowing models to query APIs, search databases, and self-correct during dynamic execution.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Influential framework combining classical heuristic search (BFS/DFS/A*) with LLM self-evaluation to solve complex planning tasks like Game of 24 and Crosswords.
Reflexion: Language Agents with Verbal Reinforcement Learning
Influential agent paper demonstrating verbal self-reflection and episodic memory, boosting HumanEval coding performance from 68% to 91% via iterative debugging.
Toolformer: Language Models Can Teach Themselves to Use Tools
Landmark Meta paper demonstrating that language models can autonomously learn when, which, and how to invoke external APIs without manual tool annotations.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
The landmark paper that coined and formalized Retrieval-Augmented Generation (RAG), establishing the dominant enterprise paradigm for grounding LLMs in external knowledge.
Dense Passage Retrieval for Open-Domain Question Answering (DPR)
Landmark Facebook AI research establishing dense vector retrieval for question answering, launching the modern vector database and semantic search industry.
Precise Zero-Shot Dense Retrieval without Relevance Labels (HyDE)
Invented Hypothetical Document Embeddings (HyDE), bridging the semantic modality gap between short abstract questions and lengthy target documents without relevance tuning.
LoRA: Low-Rank Adaptation of Large Language Models
The landmark Microsoft paper establishing Low-Rank Adaptation (LoRA), the ubiquitous industry standard for fine-tuning multi-billion parameter foundation models with zero inference latency overhead.
QLoRA: Efficient Finetuning of Quantized LLMs
Monumental NeurIPS paper democratizing LLM fine-tuning by allowing consumer-grade and single enterprise GPUs to fine-tune massive foundation models.
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Landmark quantization paper establishing GPTQ, the standard method for post-training weight quantization enabling high-throughput LLM deployment on resource-constrained GPUs.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Major MIT research establishing AWQ, widely deployed across vLLM, TensorRT-LLM, and TGI for ultra-low latency, memory-efficient production LLM serving.
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Landmark ICML paper enabling 8-bit weight and 8-bit activation (W8A8) quantization for LLMs, doubling serving throughput and halving memory footprints on production GPUs.
RoFormer: Enhanced Transformer with Rotary Position Embedding
The foundational mathematical paper creating RoPE, adopted by almost all modern LLMs (Llama, Mistral, Qwen, DeepSeek, Gemma) for seamless context extrapolation.
Fast Transformer Decoding: One Write-Head is All You Need
Seminal paper by Noam Shazeer that identified the memory-bandwidth bottleneck of autoregressive transformer decoding, proposing MQA which slashed KV cache size by the number of heads.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
The landmark EMNLP paper establishing GQA, now universally adopted across production LLMs (Llama 2/3, Mistral, DeepSeek) for optimal serving throughput.
