Skip to main content

> ML_LITERATURE_ATLAS_v1.0

Research Literature Atlas

253 qualified literature records from foundational statistical learning to frontier reasoning LLMs: verified DOIs, arXiv IDs, and original bilingual syntheses.

Showing 18 of 253 Qualified Records (Page 6 of 15)Verified Academic Citations
Literature Record
2019 · ACMstandard-technical-report

Model Cards for Model Reporting

Foundational transparency paper establishing Model Cards, adopted globally by Hugging Face, Google, Meta, and ISO/IEC AI governance frameworks.

model governancefairness audit
Literature Record
2022 · arXivsafety-fairness

Constitutional AI: Harmlessness from AI Feedback (RLAIF)

Anthropic landmark paper establishing Constitutional AI and RLAIF, the primary paradigm powering safe, helpful, and honest assistants like Claude.

llm alignment rlhfai safety
Literature Record
2021 · USENIXsafety-fairness

Extracting Training Data from Large Language Models

USENIX Security award-winning paper proving that large language models leak memorized sensitive training data, establishing memorization risks in generative AI.

privacy preserving ml
Literature Record
2009 · IEEEdataset

ImageNet: A Large-Scale Hierarchical Image Database

The most influential dataset in computer science history, whose annual challenge (ILSVRC) directly catalyzed modern deep learning and GPU acceleration.

image classification
Literature Record
2018 · BlackboxNLPbenchmark

GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Foundational evaluation benchmark paper that drove the development and comparison of BERT, RoBERTa, and early pre-trained Transformer language models.

text classificationnatural language inference
Literature Record
2019 · Advancesbenchmark

SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Major benchmark successor ensuring rigorous tracking of linguistic reasoning and multi-sentence comprehension in advanced foundation models.

natural language inferencequestion answering
Literature Record
2020 · Internationalbenchmark

Measuring Massive Multitask Language Understanding (MMLU)

The global standard benchmark used in every major LLM release (GPT-4, Claude 3.5, Gemini, Llama-3) to measure broad world knowledge and academic reasoning.

text generationknowledge evaluation
Literature Record
2021 · arXivbenchmark

Training Verifiers to Solve Math Word Problems (GSM8K)

Foundational reasoning benchmark establishing grade-school math as the proving ground for chain-of-thought prompting and mathematical verification.

mathematical reasoningtext generation
Literature Record
2021 · arXivbenchmark

Evaluating Large Language Models Trained on Code (Codex / HumanEval)

Landmark paper formalizing code generation evaluation through unit test execution rather than surface-level BLEU matching, powering developer AI tooling.

code generation
Literature Record
2024 · Internationalbenchmark

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

The defining benchmark for autonomous coding agents, measuring whether LLMs can resolve complex repository-level GitHub issues and pass unit test regressions.

code generationautonomous agents
Literature Record
2023 · arXivseminal-architecture

Llama 2: Open Foundation and Fine-Tuned Chat Models

Meta foundation paper detailing production-grade pretraining at 2T tokens, Grouped-Query Attention (GQA), and safety tuning for conversational assistants.

text generationquestion answering
Literature Record
2024 · arXivseminal-architecture

The Llama 3 Herd of Models

Definitive comprehensive technical report detailing large-scale GPU cluster infrastructure, pipeline parallelism, high-quality synthetic data pipelines, and post-training alignment.

text generationcode generationmultimodal
Literature Record
2023 · arXivseminal-architecture

Mistral 7B

Seminal paper establishing Mistral 7B as the premier compact open model, pioneering efficient long-context processing with sub-quadratic attention memory.

text generationcode generation
Literature Record
2024 · arXivseminal-architecture

Mixtral of Experts

The landmark sparse MoE paper that democratized high-capacity, low-latency inference for open models on consumer and cloud hardware.

text generationcode generation
Literature Record
2024 · arXivseminal-architecture

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Pioneering architecture drastically cutting generation inference memory footprint via low-rank latent KV projections while achieving top-tier benchmark efficiency.

text generationcode generation
Literature Record
2025 · arXivalgorithm

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Revolutionary 2025 reasoning paper detailing Group Relative Policy Optimization (GRPO), cold-start distillation, and multi-stage alignment matching proprietary frontier reasoning performance.

text generationcode generation
Literature Record
2024 · arXivseminal-architecture

Gemma: Open Models Based on Gemini Research and Technology

Google open weights foundation release bringing frontier research architecture and responsible AI safety evaluation to open-source developer workflows.

text generationquestion answering
Literature Record
2024 · Internationalalgorithm

KTO: Model Alignment as Prospect Theoretic Optimization

Influential ICML paper grounding language model alignment in behavioural economics (Prospect Theory), matching DPO performance using realistic unpaired real-world feedback.

text generation
Showing 91–108 of 253 Literature Records