> ML_DATASETS_ATLAS_v1.0
AI Datasets Atlas
103 qualified AI datasets across tabular, vision, audio, NLP, and multimodal domains: verified licensing tiers, privacy risks, bias, and contamination audits.
Fashion-MNIST (Xiao, Rasul & Vollgraf 2017)
E-commerce & Computer Vision · 70,000 28x28 grayscale images across 10 clothing categories · Zalando Research
CelebFaces Attributes Dataset (CelebA)
Face Recognition & Attribute Prediction · 202,599 face images of 10,177 celebrity identities with 40 binary attributes · Chinese University of Hong Kong (Liu et al.)
Cityscapes: Semantic Understanding of Urban Street Scenes
Autonomous Driving & Urban Perception · 5,000 fine-annotated stereo images across 50 cities, 20,000 coarsely annotated frames · Daimler AG / Max Planck Institute / TU Darmstadt (Cordts et al.)
ADE20K Scene Parsing Benchmark
Computer Vision & Scene Understanding · 22,210 densely annotated images spanning 150 semantic object and stuff categories · MIT CSAIL (Zhou et al.)
KITTI Vision Benchmark Suite (Geiger et al. 2012)
Autonomous Driving & Robotics · 6 hours of driving data with synchronized color cameras, Velodyne LiDAR, and GPS/IMU · Karlsruhe Institute of Technology (KIT) / Toyota Technological Institute (TTIC)
SA-1B: Segment Anything 1-Billion Masks (Meta AI)
Foundation Vision Segmentation · 11 million high-resolution licensed images, 1.1 billion segmentation masks · Meta AI Research (Kirillov et al.)
Open Images Dataset V7 (Google Research)
Large-Scale Computer Vision · 9 million images with 16 million bounding boxes across 600 classes and localized narratives · Google Research
Labeled Faces in the Wild (LFW)
Face Verification Benchmark · 13,233 face images of 5,749 people collected from the web · University of Massachusetts, Amherst (Huang et al.)
The Street View House Numbers (SVHN)
Real-World Optical Character Recognition · 600,000+ 32x32 color digit images cropped from Google Street View house numbers · Stanford University / Google (Netzer et al.)
SQuAD 2.0 (Stanford Question Answering Dataset)
Natural Language Processing & Reading Comprehension · 150,000+ questions across 500+ Wikipedia articles · Stanford NLP Group (Rajpurkar et al.)
GSM8K (Grade School Math 8K)
Mathematical Reasoning & Code Generation · 8,500 grade school math word problems with step-by-step solutions · OpenAI (Cobbe et al.)
C4 (Colossal Clean Crawled Corpus)
Foundation Model Pretraining · 750 GB clean English text (roughly 156 billion tokens) · Google Research / Allen Institute for AI (Raffel et al.)
FineWeb (15 Trillion High-Quality Web Tokens)
Foundation LLM Pretraining · 15 trillion tokens (44 TB of compressed text across 96 Common Crawl snapshots) · Hugging Face (Lozhkov et al.)
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Foundation Model Pretraining · 825 GB diverse text (roughly 300 billion tokens) across 22 diverse subsets · EleutherAI (Gao et al.)
The Stack v2: The Next Generation Code Pretraining Dataset
Code Intelligence & Program Synthesis · 67.5 TB raw source code across 600+ programming languages · BigCode Project / Software Heritage / Hugging Face
UltraFeedback: High-Quality Multi-Domain Preference Dataset
RLHF & Preference Alignment · 64,000 prompts with 256,000 model completions evaluated by GPT-4 · OpenBMB / Tsinghua University (Cui et al.)
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
Conversational AI & User Intent · 1 million real-world user conversations with 25 frontier LLMs collected from Chatbot Arena · LMSYS Org / UC Berkeley (Zheng et al.)
MATH: Measuring Mathematical Problem Solving (Hendrycks et al. 2021)
Advanced Mathematical Reasoning · 12,500 high school competition math problems (AMC 10, AMC 12, AIME) · UC Berkeley (Dan Hendrycks et al.)
