Skip to main content

> ML_DATASETS_ATLAS_v1.0

AI Datasets Atlas

103 qualified AI datasets across tabular, vision, audio, NLP, and multimodal domains: verified licensing tiers, privacy risks, bias, and contamination audits.

Showing 18 of 103 Qualified Datasets (Page 2 of 6)Audited Licensing & Privacy Bounds
Dataset
MITE-commerce & Computer Vision

Fashion-MNIST (Xiao, Rasul & Vollgraf 2017)

E-commerce & Computer Vision · 70,000 28x28 grayscale images across 10 clothing categories · Zalando Research

image
Dataset
Non-Commercial Research LicenseFace Recognition & Attribute Prediction

CelebFaces Attributes Dataset (CelebA)

Face Recognition & Attribute Prediction · 202,599 face images of 10,177 celebrity identities with 40 binary attributes · Chinese University of Hong Kong (Liu et al.)

Implementing Tools:
image
Dataset
Custom Research Non-Commercial LicenseAutonomous Driving & Urban Perception

Cityscapes: Semantic Understanding of Urban Street Scenes

Autonomous Driving & Urban Perception · 5,000 fine-annotated stereo images across 50 cities, 20,000 coarsely annotated frames · Daimler AG / Max Planck Institute / TU Darmstadt (Cordts et al.)

Implementing Tools:
imagevideo
Dataset
BSD-3-Clause (Annotations)Computer Vision & Scene Understanding

ADE20K Scene Parsing Benchmark

Computer Vision & Scene Understanding · 22,210 densely annotated images spanning 150 semantic object and stuff categories · MIT CSAIL (Zhou et al.)

image
Dataset
CC-BY-NC-SA-3.0Autonomous Driving & Robotics

KITTI Vision Benchmark Suite (Geiger et al. 2012)

Autonomous Driving & Robotics · 6 hours of driving data with synchronized color cameras, Velodyne LiDAR, and GPS/IMU · Karlsruhe Institute of Technology (KIT) / Toyota Technological Institute (TTIC)

Implementing Tools:
imagevideo3d pointclouds
Dataset
Custom Research LicenseFoundation Vision Segmentation

SA-1B: Segment Anything 1-Billion Masks (Meta AI)

Foundation Vision Segmentation · 11 million high-resolution licensed images, 1.1 billion segmentation masks · Meta AI Research (Kirillov et al.)

Implementing Tools:
image
Dataset
CC-BY-4.0 (Annotations) / CC-BY-2.0 (Flickr Images)Large-Scale Computer Vision

Open Images Dataset V7 (Google Research)

Large-Scale Computer Vision · 9 million images with 16 million bounding boxes across 600 classes and localized narratives · Google Research

Implementing Tools:
image
Dataset
Open for Non-Commercial ResearchFace Verification Benchmark

Labeled Faces in the Wild (LFW)

Face Verification Benchmark · 13,233 face images of 5,749 people collected from the web · University of Massachusetts, Amherst (Huang et al.)

Implementing Tools:
image
Dataset
Non-commercial research onlyReal-World Optical Character Recognition

The Street View House Numbers (SVHN)

Real-World Optical Character Recognition · 600,000+ 32x32 color digit images cropped from Google Street View house numbers · Stanford University / Google (Netzer et al.)

Implementing Tools:
image
Dataset
CC-BY-SA-4.0Natural Language Processing & Reading Comprehension

SQuAD 2.0 (Stanford Question Answering Dataset)

Natural Language Processing & Reading Comprehension · 150,000+ questions across 500+ Wikipedia articles · Stanford NLP Group (Rajpurkar et al.)

Implementing Tools:
text
Dataset
MITMathematical Reasoning & Code Generation

GSM8K (Grade School Math 8K)

Mathematical Reasoning & Code Generation · 8,500 grade school math word problems with step-by-step solutions · OpenAI (Cobbe et al.)

Implementing Tools:
text
Dataset
ODC-BY / Common Crawl TermsFoundation Model Pretraining

C4 (Colossal Clean Crawled Corpus)

Foundation Model Pretraining · 750 GB clean English text (roughly 156 billion tokens) · Google Research / Allen Institute for AI (Raffel et al.)

Implementing Tools:
text
Dataset
Open Data Commons Attribution License (ODC-By)Foundation LLM Pretraining

FineWeb (15 Trillion High-Quality Web Tokens)

Foundation LLM Pretraining · 15 trillion tokens (44 TB of compressed text across 96 Common Crawl snapshots) · Hugging Face (Lozhkov et al.)

Implementing Tools:
text
Dataset
Custom Academic Open / Component-specific licensesFoundation Model Pretraining

The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Foundation Model Pretraining · 825 GB diverse text (roughly 300 billion tokens) across 22 diverse subsets · EleutherAI (Gao et al.)

Implementing Tools:
text
Dataset
Apache-2.0 (Dataset Tooling) / Permissive Open Source (Code Licenses)Code Intelligence & Program Synthesis

The Stack v2: The Next Generation Code Pretraining Dataset

Code Intelligence & Program Synthesis · 67.5 TB raw source code across 600+ programming languages · BigCode Project / Software Heritage / Hugging Face

Implementing Tools:
textcode
Dataset
MITRLHF & Preference Alignment

UltraFeedback: High-Quality Multi-Domain Preference Dataset

RLHF & Preference Alignment · 64,000 prompts with 256,000 model completions evaluated by GPT-4 · OpenBMB / Tsinghua University (Cui et al.)

Implementing Tools:
text
Dataset
LMSYS Dataset License (Non-Commercial Research)Conversational AI & User Intent

LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Conversational AI & User Intent · 1 million real-world user conversations with 25 frontier LLMs collected from Chatbot Arena · LMSYS Org / UC Berkeley (Zheng et al.)

Implementing Tools:
text
Dataset
MITAdvanced Mathematical Reasoning

MATH: Measuring Mathematical Problem Solving (Hendrycks et al. 2021)

Advanced Mathematical Reasoning · 12,500 high school competition math problems (AMC 10, AMC 12, AIME) · UC Berkeley (Dan Hendrycks et al.)

Implementing Tools:
text