> ML_DATASET // C4-COLOSSAL-CLEAN-CRAWLED-CORPUS_v1.0
C4 (Colossal Clean Crawled Corpus)
Google Research / Allen Institute for AI (Raffel et al.) · Foundation Model Pretraining · 750 GB clean English text (roughly 156 billion tokens)
Foundation Model PretrainingODC-BY / Common Crawl Terms750 GB clean English text (roughly 156 billion tokens)open
Dataset Profile & Characteristics
Label Type:Unsupervised pre-training text
Languages:en
License Tier:permissive-open-source
Modalities:text
Intended Use
- Pretraining large foundation language models (T5, LLaMA baselines)
Prohibited / Discouraged Use
- Targeting minority dialect representation without additional data
Bias, Leakage & Privacy Risk Analysis
Privacy / Sensitive Data Risks:
Extracted from Common Crawl; may contain residual public internet personal information.
Known Bias:
Heavy English web bias; heuristic filtering removed non-standard dialects and colloquial text.
Known Benchmark Leakage:
Contained evaluation benchmark questions prior to modern de-duplication efforts.
