> ML_DATASET // THE-STACK-V2-CODE_v1.0
The Stack v2: The Next Generation Code Pretraining Dataset
BigCode Project / Software Heritage / Hugging Face · Code Intelligence & Program Synthesis · 67.5 TB raw source code across 600+ programming languages
Code Intelligence & Program SynthesisApache-2.0 (Dataset Tooling) / Permissive Open Source (Code Licenses)67.5 TB raw source code across 600+ programming languagesopen
Dataset Profile & Characteristics
Label Type:Permissively licensed code repositories with metadata, Git commits, and docstrings
Languages:en, code
License Tier:permissive-open-source
Modalities:text, code
Intended Use
- Training code foundation models (StarCoder 2) with strict opt-out compliance
Prohibited / Discouraged Use
- Generating malware or exploiting zero-day vulnerabilities
Bias, Leakage & Privacy Risk Analysis
Privacy / Sensitive Data Risks:
Scanned for hardcoded secrets, API tokens, private keys, and developer emails.
Known Bias:
Dominated by popular languages (JavaScript, Python, Java, C++, Go, Rust).
Known Benchmark Leakage:
Contamination filters applied against HumanEval, MBPP, and SWE-bench.
