> ML_DATASET // MOZILLA-COMMON-VOICE_v1.0
Mozilla Common Voice (Multilingual Speech Corpus)
Mozilla Foundation · Multilingual Speech Recognition · 30,000+ hours across 120+ languages recorded by global crowdsourced volunteers
Multilingual Speech RecognitionCC0-1.0 (Public Domain Dedication)30,000+ hours across 120+ languages recorded by global crowdsourced volunteersopen
Dataset Profile & Characteristics
Label Type:Validated crowdsourced transcripts with demographic metadata (age, sex, accent)
Languages:en, es, fr, de, tr, ar, zh, ja, ru, sw
License Tier:permissive-open-source
Modalities:audio, text
Intended Use
- Training diverse, accent-robust multilingual speech recognition and voice AI systems
Prohibited / Discouraged Use
- Voice cloning without explicit speaker authorization
Bias, Leakage & Privacy Risk Analysis
Privacy / Sensitive Data Risks:
Volunteer voice recordings with voluntary self-reported demographic information.
Known Bias:
Varying microphone qualities; male voice skew across certain low-resource language subsets.
Known Benchmark Leakage:
Clean split rules prevent single speaker voice leakage between train and test sets.
