Skip to main content

> ML_DATASET // MOZILLA-COMMON-VOICE_v1.0

Mozilla Common Voice (Multilingual Speech Corpus)

Mozilla Foundation · Multilingual Speech Recognition · 30,000+ hours across 120+ languages recorded by global crowdsourced volunteers

Multilingual Speech RecognitionCC0-1.0 (Public Domain Dedication)30,000+ hours across 120+ languages recorded by global crowdsourced volunteersopen

Dataset Profile & Characteristics

Label Type:Validated crowdsourced transcripts with demographic metadata (age, sex, accent)
Languages:en, es, fr, de, tr, ar, zh, ja, ru, sw
License Tier:permissive-open-source
Modalities:audio, text

Intended Use

  • Training diverse, accent-robust multilingual speech recognition and voice AI systems

Prohibited / Discouraged Use

  • Voice cloning without explicit speaker authorization

Bias, Leakage & Privacy Risk Analysis

Privacy / Sensitive Data Risks:

Volunteer voice recordings with voluntary self-reported demographic information.

Known Bias:

Varying microphone qualities; male voice skew across certain low-resource language subsets.

Known Benchmark Leakage:

Clean split rules prevent single speaker voice leakage between train and test sets.

Compatible Tools & Libraries