> ML_DATASET // LIBRISPEECH-ASR-CORPUS_v1.0
LibriSpeech ASR Corpus (Panayotov et al. 2015)
Johns Hopkins University / Daniel Povey · Speech Recognition & Acoustic Modeling · 1,000 hours of 16kHz read English audiobooks with aligned text transcripts
Speech Recognition & Acoustic ModelingCC-BY-4.01,000 hours of 16kHz read English audiobooks with aligned text transcriptsopen
Dataset Profile & Characteristics
Label Type:Word-level and sentence-level clean/other acoustic transcripts
Languages:en
License Tier:permissive-open-source
Modalities:audio, text
Intended Use
- Standard benchmark for automatic speech recognition (ASR) acoustic models and wav2vec 2.0 / Whisper evaluations
Prohibited / Discouraged Use
- Conversational spontaneous dialogue recognition in high-noise industrial settings
Bias, Leakage & Privacy Risk Analysis
Privacy / Sensitive Data Risks:
Volunteer audio recordings from LibriVox project; public domain audiobooks.
Known Bias:
Read literary English audio; low background noise; older demographic accents.
Known Benchmark Leakage:
Official test-clean and test-other splits strictly partitioned by speaker identity.
