Skip to main content

> ML_DATASET // LIBRISPEECH-ASR-CORPUS_v1.0

LibriSpeech ASR Corpus (Panayotov et al. 2015)

Johns Hopkins University / Daniel Povey · Speech Recognition & Acoustic Modeling · 1,000 hours of 16kHz read English audiobooks with aligned text transcripts

Speech Recognition & Acoustic ModelingCC-BY-4.01,000 hours of 16kHz read English audiobooks with aligned text transcriptsopen

Dataset Profile & Characteristics

Label Type:Word-level and sentence-level clean/other acoustic transcripts
Languages:en
License Tier:permissive-open-source
Modalities:audio, text

Intended Use

  • Standard benchmark for automatic speech recognition (ASR) acoustic models and wav2vec 2.0 / Whisper evaluations

Prohibited / Discouraged Use

  • Conversational spontaneous dialogue recognition in high-noise industrial settings

Bias, Leakage & Privacy Risk Analysis

Privacy / Sensitive Data Risks:

Volunteer audio recordings from LibriVox project; public domain audiobooks.

Known Bias:

Read literary English audio; low background noise; older demographic accents.

Known Benchmark Leakage:

Official test-clean and test-other splits strictly partitioned by speaker identity.

Compatible Tools & Libraries