Skip to main content

> ML_LITERATURE // BAEVSKI-2020-WAV2VEC-2-FRAMEWORK-SELF-SUPERVISED-LEARNING-SPEECH_v1.0

wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael Auli · Advances in Neural Information Processing Systems (NeurIPS) (2020)

seminal-architecture2020industry-standardthirdPartyReproduced

Principal Contribution

Masked latent speech representations across time steps and trained a transformer with contrastive loss over quantized representations, achieving state-of-the-art ASR with 10 minutes of labels.

Operational Relevance

Serves as qualified reference for implementing task-speech-recognition in production systems.

Assumptions

  • Underlying spatio-temporal continuity and domain distributional stability hold

Limitations

  • Performance scaling and computational footprint depend on receptive field depth, sequence length, and resolution

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: