> ML_LITERATURE // BAEVSKI-2020-WAV2VEC-2-FRAMEWORK-SELF-SUPERVISED-LEARNING-SPEECH_v1.0
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael Auli · Advances in Neural Information Processing Systems (NeurIPS) (2020)
seminal-architecture2020industry-standardthirdPartyReproduced
Principal Contribution
Masked latent speech representations across time steps and trained a transformer with contrastive loss over quantized representations, achieving state-of-the-art ASR with 10 minutes of labels.
Operational Relevance
Serves as qualified reference for implementing task-speech-recognition in production systems.
Assumptions
- Underlying spatio-temporal continuity and domain distributional stability hold
Limitations
- Performance scaling and computational footprint depend on receptive field depth, sequence length, and resolution
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
