> ML_LITERATURE // KONG-2020-HIFI-GAN-GENERATIVE-ADVERSARIAL-NETWORKS-SPEECH-SYNTHESIS_v1.0
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Jungil Kong, Jaehyeon Kim, Jaekyoung Bae · Advances in Neural Information Processing Systems (NeurIPS) (2020)
seminal-architecture2020industry-standardthirdPartyReproduced
Principal Contribution
Designed a multi-period discriminator (MPD) and multi-scale discriminator (MSD) for GAN-based neural vocoding, synthesizing 22.05 kHz audio faster than real-time on CPU.
Operational Relevance
Serves as qualified reference for implementing task-text-to-speech in production systems.
Assumptions
- Underlying spatio-temporal continuity and domain distributional stability hold
Limitations
- Performance scaling and computational footprint depend on receptive field depth, sequence length, and resolution
