> ML_LITERATURE // ROMBACH-2022-HIGH-RESOLUTION-IMAGE-SYNTHESIS-LATENT-DIFFUSION-MODELS_v1.0
High-Resolution Image Synthesis with Latent Diffusion Models (Stable Diffusion)
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Björn Ommer · IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2022)
seminal-architecture2022industry-standardthirdPartyReproduced
Principal Contribution
Operated diffusion in the lower-dimensional latent space of a pretrained autoencoder with cross-attention conditioning, launching Stable Diffusion.
Operational Relevance
The foundation for the global open-source image generation industry (Stable Diffusion 1.5, SDXL, ControlNet, LoRA).
Assumptions
- Perceptual compression separates high-frequency pixel imperceptible details from semantic conceptual composition in images
Limitations
- Autoencoder decoders can introduce minor facial or text detail artifacts; spatial text rendering struggles without T5 text encoder
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
