Skip to main content

> ML_LITERATURE // ROMBACH-2022-HIGH-RESOLUTION-IMAGE-SYNTHESIS-LATENT-DIFFUSION-MODELS_v1.0

High-Resolution Image Synthesis with Latent Diffusion Models (Stable Diffusion)

Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Björn Ommer · IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2022)

seminal-architecture2022industry-standardthirdPartyReproduced

Principal Contribution

Operated diffusion in the lower-dimensional latent space of a pretrained autoencoder with cross-attention conditioning, launching Stable Diffusion.

Operational Relevance

The foundation for the global open-source image generation industry (Stable Diffusion 1.5, SDXL, ControlNet, LoRA).

Assumptions

  • Perceptual compression separates high-frequency pixel imperceptible details from semantic conceptual composition in images

Limitations

  • Autoencoder decoders can introduce minor facial or text detail artifacts; spatial text rendering struggles without T5 text encoder

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: