> ML_LITERATURE // RAMESH-2022-HIERARCHICAL-TEXT-CONDITIONAL-IMAGE-GENERATION-DALLE2_v1.0
Hierarchical Text-Conditional Image Generation with CLIP Latents (DALL-E 2 / unCLIP)
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, Mark Chen · arXiv preprint (2022)
seminal-architecture2022foundationalthirdPartyReproduced
Principal Contribution
Inverted the CLIP image encoder using a diffusion prior mapping text captions to CLIP image embeddings followed by a diffusion decoder generating high-resolution pixels.
Operational Relevance
Serves as qualified reference for deploying task-image-generation in production.
Assumptions
- Spatial feature coherence and data manifold structure adhere to continuous representation hypotheses
Limitations
- Computational complexity scales with spatial resolution and parameter capacity
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
