Skip to main content

> ML_LITERATURE // RAMESH-2022-HIERARCHICAL-TEXT-CONDITIONAL-IMAGE-GENERATION-DALLE2_v1.0

Hierarchical Text-Conditional Image Generation with CLIP Latents (DALL-E 2 / unCLIP)

Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, Mark Chen · arXiv preprint (2022)

seminal-architecture2022foundationalthirdPartyReproduced

Principal Contribution

Inverted the CLIP image encoder using a diffusion prior mapping text captions to CLIP image embeddings followed by a diffusion decoder generating high-resolution pixels.

Operational Relevance

Serves as qualified reference for deploying task-image-generation in production.

Assumptions

  • Spatial feature coherence and data manifold structure adhere to continuous representation hypotheses

Limitations

  • Computational complexity scales with spatial resolution and parameter capacity

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: