Skip to main content

> ML_LITERATURE // RAMESH-2021-ZERO-SHOT-TEXT-TO-IMAGE-GENERATION-DALLE_v1.0

Zero-Shot Text-to-Image Generation (DALL-E)

Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, Ilya Sutskever · International Conference on Machine Learning (ICML) (2021)

seminal-architecture2021foundationalthirdPartyReproduced

Principal Contribution

Trained a discrete VAE (dVAE) to compress images into 32x32 tokens and modeled concatenated text-image token sequences autoregressively using a 12-billion parameter Transformer.

Operational Relevance

Serves as qualified reference for deploying task-image-generation in production.

Assumptions

  • Spatial feature coherence and data manifold structure adhere to continuous representation hypotheses

Limitations

  • Computational complexity scales with spatial resolution and parameter capacity

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: