> ML_LITERATURE // RAMESH-2021-ZERO-SHOT-TEXT-TO-IMAGE-GENERATION-DALLE_v1.0
Zero-Shot Text-to-Image Generation (DALL-E)
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, Ilya Sutskever · International Conference on Machine Learning (ICML) (2021)
seminal-architecture2021foundationalthirdPartyReproduced
Principal Contribution
Trained a discrete VAE (dVAE) to compress images into 32x32 tokens and modeled concatenated text-image token sequences autoregressively using a 12-billion parameter Transformer.
Operational Relevance
Serves as qualified reference for deploying task-image-generation in production.
Assumptions
- Spatial feature coherence and data manifold structure adhere to continuous representation hypotheses
Limitations
- Computational complexity scales with spatial resolution and parameter capacity
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
