> ML_LITERATURE // DOSOVITSKIY-2020-IMAGE-WORTH-16X16-WORDS-VIT_v1.0
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT)
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby · International Conference on Learning Representations (ICLR) (2020)
seminal-architecture2020foundationalthirdPartyReproduced
Principal Contribution
Applied a standard Transformer encoder directly to sequences of flattened 16x16 image patches, breaking the necessity of inductive convolutional bias.
Operational Relevance
The standard vision backbone across multimodal LLMs (GPT-4V, Gemini, Claude 3.5 Sonnet, LLaVA), self-supervised models (DINOv2, MAE), and CLIP.
Assumptions
- When pre-trained on massive datasets (JFT-300M / ImageNet-21k), general self-attention out-scales spatial convolutional inductive biases
Limitations
- Severe performance degradation when trained from scratch on small datasets (ImageNet-1k) without heavy data augmentation or pre-training
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
