Skip to main content

> ML_LITERATURE // DOSOVITSKIY-2020-IMAGE-WORTH-16X16-WORDS-VIT_v1.0

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT)

Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby · International Conference on Learning Representations (ICLR) (2020)

seminal-architecture2020foundationalthirdPartyReproduced

Principal Contribution

Applied a standard Transformer encoder directly to sequences of flattened 16x16 image patches, breaking the necessity of inductive convolutional bias.

Operational Relevance

The standard vision backbone across multimodal LLMs (GPT-4V, Gemini, Claude 3.5 Sonnet, LLaVA), self-supervised models (DINOv2, MAE), and CLIP.

Assumptions

  • When pre-trained on massive datasets (JFT-300M / ImageNet-21k), general self-attention out-scales spatial convolutional inductive biases

Limitations

  • Severe performance degradation when trained from scratch on small datasets (ImageNet-1k) without heavy data augmentation or pre-training

Connected Algorithms, Architectures & Tools

Related Algorithms:
Related Architectures:
Implementing Libraries: