Skip to main content

> ML_ARCHITECTURE // VISION-TRANSFORMER-VIT_v1.0

Vision Transformer (ViT / DINOv2 / MAE)

Pure transformer architecture operating directly on flattened 16x16 image patches without convolutional inductive biases, unlocking massive scale for visual representation pretraining.

Vision Transformersimage
Back to All Architectures

Architecture Overview

Pure transformer architecture operating directly on flattened 16x16 image patches without convolutional inductive biases, unlocking massive scale for visual representation pretraining.

Implementing Libraries

PyTorchLinux Foundation / PyTorch Foundation · v2.4.1
View Spec
torchvisionPyTorch Foundation / Meta · v0.19.1
View Spec
timm (PyTorch Image Models)Ross Wightman / Hugging Face · v1.0.9
View Spec

Seminal Papers

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT)Alexey Dosovitskiy, Lucas Beyer (2020) · International Conference on Learning Representations (ICLR)
Masked Autoencoders Are Scalable Vision Learners (MAE)Kaiming He, Xinlei Chen (2022) · IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
DINOv2: Learning Robust Visual Features without SupervisionMaxime Oquab, Timothée Darcet (2023) · arXiv preprint
Architectural Limitations & Constraints
  • Requires compatible deep learning framework and hardware acceleration for efficient execution.