Skip to main content

> ML_ALGORITHM // MASKED-AUTOENCODER-MAE_v1.0

Masked Autoencoders (MAE)

Scalable self-supervised vision model that masks a high proportion (75%) of input image patches and trains an asymmetric autoencoder to reconstruct pixel values.

Masked Image Modelingself-supervisedblack-boxlarge (>100k)
Back to All Algorithms
Computational Complexity
Training Complexity:O(epochs * (0.25 * ViT_encoder + ViT_lightweight_decoder))
Inference Complexity:O(ViT_encoder)
Hardware Profile
CPU Friendly:No
Requires GPU:Yes
Memory Footprint:high
Interpretability & Data
Interpretability Tier:black-box
Training Data Needs:large (>100k)

Interpretability Assessment

Reconstruction outputs can be inspected directly to verify semantic visual understanding.

Suitable Tasks & Supported Modalities

Suitable Tasks:
feature extractionimage classificationobject detection
Supported Modalities:
imagevideo

Implementing Libraries

PyTorchLinux Foundation / PyTorch Foundation · v2.4.1
View Spec
torchvision
TransformersHugging Face · v4.44.2
View Spec

Foundational Literature

Masked Autoencoders Are Scalable Vision Learners (MAE)Kaiming He, Xinlei Chen (2022) · IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Common Pitfalls & Warnings
  • Using low masking ratios (<50%) allows trivial interpolation from neighboring pixels rather than semantic feature learning