> ML_ALGORITHM // MASKED-AUTOENCODER-MAE_v1.0
Masked Autoencoders (MAE)
Scalable self-supervised vision model that masks a high proportion (75%) of input image patches and trains an asymmetric autoencoder to reconstruct pixel values.
Masked Image Modelingself-supervisedblack-boxlarge (>100k)
Back to All AlgorithmsComputational Complexity
Training Complexity:O(epochs * (0.25 * ViT_encoder + ViT_lightweight_decoder))
Inference Complexity:O(ViT_encoder)
Hardware Profile
CPU Friendly:No
Requires GPU:Yes
Memory Footprint:high
Interpretability & Data
Interpretability Tier:black-box
Training Data Needs:large (>100k)
Interpretability Assessment
Reconstruction outputs can be inspected directly to verify semantic visual understanding.
Suitable Tasks & Supported Modalities
Suitable Tasks:
feature extractionimage classificationobject detection
Supported Modalities:
imagevideo
Implementing Libraries
Foundational Literature
Masked Autoencoders Are Scalable Vision Learners (MAE)Kaiming He, Xinlei Chen (2022) · IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Common Pitfalls & Warnings
- Using low masking ratios (<50%) allows trivial interpolation from neighboring pixels rather than semantic feature learning
