> ML_ALGORITHM // CONTRASTIVE-LANGUAGE-IMAGE-PRETRAINING-CLIP_v1.0
Contrastive Language-Image Pre-training (CLIP)
Dual-encoder foundation model trained on 400M image-text pairs via contrastive loss, mapping text and visual concepts into a shared embedding space for zero-shot tasks.
Multimodal Contrastive Learningself-supervisedmoderate-posthocmassive (>10M)
Back to All AlgorithmsComputational Complexity
Training Complexity:O(epochs * batch_size^2 * (text_enc + vision_enc))
Inference Complexity:O(text_enc + vision_enc)
Hardware Profile
CPU Friendly:No
Requires GPU:Yes
Memory Footprint:high
Interpretability & Data
Interpretability Tier:moderate-posthoc
Training Data Needs:massive (>10M)
Interpretability Assessment
Enables zero-shot classification via natural language prompt cosine similarity without specialized training heads.
Suitable Tasks & Supported Modalities
Suitable Tasks:
zero shot classificationimage text retrievalfeature extraction
Supported Modalities:
imagetextmultimodal
Implementing Libraries
Foundational Literature
Learning Transferable Visual Models From Natural Language Supervision (CLIP)Alec Radford, Jong Wook Kim (2021) · International Conference on Machine Learning (ICML)
Common Pitfalls & Warnings
- Poor compositional reasoning (struggles to distinguish "a red cube on a blue sphere" vs "a blue cube on a red sphere")
- Weak counting and spatial attribute understanding in complex scenes
