> ML_LITERATURE // RADFORD-2021-LEARNING-TRANSFERABLE-VISUAL-MODELS-NATURAL-LANGUAGE-SUPERVISION_v1.0
Learning Transferable Visual Models From Natural Language Supervision (CLIP)
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, Ilya Sutskever · International Conference on Machine Learning (ICML) (2021)
seminal-architecture2021foundationalthirdPartyReproduced
Principal Contribution
Pretrained dual encoders on 400M internet image-text pairs with symmetric InfoNCE contrastive loss, unlocking robust zero-shot classification.
Operational Relevance
The universal vision-language bridge powering Stable Diffusion, Midjourney, vector multimodal search, and zero-shot image classification.
Assumptions
- Natural language captions provide rich, scalable semantic supervision that maps visual concepts into a shared embedding space
Limitations
- Weak compositional binding and spatial counting; vulnerable to typographic adversarial attacks where text labels override visual contents
Connected Algorithms, Architectures & Tools
Related Algorithms:
Related Architectures:
Implementing Libraries:
