> ML_ARCHITECTURE // CLIP-CONTRASTIVE-DUAL-ENCODER_v1.0
Contrastive Dual-Encoder (CLIP / OpenCLIP / SigLIP)
Dual-encoder architecture projecting images and text into a shared normalized embedding space trained via symmetric cross-entropy contrastive loss, the backbone of visual search and text-to-image conditioning.
Multimodal Foundation Modelsimagetext
Back to All ArchitecturesArchitecture Overview
Dual-encoder architecture projecting images and text into a shared normalized embedding space trained via symmetric cross-entropy contrastive loss, the backbone of visual search and text-to-image conditioning.
Implementing Libraries
Seminal Papers
Learning Transferable Visual Models From Natural Language Supervision (CLIP)Alec Radford, Jong Wook Kim (2021) · International Conference on Machine Learning (ICML)
Architectural Limitations & Constraints
- Requires compatible deep learning framework and hardware acceleration for efficient execution.
