Skip to main content

> ML_ARCHITECTURE // CLIP-CONTRASTIVE-DUAL-ENCODER_v1.0

Contrastive Dual-Encoder (CLIP / OpenCLIP / SigLIP)

Dual-encoder architecture projecting images and text into a shared normalized embedding space trained via symmetric cross-entropy contrastive loss, the backbone of visual search and text-to-image conditioning.

Multimodal Foundation Modelsimagetext
Back to All Architectures

Architecture Overview

Dual-encoder architecture projecting images and text into a shared normalized embedding space trained via symmetric cross-entropy contrastive loss, the backbone of visual search and text-to-image conditioning.

Implementing Libraries

TransformersHugging Face · v4.44.2
View Spec
PyTorchLinux Foundation / PyTorch Foundation · v2.4.1
View Spec
torchvisionPyTorch Foundation / Meta · v0.19.1
View Spec

Seminal Papers

Learning Transferable Visual Models From Natural Language Supervision (CLIP)Alec Radford, Jong Wook Kim (2021) · International Conference on Machine Learning (ICML)
Architectural Limitations & Constraints
  • Requires compatible deep learning framework and hardware acceleration for efficient execution.