A multimodal model that maps images and text into one shared meaning space.
CLIP is like the name-tag helper on school picture day. It knows “golden retriever” matches the dog photo, not your math teacher.
It can match images with captions and find pictures from words. It also helps some text-to-image AIs read prompts.
Multimodal AI
CLIP is a classic way to match pictures with words.
Embedding
CLIP turns images and text into number lists you can compare.
Diffusion
Many text-to-image systems use CLIP to understand prompts better.
Pretraining
CLIP learns from huge sets of image-caption pairs first.