Back to section
Event · 2021-01

Multimodal understanding

CLIP

A shared space for text and images — the foundation all image generation grew on.

CLIP learned to match images with internet text descriptions, placing two modalities in a shared representation space. It enabled visual categories to be specified with language instead of task-specific labels and became an important foundation for later generative and multimodal systems.

Sources
OpenAIOpen primary source