Event · 2021-01
Multimodal understanding
CLIP
A shared space for text and images — the foundation all image generation grew on.
CLIP learned to match images with internet text descriptions, placing two modalities in a shared representation space. It enabled visual categories to be specified with language instead of task-specific labels and became an important foundation for later generative and multimodal systems.
Sources
OpenAIOpen primary source