2021
CLIP
A shared space for text and images — the foundation all image generation grew on.
You can now show the same model a picture, play it a recording or hand it a video — and it understands them together rather than separately.
A single model takes text, image, audio and video and answers in real time. The computer got senses instead of a text box.
A shared space for text and images — the foundation all image generation grew on.
Speech recognition reaches the level where audio becomes an ordinary input.
Vision as a standard part of a frontier model.
Conversation latency drops to human level.
An hour of footage fits, but following the plot and causes inside it does not work yet.
Modalities with neither cheap sensors nor large datasets so far.