Attention Is All You Need
The paper that defined the next decade.
A universal network design that handles text, images, audio and code equally well — and keeps getting better when you simply make it bigger.
Attention gave us an architecture that scales almost without limit. The same block works for text, code, audio and video.
The paper that defined the next decade.
Language pre-training becomes the industry default.
Coherent text generation stops being a toy.
The same architecture works on images — the transformer turns out to be universal.
Sparse models: more parameters without a proportional compute increase.
Attention costs grow as the square of the input length. Linear and hybrid schemes are the main engineering work here.
The design has held for nine years. Every major lab is looking for a replacement, but no alternative has yet won at scale.