GPT-3, 175B
In-context learning: a couple of examples is enough.
It turned out you can predict a model's quality in advance from how much compute and data go into it. Progress stopped being luck and became a budgeting question.
Quality grows predictably with compute, data and parameters. That turned research into industrial planning and started the data-centre race.
In-context learning: a couple of examples is enough.
The curves clusters have been planned against ever since.
Data beats size — every training budget gets rebuilt.
Some skills appear in a jump past a scale threshold rather than growing smoothly.
High-quality human text on the internet is close to exhausted. Next come synthetic data, video and companies' own corpora.
The next growth axis is not model size but the volume of training on the model's own attempts and mistakes.