Milestones on the line · THINKING

Reasoning

Models used to blurt out an answer. Now they get to think: break the task into steps, check themselves and spend more time on hard problems — like a person reaching for scratch paper.

Models learned to spend compute at answer time: think longer, solve harder. A second scaling axis appeared alongside training.

StatusPASSED
TypeMilestones on the line
Marker2024
What exists today7

Researched

2022

Chain of thought

Asking the model to think step by step raised accuracy — the first hint that thinking at answer time pays.

verified
2024

o1

Chain of thought as a product, not a prompt trick.

verified
2025

The road to olympiad gold

AlphaProof took IMO 2024 silver, Aristotle reached 2025 gold. Every system solving problems formally worked through Lean.

verifiedNature
2026-07

AxiomProver: 42 out of 42

A multi-agent ensemble solved all six problems of IMO 2026 with formal Lean 4 proofs.

NEWreportedGitHub

In progress

сейчас

Reasoning plus tools

Thinking and calling a calculator, a search or code merged into one loop instead of two separate modes.

NEW

Planned

в планах

Knowing when to think

Long reasoning is expensive. The model must decide for itself where scratch paper is needed and where a reflex is enough.

NEW

Distant horizons

впереди

Reasoning that lasts months

Problems where thinking and verification take weeks of continuous work rather than minutes.

NEW
Branches
Milestones on the line
To the line

Track the line as it moves

Once a week: which branches advanced, what unlocked, and what turned out to be overstated.