To the line
Capabilities on the main line · POST-TRAINING

Following human intent

In one viewIn 2022, models became much better at understanding requests and responding through dialogue. They were tuned on instruction examples and human evaluations.

ChatGPT turned a language model from a research tool into a mainstream interface. Confident errors and the tendency to agree with users remained major limitations.

StatusPASSED
TypeCapabilities on the main line
Marker2022
Events in dossier7
Development chronology

Researched

2022-03

InstructGPT

About this eventInstruction tuning — the model starts following orders.

2022-11

ChatGPT

About this eventA dialogue interface made instruction following a mainstream way to use a language model.

2022-12

Constitutional AI

About this eventTraining against a written set of principles instead of hand-labelling every answer.

2023-03

GPT-4

About this eventHuman-level results on professional exams move AI out of the toy category.

2024-03

Claude 3

About this eventAnthropic's family established sustained frontier competition across reasoning, coding and image analysis.

In progress

now

Sycophancy as a defect

About this eventA model trained to please tends to agree. Teaching it to push back is an open alignment problem.

Planned

ahead

Learning from outcomes

About this eventReplacing "did the answer feel good" with "did the result actually work".

Sources and research

Primary material behind this dossier: papers, lab publications and official reports.

Directions
Capabilities on the main line