Back to section
Event · 2022-03

Following human intent

InstructGPT

Instruction tuning — the model starts following orders.

InstructGPT combined demonstrations of instructions with human preference feedback and showed that a smaller aligned model could be more useful than a larger base model. It shifted model development beyond pretraining toward post-training for response format, refusals, style and user intent.

Sources
arXivOpen primary source