Event · 2022-03
Following human intent
InstructGPT
Instruction tuning — the model starts following orders.
InstructGPT combined demonstrations of instructions with human preference feedback and showed that a smaller aligned model could be more useful than a larger base model. It shifted model development beyond pretraining toward post-training for response format, refusals, style and user intent.
Sources
arXivOpen primary source