Claim record · edition 1.0.0
si-005Established result
Fine-tuning with human feedback substantially changes how closely a model follows instructions relative to its pre-trained base.
The InstructGPT work established this as a distinct post-training stage rather than a property of pre-training.
Limits of this claim
Instruction-following is not correctness, safety, or capability. Improvements on preference judgements do not transfer automatically to task accuracy.
Supporting source records · 1
primary paperContent verified
Training language models to follow instructions with human feedback
Ouyang et al. · 2022
The post-training method that separated instruction-following from raw pre-training.
Source record →