Claim record · edition 1.0.0

si-005Established result

Fine-tuning with human feedback substantially changes how closely a model follows instructions relative to its pre-trained base.

The InstructGPT work established this as a distinct post-training stage rather than a property of pre-training.

Limits of this claim

Instruction-following is not correctness, safety, or capability. Improvements on preference judgements do not transfer automatically to task accuracy.

Supporting source records · 1

primary paperContent verified

Training language models to follow instructions with human feedback

Ouyang et al. · 2022

The post-training method that separated instruction-following from raw pre-training.

Source record →

Related concepts