Source record · content verified

Training language models to follow instructions with human feedback

primary paper2022

Why this source is here

The post-training method that separated instruction-following from raw pre-training.

Verification
Content-verified on 2026-07-27: the canonical source and its title were resolved during the Atlas review. This is not an endorsement of the source’s argument.
Authors
Ouyang et al.
Identifier
arXiv:2203.02155
Open source destination ↗

Claims citing this source · 1

Concepts citing this source · 1