Source record · content verified
Training language models to follow instructions with human feedback
primary paper2022
Why this source is here
The post-training method that separated instruction-following from raw pre-training.
- Verification
- Content-verified on 2026-07-27: the canonical source and its title were resolved during the Atlas review. This is not an endorsement of the source’s argument.
- Authors
- Ouyang et al.
- Identifier
- arXiv:2203.02155