RLHF (Reinforcement Learning from Human Feedback) is the step that turns a base model which merely continues text into a chat model that follows instructions and matches human preferences. OpenAI's public InstructGPT work around 2022 brought the pipeline wide attention, and it has since become a standard part of aligning mainstream large models.
1. Why pretraining alone is not enough
A pretrained model is trained to predict the next token, so it learns what internet text looks like rather than what a user actually wants. Talking to a raw base model often produces ignored instructions, off-topic answers, harmful output or an unsuitable tone. Supervised fine-tuning (SFT) fixes part of this, but SFT mainly teaches the model to imitate reference answers; it cannot easily express "both answers look fine, but A is clearly better than B". RLHF fills exactly that gap.
2. The classic three steps
(1) Supervised fine-tuning (SFT): fine-tune the base model on high-quality human-written prompt-answer pairs to get a model that roughly follows instructions.
(2) Reward model training: generate several responses to the same prompt, have human annotators rank them by quality, then train a scoring model to predict which response humans prefer. In practice pairwise comparisons are more stable than asking annotators for absolute scores.
(3) Reinforcement learning (PPO): treat the reward model as a judge and update the language model with an algorithm such as PPO (Proximal Policy Optimization) so its responses earn higher rewards. A KL penalty is usually added to keep the new model from drifting too far from the SFT model, preventing strange text produced purely to game the score.
3. Simpler alternatives
DPO (Direct Preference Optimization) skips the separate reward model and the reinforcement learning loop, optimizing the language model directly on preference data. It folds "rankings to reward to policy" into one simpler objective, sharply reducing engineering complexity, which is why it is widely used in the open-source community.
RLAIF / Constitutional AI replaces part of the human annotation with an AI that scores or revises answers against a written set of principles, cutting labeling cost but shifting the question of "who judges" to how those principles are designed.
4. Limits and criticism
Preference data comes from annotators, so their biases end up encoded in the model. RLHF optimizes what humans like, which is not the same as what is correct, and can encourage sycophantic answers. Reward models can be over-optimized (reward hacking), letting the model game the metric instead of genuinely improving. The whole pipeline is long, expensive and sensitive to hyperparameters, and training stability remains an engineering challenge.
5. How to think about its role
Pretraining is like reading widely, SFT is learning to answer on command, and RLHF is learning to tell which answer is better. Together they form the chat models we use today. RLHF generally does not change how much a model knows; it changes how the model expresses and selects its answers.
Reference: compiled from publicly released industry information (public papers and documentation on InstructGPT, DPO, Constitutional AI and related work).