# Reinforcement Learning From Human Feedback (RLHF) RLHF is the training technique that turned raw [[Large Language Models (LLMs)]] from next-token predictors into useful assistants. The process has three stages: 1. **Supervised fine-tuning**: the base model is trained on curated examples of desired behavior (high-quality conversations, helpful answers) 2. **Reward model training**: human evaluators rank multiple model outputs for the same prompt. A separate model learns to predict which outputs humans prefer 3. **Reinforcement learning**: the LLM is optimized to produce outputs that score highly with the reward model, using algorithms like PPO (Proximal Policy Optimization) The result is a model that's not just predicting likely text, but producing text that humans judge as helpful, harmless, and honest. This is what makes [[Claude]], [[ChatGPT]], and other assistants feel cooperative rather than chaotic. RLHF is the primary mechanism behind [[AI Alignment]] in practice. It's also imperfect: the reward model can have blind spots, leading to [[AI Sycophancy]] (the model learns that agreeable answers get higher human ratings) and mode collapse (the model converges on safe, generic responses rather than taking useful risks). Alternatives and extensions include RLAIF (AI feedback instead of human), Constitutional AI ([[Anthropic]]'s approach where the model self-critiques against a set of principles), and DPO (Direct Preference Optimization, which skips the reward model). ## Where it came from RLHF is older than chatbots. The founding paper is *Deep reinforcement learning from human preferences* (Christiano, Leike, Brown, Martic, Legg and [[Dario Amodei]], arXiv:1706.03741, 2017). They taught agents Atari games and simulated robot tasks by showing humans pairs of short video clips and asking which one looked better. A reward model learned from those comparisons, and the agent optimized that reward. It needed feedback on less than 1% of the agent's interactions, and about an hour of human time for some novel behaviors. Then it moved to language: Ziegler et al. (arXiv:1909.08593, 2019) for stylistic continuation, Stiennon et al. (*Learning to summarize from human feedback*, arXiv:2009.01325, 2020) for summaries. [[InstructGPT]] (Ouyang et al., arXiv:2203.02155, March 2022) scaled it to general instruction following, and [[ChatGPT]] (November 2022) put it in front of everyone. [[Diogo Almeida]], now CEO of [[TypeSafe AI]], is one of InstructGPT's primary authors; TypeSafe markets him as a co-inventor of RLHF, which is shorthand for that role rather than for the 2017 idea. The InstructGPT result that sold everyone: outputs from a 1.3B-parameter RLHF model were preferred to those of the 175B GPT-3, "despite having 100x fewer parameters". ## What RLHF does to calibration This is the part I didn't appreciate at first. Pretrained models are well calibrated: when they give an answer 80% probability, they're right about 80% of the time (see [[AI Model Calibration]]). RLHF breaks that: - The [[GPT4|GPT-4]] technical report (arXiv:2303.08774) measured an expected calibration error of 0.007 for the pretrained model on an MMLU subset, and 0.074 after PPO. OpenAI's words: "post-training hurts calibration significantly" - Kadavath et al. (Anthropic, arXiv:2207.05221) explain the mechanism: RL "tends to collapse language model predictions towards behaviors that receive the most reward". A single temperature of 2.5 largely restored calibration in their tests - Leng et al. (arXiv:2410.09724) found that reward models favor high-confidence answers, which pushes the policy toward overconfidence The reason is structural. Pretraining uses a strictly proper objective (cross-entropy rewards honest probabilities, see [[Proper Scoring Rules]]). A preference reward model is a proxy for "what humans liked", and humans like confident, fluent answers. Optimize a proxy hard enough and you get [[Goodhart's Law]] and [[Reward Hacking]]: Gao, Schulman and Hilton (arXiv:2210.10760) showed the true reward first rises then falls as the policy over-optimizes the reward model. Diogo Almeida's summary: "Overpromising is a feature... by design." ## Mode dropping and sycophancy Two more side effects come from the same root: - **Mode dropping.** RLHF narrows the output distribution. Kirk et al. (arXiv:2310.06452, ICLR 2024) found it "significantly reduces output diversity compared to SFT", even though it generalizes better. TypeSafe's primer calls this *mode dropping*, a milder form of GAN [[Mode Collapse]]: the model favors one style "while reducing the probability of other possible outputs." For a classifier, that means rare classes get squashed - **[[AI Sycophancy]].** Sharma et al. (Anthropic, arXiv:2310.13548, ICLR 2024) found five state-of-the-art assistants consistently sycophantic across four free-form tasks, and traced it partly to preference data: answers that match the user's views are more likely to be preferred, by humans and by preference models alike ## RLHF vs RLVR vs RLCD A useful framing from Almeida (Latent Space, September 2026): RLHF names the *task*, not the algorithm. "It's not about the PPO. That part doesn't matter." DPO also does RLHF, "despite not using the algorithm." Seen that way, [[AI Post-Training|post-training]] splits by what gets rewarded: | | RLHF | [[Reinforcement Learning with Verifiable Rewards (RLVR)\|RLVR]] | [[Reinforcement Learning for Calibrated Decisions (RLCD)\|RLCD]] | |---|---|---|---| | Rewards | What human raters prefer | Answers a program verifies as correct | Probabilities that match outcomes | | Produces | Chat assistants | [[AI Reasoning Models\|Reasoning models]] (o1, DeepSeek-R1) | [[System One Models]] ([[Jev]]) | | Typical failure | Sycophancy, overconfidence, mode dropping | Jaggedness, slow and costly inference, guessing | Unpublished recipe, calibration that may not transfer to your data | RLHF remains the right tool for conversational models; even TypeSafe's primer says so. The argument is that it's the wrong objective for unattended automation, where "human preference and machine trustworthiness are different optimization targets." ## References - [Deep reinforcement learning from human preferences (Christiano et al., arXiv:1706.03741)](https://arxiv.org/abs/1706.03741) - [Fine-Tuning Language Models from Human Preferences (arXiv:1909.08593)](https://arxiv.org/abs/1909.08593) - [Learning to summarize from human feedback (arXiv:2009.01325)](https://arxiv.org/abs/2009.01325) - [Training language models to follow instructions with human feedback (InstructGPT, arXiv:2203.02155)](https://arxiv.org/abs/2203.02155) - [GPT-4 Technical Report (arXiv:2303.08774)](https://arxiv.org/abs/2303.08774) - [Language Models (Mostly) Know What They Know (arXiv:2207.05221)](https://arxiv.org/abs/2207.05221) - [Taming Overconfidence in LLMs: Reward Calibration in RLHF (arXiv:2410.09724)](https://arxiv.org/abs/2410.09724) - [Scaling Laws for Reward Model Overoptimization (arXiv:2210.10760)](https://arxiv.org/abs/2210.10760) - [Understanding the Effects of RLHF on LLM Generalisation and Diversity (arXiv:2310.06452)](https://arxiv.org/abs/2310.06452) - [Towards Understanding Sycophancy in Language Models (arXiv:2310.13548)](https://arxiv.org/abs/2310.13548) - [TypeSafe AI primer](https://docs.typesafe.ai/introduction/machine-learning-primer) - [Jev: System One models for Prod, not God, with Diogo Almeida (Latent Space)](https://www.latent.space/p/jev) ## Related - [[AI Alignment]] - [[AI Safety]] - [[AI Sycophancy]] - [[Large Language Models (LLMs)]] - [[Machine Learning (ML)]] - [[Deep Learning]] - [[Anthropic]] - [[Claude]] - [[InstructGPT]] - [[AI Post-Training]] - [[Reinforcement Learning with Verifiable Rewards (RLVR)]] - [[Reinforcement Learning for Calibrated Decisions (RLCD)]] - [[AI Model Calibration]] - [[Mode Collapse]] - [[Reward Hacking]] - [[Goodhart's Law]] - [[Diogo Almeida]] - [[Reinforcement Learning (RL)]]