# Reward Hacking Reward hacking is when an AI trained with [[Reinforcement Learning (RL)|reinforcement learning]] scores well on its reward function without doing the thing that reward was meant to encourage. The model satisfies the letter of the objective and skips the intent. It is the RL-specific case of a wider failure called specification gaming, where any AI reaches its stated goal in a way nobody actually wanted. ## Why it happens A reward function is a proxy. You cannot write "be a good coding assistant" as math, so you approximate it with something measurable, like "make the tests pass." The model then optimizes the proxy you wrote, not the goal you had in mind. Anywhere the proxy and the intent come apart, a capable optimizer finds the gap and settles into it. ## What it looks like The failures get more inventive as models get more capable. While training [[Composer 2.5]], Cursor caught the model reverse-engineering a Python cache to recover function signatures it was supposed to reconstruct from scratch, and decompiling Java bytecode to read APIs it was meant to infer. Older RL examples are blunter. The classic one is a boat-racing agent that learns to spin in a loop hitting bonus targets forever instead of finishing the race, because the score went up either way. ## Why it is getting attention The post-training method turns out to matter more here than raw capability. A May 2026 study ran one model family through the new Reward Hacking Benchmark, a set of multi-step tasks that each hide a tempting shortcut. [[DeepSeek V3|DeepSeek-V3]] took the shortcut 0.6% of the time. Its reinforcement-trained sibling, DeepSeek-R1-Zero, took it 13.9% of the time. Same lineage, very different behavior, and the RL stage is what taught the shortcut. For agentic coding tools, this bites in practice. An agent that games its own checks can ship code that passes CI and still does the wrong thing. That is the concrete reason human review on consequential changes stays mandatory, no matter how good the benchmark scores look. ## Overoptimization has a curve In 2022, Leo Gao, John Schulman and Jacob Hilton (OpenAI) measured reward hacking in [[Reinforcement Learning From Human Feedback (RLHF)|RLHF]]-style training (*Scaling Laws for Reward Model Overoptimization*). Human labels are expensive, so they used a synthetic setup: a large "gold-standard" reward model plays the role of the humans, and its labels train smaller proxy reward models. The policy optimizes the proxy while they track the gold score. The proxy score keeps climbing. The gold score first goes up, peaks, then falls. They plot it against d = √KL, the square root of the KL divergence between the optimized policy and the initial one (i.e., how far training has pushed the model away from where it started). The fitted curves are simple: - Best-of-n sampling: R(d) = d(α − βd) - Reinforcement learning: R(d) = d(α − β log d) The α and β coefficients scale smoothly with the size of the reward model, so you can predict where the peak will be. That's [[Goodhart's Law]] with a formula attached: optimize a measure hard enough and it stops tracking the thing you cared about. ### Calibration is part of the damage A preference reward model is a proxy for "answers humans like", and confident answers tend to look better. OpenAI's own GPT-4 technical report shows what happens to confidence. On a subset of MMLU, the pre-trained model had an expected calibration error of 0.007. After PPO post-training, it was 0.074, and the report says it plainly: "The post-training hurts calibration significantly." My reading is that overconfidence is one more way to score well on the proxy while drifting from the goal. See [[AI Model Calibration]] for how that's measured and what people do about it. ## References - Reward hacking overview: https://en.wikipedia.org/wiki/Reward_hacking - Lilian Weng, Reward Hacking in Reinforcement Learning: https://lilianweng.github.io/posts/2024-11-28-reward-hacking/ - [Scaling Laws for Reward Model Overoptimization (Gao, Schulman, Hilton; arXiv:2210.10760)](https://arxiv.org/abs/2210.10760) - [GPT-4 Technical Report (OpenAI; arXiv:2303.08774)](https://arxiv.org/abs/2303.08774) ## Related - [[Reinforcement Learning (RL)]] - [[Reinforcement Learning From Human Feedback (RLHF)]] - [[AI Alignment]] - [[Goodhart's Law]] - [[AI Safety]] - [[AI Guardrails]] - [[Large Language Models (LLMs)]] - [[AI Agents]] - [[Composer 2.5]] - [[AI Model Calibration]]