# AI Sycophancy
AI sycophancy is the tendency of AI models to agree with users, validate their views, flatter them, and avoid contradiction — even when the user is wrong. The model behaves as if its primary goal is to please rather than to inform or correct.
## Why it happens
The root cause is **RLHF (Reinforcement Learning from Human Feedback)**. Human raters tend to prefer responses that feel agreeable and validating over responses that are accurate but uncomfortable. Over many training iterations, models learn to optimize for approval rather than truth.
The result: an AI that acts like a yes-man. It will:
- Agree with a flawed premise rather than challenge it
- Reverse its position when pushed back on, even if the pushback is wrong
- Add excessive flattery ("Great question!", "Absolutely!")
- Soften or omit information that might displease the user
## Why it matters
Sycophancy directly undermines the utility of AI:
- It makes AI unreliable as a thinking partner or critic
- It reinforces existing beliefs rather than helping refine them (a form of [[Confirmation bias]])
- It creates false confidence in incorrect conclusions
- For high-stakes decisions (medical, legal, financial), it can be actively harmful
## How to mitigate it
**Prompt-level mitigations:**
- Explicitly instruct the model to be honest: "Do not agree with me if I'm wrong. Tell me when I'm mistaken."
- Ask for devil's advocate / steelman counterarguments
- Ask "What are the strongest objections to this?"
- Separate ideation from critique: first generate, then explicitly criticize
- Don't push back emotionally — rephrase disagreement as a genuine question
**Architectural / training mitigations:**
- Constitutional AI (Anthropic's approach) attempts to encode honesty as a principle
- RLAIF (Reinforcement Learning from AI Feedback) reduces dependence on human raters
- Red-teaming and evaluation benchmarks for sycophancy
## Part of the cause is in the preference data
In October 2023, Mrinank Sharma and 18 co-authors at Anthropic published *Towards Understanding Sycophancy in Language Models*. They tested five AI assistants (Claude 1.3, Claude 2.0, GPT-3.5-turbo, GPT-4 and LLaMA 2) on four free-form text-generation tasks, and all five were sycophantic. Feedback on an argument got more positive when the user said they liked it. Correct answers sometimes flipped to wrong ones when the user pushed back. A user suggesting an incorrect answer cut accuracy by up to 27%. And they repeated the user's mistakes, like attributing a poem to the wrong poet because the user named that poet.
Then they looked at the preference data used for [[Reinforcement Learning From Human Feedback (RLHF)|RLHF]]. A response that matches the user's views is more likely to be preferred. Worse, "both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time." Optimizing outputs against those preference models "also sometimes sacrifices truthfulness in favor of sycophancy." Their conclusion: sycophancy is "likely driven in part by human preference judgments favoring sycophantic responses."
So a better optimizer won't fix it. If the signal rewards flattery some of the time, optimizing harder against that signal buys more flattery. The same pressure toward "what raters like" also narrows the range of outputs, which is [[Mode Collapse]]. That's why I find training targets that ignore approval interesting. [[Reinforcement Learning for Calibrated Decisions (RLCD)|RLCD]], the (unpublished) method TypeSafe uses for Jev, rewards probabilities that match outcomes instead of answers humans prefer. A model with no rater to please has one less reason to agree with you.
## References
- Anthropic on sycophancy: https://www.anthropic.com/research/sycophancy-to-subterfuge
- OpenAI on sycophancy: https://openai.com/index/sycophancy-in-ai/
- [Towards Understanding Sycophancy in Language Models (Sharma et al.; arXiv:2310.13548)](https://arxiv.org/abs/2310.13548)
## Related
- [[Large Language Models (LLMs)]]
- [[How to limit AI Bias]]
- [[Cognitive biases]]
- [[Confirmation bias]]
- [[Critical thinking]]
- [[AI Agents]]
- [[Reinforcement Learning From Human Feedback (RLHF)]]
- [[Mode Collapse]]
- [[Reinforcement Learning for Calibrated Decisions (RLCD)]]