# AI Hallucination An AI hallucination is when a [[Large Language Models (LLMs)|Large Language Model]] generates content that is factually incorrect, fabricated, or nonsensical while presenting it with full confidence. The model doesn't "know" it's wrong; it's producing the most statistically likely continuation of its input, and sometimes that continuation is fiction dressed as fact. Hallucinations are not bugs in the traditional sense. They're a structural property of how LLMs work: the model predicts tokens based on patterns, not based on verified knowledge. It can confidently cite papers that don't exist, invent function signatures, or describe historical events that never happened. Common hallucination types: - **Factual fabrication**: invented facts, dates, names, citations - **Code hallucination**: plausible-looking code that references nonexistent APIs, functions, or parameters - **Confident confabulation**: filling knowledge gaps with plausible-sounding but false information - **Source hallucination**: inventing URLs, paper titles, or quotes Hallucination risk increases with [[Context Bloat]] (noisy context confuses the model), poor [[Context Engineering]] (insufficient grounding information), and long generation without verification. [[Retrieval-Augmented Generation (RAG)]] reduces hallucinations by grounding output in actual documents. [[Agentic TDD]] catches code hallucinations through execution and testing. [[AI Guardrails]] can flag likely hallucinations before they reach users. The practical stance: never assume LLM output is correct until verified. This is why [[Agentic Engineering]] emphasizes code execution, testing, and human review over blind trust. ## Why models guess instead of saying "I don't know" In September 2025, four OpenAI researchers (Adam Tauman Kalai, Ofir Nachum, Santosh Vempala and Edwin Zhang) published *Why Language Models Hallucinate*. The paper takes a lot of the mystery out of the topic, and the argument has two parts. **Hallucinations start in pretraining, as ordinary classification errors.** In their words, hallucinations "originate simply as errors in binary classification". Picture every candidate statement as a yes/no question: is this valid or not? When a model can't reliably tell the true statements from the plausible false ones, the statistical pressure of pretraining makes it produce some of those false ones. No exotic bug is needed. It's the same kind of error any classifier makes on hard examples. **They survive post-training because the evals reward guessing.** Most benchmarks grade answers as right or wrong. "I don't know" scores zero, exactly like a wrong answer, so a guess always has the better expected score. The paper puts it bluntly: "Under binary grading, abstaining is strictly sub-optimal." Models end up tuned to be good test-takers, like a student who fills in every multiple-choice bubble because blanks earn nothing. Their fix is to change how the mainstream benchmarks are scored (instead of adding yet another hallucination benchmark). Each question states an explicit confidence target: "Answer only if you are >t confident, since mistakes are penalized t/(1−t) points, while correct answers receive 1 point, and an answer of 'I don't know' receives 0 points." With t = 0.75, a wrong answer costs 3 points. Do the math and answering only pays off when your chance of being right is really above 75%. The best strategy becomes reporting what you actually believe, which is the logic behind [[Proper Scoring Rules|proper scoring rules]]. I find this framing really useful. It turns hallucination into an incentive problem at least as much as a knowledge problem. A model whose confidence matches its accuracy ([[AI Model Calibration]]) can decline to answer when it should, and then your code decides what happens next, e.g. [[Confidence-Gated Routing|routing on confidence]] to a human or a bigger model below a threshold. [[Jev]] is built around that idea: it returns probabilities over typed answers instead of text, and TypeSafe trains it to express its uncertainty. It can still be confidently wrong, though, so calibration is something I'd verify on my own data before trusting any threshold. ## References - [Why Language Models Hallucinate (Kalai, Nachum, Vempala, Zhang; arXiv:2509.04664)](https://arxiv.org/abs/2509.04664) ## Related - [[Large Language Models (LLMs)]] - [[AI Limitations]] - [[AI Sycophancy]] - [[AI Safety]] - [[AI Guardrails]] - [[Context Engineering]] - [[Context Bloat]] - [[Retrieval-Augmented Generation (RAG)]] - [[Agentic TDD]] - [[Agentic Engineering]] - [[Prompt Engineering]] - [[Slopsquatting]] - [[Obsidian Starter Kit - Tutorial - Managing AI sessions]] - Operational guidance to reduce hallucination risk in OSK sessions - [[AI Model Calibration]] - [[Proper Scoring Rules]] - [[Confidence-Gated Routing]] - [[Jev]]