Why It Invents Things So Confidently
A hallucination isn't a malfunction — it's the same operation as a correct answer, minus the luck. Which is why confidence tells you nothing, and why you check by cost.
You ask for ten sources on a topic and get ten: authors, titles, years, journals. Eight are real. Two do not exist — and if anything, the two fakes look tidier than the rest.
This is the failure mode that burns people worst, because it doesn't announce itself. Nothing stutters, nothing hedges, no warning appears. And it's not a glitch that will be patched away next year: it comes from the same machinery that produces the good answers.
The same operation, minus the luck
Go back to the landscape. The model produces a plausible continuation — that's all it does. When the terrain under the question is deeply worn by material that happens to be true, the plausible continuation is also the correct one. When it isn't, the ball still rolls. It always rolls. There is no cliff marked "no data here."
So a correct answer and a fabricated one are the same act: continue plausibly. One landed on reality; one didn't. And nothing inside the model distinguishes them, because the model has no separate faculty for "checking whether I know this" — remember, recall and invention are one operation from the inside.
That has a consequence worth putting in bold: the confidence in an answer is a property of the writing style, not of the evidence. Fluent, structured and assured is how this system writes when it's right and how it writes when it's making things up. You cannot read reliability off the surface, and your instinct to try is trained on humans, where hesitation genuinely correlates with uncertainty.
The student who can't leave a blank
Here's the picture to carry.
Imagine a straight-A student sitting an exam with one unusual rule: leaving a cell blank is forbidden, and there's no penalty for a wrong answer. Blank scores zero; a wrong guess scores zero; a lucky guess scores full marks. What's the optimal strategy? Fill in every cell, always, with the most plausible-looking thing available. And do it in the same neat handwriting as the answers you actually know — because handwriting isn't graded.
The model has no blank cell. Only a filled one.
This isn't just a metaphor. In 2025 Adam Kalai and colleagues published an analysis with the blunt title Why Language Models Hallucinate, arguing exactly this: hallucinations persist not because nobody has fixed a bug, but because the way models are trained and scored rewards guessing over admitting uncertainty. Standard benchmarks grade answers right or wrong; "I don't know" earns the same zero as a wrong answer, so a model tuned to score well learns to always produce something. The exam made the student this way.
Which also tells you what helps: change the local incentive. If saying "I don't know" is explicitly allowed and asked for, you get more of it. That's a prompt you can write, and we'll get to it.
What it costs when nobody checks
In 2023 a New York lawyer preparing a filing in Mata v. Avianca asked ChatGPT for supporting case law and got it: case names, court decisions, quotations, citations. Six of the cases did not exist. When opposing counsel couldn't find them, he asked the model whether the cases were real; it assured him they were. The brief was filed. Judge P. Kevin Castel sanctioned the lawyers $5,000 that June, and the episode has been repeated by other professionals many times since.
The instructive part isn't the fabrication — it's the verification step. He did check. He checked by asking the same system that produced the answer, which is not a check at all. Which brings us to the two moves that feel like verification and aren't.
Two fake checks
"Are you sure?" By now you know why this fails: the model was trained on human preference, and pushback is the strongest signal there is that the human wants a different answer. You'll usually get a reversal — sometimes off a wrong answer, sometimes off a right one. The reversal carries almost no information about the truth. If you want a real second look, ask for reasoning, not for a verdict: "work through whether this citation exists, step by step, and say what you'd need to confirm it."
"Give me a source." The source is generated by the same process as the claim. Asking a model that just invented a fact to produce a reference for it very often yields an invented reference — plausible author, plausible journal, plausible year, sometimes a real DOI belonging to a different paper. A citation is only evidence when you have followed it. An unfollowed link is decoration.
Where the risk lives
Hallucination isn't spread evenly. It concentrates on thin terrain — anything specific, rare, recent or precise:
- exact numbers, dates, statistics, prices, dosages;
- verbatim quotations — the shape of a quote is easy, the exact words are not;
- citations, URLs, page numbers, file names, court cases, law and regulation specifics;
- facts about specific people, especially non-famous ones, and biographies generally;
- API details, function parameters, library versions, configuration flags — the format is well worn, the specifics are not;
- anything after the knowledge cutoff;
- narrow, niche or highly technical subfields, where the ground was barely trodden at all.
And one trigger that's less obvious but catches people constantly: a false premise in your own question. Ask "why did Einstein fail mathematics at school?" and a model that isn't careful will explain something that never happened, because your question already put that valley in front of the ball. When you're fishing for information, don't smuggle the assumption into the question. Ask "did X happen?" before "why did X happen?"
What actually reduces it
None of these are perfect. All of them beat hoping.
Put the material on the desk. A model answering from a document you supplied is doing something far closer to reading than to recall. Anchor it further by demanding traceability: "quote the exact sentence from the text that supports each claim, and if it isn't there, say so."
Turn on search for anything time-sensitive — and still open the links. Grounding in retrieved pages reduces invention substantially; it doesn't eliminate misreading, and a real link can still be attached to a claim it doesn't support.
Authorise ignorance out loud. "If you're not sure, say so. It's better to answer 'I don't know' than to guess." A striking amount of fabrication is politeness under pressure to produce.
Ask for the confidence split. "Separate this into: what you're confident about, what you're unsure about, and what should be verified externally." Models are considerably better at flagging the shaky parts when asked to sort them than at spontaneously hedging inside prose.
Verify outside the loop. Any check that stays inside the same conversation isn't a check. Search independently, open the source, run the code, ask a person.
Calibration, not distrust
The point of all this isn't to distrust the model — a colleague you double-check on everything saves you nothing. The point is to spend your checking where it pays.
Sort the work by the cost of an error. On cheap-error work — brainstorming, first drafts, explaining a concept you'll test anyway, code you're about to run, structuring your own messy thoughts — a wrong answer costs you thirty seconds and you'll notice it naturally. Don't verify; just work. On expensive-error work — a number that goes in a report, a citation in something you publish, anything legal, medical or financial, anything that triggers an irreversible action, anything a client will read — a wrong answer costs reputation or worse. Verify every single specific, every time.
Most people get this exactly backwards: they interrogate the model about a brainstorm and paste the citation straight into the document.
One last thing, and it links straight back to How to Learn. Fluent, well-organised text produces a strong feeling of understanding in the reader — the same illusion of competence you get from rereading a textbook with a highlighter. When the model explains something to you smoothly, you'll feel you've got it. The cure is the same as it was there: close the answer and reproduce it yourself. If you can't, you didn't get it — you got the feeling.
In practice
Never re-ask the same model as your verification step. Verify outside: source, search, execution, a human.
Follow every link before it leaves your hands. A citation you haven't opened is not a citation.
Ask "did this happen?" before "why did this happen?" Don't hand it a premise you want checked.
Say "I don't know is an acceptable answer" out loud whenever the answer matters.
Sort by cost of error, not by importance of topic. Cheap-error work: don't verify, iterate. Expensive-error work: verify every specific, always.
Check yourself
Close the article and answer in your own words:
- Why is a hallucination not a malfunction but the same operation as a correct answer?
- What does the exam-with-no-blanks metaphor explain, and what did Kalai and colleagues argue about how models are scored?
- Why are "are you sure?" and "give me a source" not verification?
- Name five zones where the risk of invention rises sharply. Why those?
- What's wrong with the question "why did Einstein fail mathematics?"
- Which four techniques genuinely reduce fabrication, and why does none of them eliminate it?
- How do you decide what to verify and what to leave alone?
In short
- A hallucination is the same act as a correct answer: produce a plausible continuation. One matched reality, one didn't; nothing inside the model tells them apart.
- Confidence is a writing style, not evidence. Fluency is identical whether the ground beneath the answer is deep or empty.
- Why Language Models Hallucinate (Kalai et al., 2025): training and benchmarks reward guessing, since "I don't know" scores the same zero as a wrong answer. The exam produced the behaviour.
- Mata v. Avianca (2023): six fabricated cases in a court filing, sanctioned — and the lawyer "verified" them by asking the same model.
- "Are you sure?" triggers sycophancy; "give me a source" can fabricate the source. Neither is a check.
- Risk concentrates on the specific, the rare, the recent and the precise: numbers, quotes, citations, links, named people, API details. A false premise in your question is a trigger of its own.
- What helps: material on the desk with traceable quoting, search plus opening the links, explicitly permitting "I don't know", and asking for a confident/unsure split. All partial.
- Verify by cost of error, not by topic. And remember that fluent explanation produces the same illusion of competence as a highlighted textbook — close it and reproduce it.
- The model has no blank cell. Only a filled one.