Why Do AI Models Hallucinate? The Science of Bluffing

Artificial Intelligence Published: 14 min read Pravesh Garcia
Why AI Models Hallucinate
Rate this post

Ask a chatbot for an obscure scholar’s birthday and you’ll often get a date. You get a specific day and year, delivered in the same calm tone it uses for the boiling point of water. Sometimes it’s right. Often it isn’t. So why do AI models hallucinate like this, and why do they sound so sure while doing it?

The usual explanation fits in one line: it’s autocomplete, so it predicts the next word. True, as far as it goes. It also explains very little. Plenty of autocomplete systems never invent a citation or a study that doesn’t exist.

A sharper answer arrived in September 2025. A team from OpenAI and Georgia Tech, Adam Tauman Kalai, Ofir Nachum, Santosh Vempala and Edwin Zhang, posted a paper bluntly titled Why Language Models Hallucinate. Their argument is uncomfortable because it isn’t really about a bug. It’s about incentives.

We built machines that score better when they bluff. Then we acted surprised when they bluffed.

That shift, from “the model is broken” to “the model is doing what we rewarded,” spreads the responsibility around. Leaderboard designers carry some of it. So does every one of us who prefers a confident answer to an honest “I’m not sure.”

What an AI hallucination actually is

“Hallucination” is a slightly misleading word borrowed from psychiatry. A person who hallucinates perceives something that isn’t there. A language model perceives nothing at all. It produces text, and sometimes that text states things that aren’t true.

A 2023 survey by Huang and colleagues at the Harbin Institute of Technology describes hallucination as “plausible yet nonfactual content.” The key word is plausible. When AI models hallucinate, they don’t produce gibberish. The output reads like an answer and fits the shape of what you asked for.

Factuality vs faithfulness

The same survey separates two kinds of failure, and the split is worth keeping in your head:

  • Factuality errors. The output contradicts the real world. A wrong date, say, or a study that was never run.
  • Faithfulness errors. The output drifts from the material you gave it. You paste in a contract and ask for a summary, and the summary mentions a clause the contract never contained.

You’ll also run into the labels intrinsic (contradicting the source) and extrinsic (adding claims the source can’t support). Vocabulary varies from paper to paper. The question that matters for you is simpler. Did the model get the world wrong, or did it get your material wrong?

The second kind is sneakier, because most of us assume the model stayed inside the document we handed it.

Why confident errors cost more than obvious ones

A wrong answer with visible doubt invites a second look. A wrong answer in flawless prose, complete with a source-shaped reference, gets pasted into a report.

Most of us treat fluency as a rough proxy for competence. With people, that shortcut works often enough. With language models it breaks down, because fluency is the one thing they produce reliably. Every tool is sometimes wrong. The danger with AI hallucinations is that the error arrives looking exactly like the truth.

Why do AI models hallucinate? Three layers of cause

The honest answer has layers, and each one makes the next worse.

  1. Pretraining. The model learns from text alone, and some facts barely appear in that text. Kalai and colleagues argue that errors here “arise through natural statistical pressures.” Learning from text alone, a model struggles to tell rare true statements from plausible false ones.
  2. Post-training. Fine-tuning makes models friendlier and more polished. The paper’s discussion suggests it can also hide symptoms without removing the underlying error floor.
  3. Evaluation. The benchmarks we use to rank models mostly give zero credit for “I don’t know.” So guessing wins.

Most explainers stop at the first layer. The interesting part, and the one with real policy weight, is the third.

Layer one: next-token prediction rewards fluency

A large language model writes one token at a time. At each step it estimates which piece of text most plausibly comes next, given everything before it. If you want the full mechanics, our guide to how large language models work walks through the architecture.

Notice what that objective rewards. Plausibility. Training asks the model to match the statistical texture of human writing. Nothing in that process checks a statement against the world.

Most of the time, plausible and true overlap. Water boils at 100°C at sea level, that sentence shows up everywhere, and the model learns it well.

The trouble starts where plausible and true come apart. Ask about a minor historical figure or a court ruling from a small jurisdiction. The model still knows what an answer should look like, right down to the citation format. It just lacks the specific fact, so it fills the shape. That gap between plausible and true is the first reason AI models hallucinate.

There’s no built-in alarm that fires when this happens, either. From the outside, a made-up answer and a well-known one come out of the same process: tokens, chosen by probability.

Layer two: facts the model saw only once

This is where the Kalai paper turns mathematical, and where it earns its keep.

The authors tie hallucination to something they call the singleton rate. That’s the share of facts that appear exactly once in the training data. Their result: for arbitrary facts, a pretrained model’s hallucination rate is at least as high as the singleton rate.

An illustration helps. Think about birthdays. A famous scientist’s birthday turns up in encyclopedias and on countless trivia pages. A private person’s birthday might appear once, in a single obituary.

Birthdays follow no rule. There’s no pattern to generalize from. Either the model saw the fact enough times to lock it in, or it didn’t.

Now suppose, hypothetically, that one in five birthdays in a training set appears only once. The bound implies you should expect the model to miss at least about one in five birthday questions. More data on famous people doesn’t help with the obscure ones.

Compare that with spelling or grammar. Those follow patterns that repeat across enormous amounts of text, so models rarely trip on them. AI models hallucinate most on arbitrary, one-off facts: names, dates, small numbers, obscure citations.

That lines up with everyday experience. Chatbots handle “explain photosynthesis” far better than “what year did this regional journal publish that paper?”

Generating an answer is harder than judging one

The paper’s second formal result is subtler. The authors compare two tasks:

  • Classification. Shown a candidate answer, decide whether it’s valid or an error.
  • Generation. Produce a valid answer yourself.

They show the generation error rate is at least twice the matching classification error rate. Put plainly, if a model can’t reliably recognize a wrong answer, it will produce wrong answers even more often.

I find this the most clarifying idea in the whole debate. It recasts hallucination as an ordinary error. Any system that sorts true from false on thin evidence, then has to commit to an output anyway, will make it.

Layer three: benchmarks that reward guessing

Picture a multiple-choice exam. A right answer earns one point. Everything else earns zero. A blank scores zero, and so does a wrong answer.

What’s the smart strategy? Never leave a blank. On a four-option question, a pure guess gives you a 25% shot at a point. An empty box guarantees nothing. Anyone who has crammed for a standardized test has worked this out.

Kalai and colleagues argue that most AI benchmarks grade like that exam. Right or wrong, nothing in between. An honest “I don’t know” earns the same zero as a confident fabrication. In their words, “language models are optimized to be good test-takers, and guessing when uncertain improves test performance.”

So the system learns, through countless small design choices by people with good intentions, that bluffing pays. Seen that way, it’s no surprise that AI models hallucinate when they could abstain.

Why labs can’t just fix it

This layer answers a question a lot of readers ask. Why don’t AI labs just train their models to say “I’m not sure” more often?

Partly because a model that abstains more will look worse on today’s leaderboards. Leaderboards shape headlines, and headlines shape which models people pick. A lab that ships a more honest model risks watching it slide down the rankings against a rival that guesses freely.

That’s why the authors don’t call for yet another hallucination benchmark. They propose changing how the existing, dominant benchmarks score answers, so those benchmarks stop penalizing uncertainty.

It echoes a point we’ve made about AI deception and the Turing test. The harder question is what the system is actually optimizing for. A hallucinating model has no intent to deceive anyone. It’s chasing a score that can’t tell confidence from knowledge.

Can post-training make hallucinations worse?

After pretraining, labs fine-tune their models with methods like reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO). We compare the main approaches in RLHF vs Constitutional AI vs RLAIF. These methods make models more helpful and much better at following instructions.

The Kalai paper’s discussion flags a catch. Post-training can suppress some visible symptoms of hallucination without touching the statistical floor set during pretraining. And if the reward signal favors confident answers over abstentions, fine-tuning has every reason to push toward confidence.

This comes from the authors’ discussion, so treat it as reasoning they haven’t measured. Still, it fits the exam logic neatly. If a human rater or an automated grader prefers a crisp answer to a hedge, the model learns crispness. Post-training leaves the root reason AI models hallucinate in place and may simply change how well they hide it.

Here’s my worry, and I can’t prove it. Some of the polish that makes modern chatbots pleasant could be the same polish that makes their mistakes harder to catch.

What’s confirmed, and what’s still debated

This topic attracts strong claims in both directions. Here’s where the evidence actually sits.

Well supported:

  • The incentive problem is real. Binary scoring gives abstention zero credit, so guessing maximizes the expected score. You can check the arithmetic yourself with the exam example above.
  • There’s a statistical floor. The singleton-rate bound and the generation-versus-classification result are formal results in the Kalai paper.
  • Grounded tools still hallucinate in real use. A Stanford RegLab and HAI study by Magesh, Surani, Dahl, Suzgun, Manning and Ho tested Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI. The tools hallucinated “between 17% and 33% of the time.” The researchers called vendor claims of hallucination-free output “overstated.”

Still debated:

  • Is hallucination inevitable? Xu, Jain and Kankanhalli offer a formal proof that language models “cannot learn all the computable functions and will therefore inevitably hallucinate if used as general problem solvers.”
  • Or is it manageable? Kalai and colleagues take a different line. In their view, some hallucination is avoidable, though training alone won’t get rid of it. They point to calibrated abstention as the practical route.

The two positions rest on different formal setups. Xu’s team studies a model acting as a general problem solver, one that must answer everything. Kalai’s team studies a model that’s allowed to say “I don’t know.”

Both can be right at once. A model forced to answer every question will sometimes be wrong. A model free to decline can, in principle, keep its stated errors low.

My read: “inevitable” is accurate and less alarming than it sounds. People are inevitably wrong too. What matters is whether a system can mark the edges of its knowledge and tell you where they are.

Can AI hallucinations be fixed?

Getting to zero looks out of reach. You can still shrink the problem and make the errors easier to spot, and the research supports a few concrete moves.

Retrieval-augmented generation helps, with limits

Retrieval-augmented generation (RAG) connects a model to a search index or a document store. Before answering, the system pulls relevant passages, and the model writes its answer from them. In theory, that grounds the output in real sources.

In practice, grounding narrows the problem without closing it.

The legal study is the cleanest evidence we have. These were commercial research tools for lawyers, built on retrieval over legal databases. They still hallucinated 17% to 33% of the time. That’s somewhere between one answer in six and one in three.

Retrieval can fetch the wrong passage. It can fetch the right one and the model can still misread it, or blend it with its own guesses. Those are faithfulness errors, and they hide very well behind a correct-looking citation.

Confidence targets and the right to abstain

The remedy Kalai and colleagues propose is surprisingly simple to state. Give the model an explicit confidence target, call it t. It answers only when its confidence exceeds t. Otherwise it abstains. Wrong answers then carry a penalty that scales as t/(1-t).

Numbers make it click:

  • At t = 0.5, the penalty ratio is 1.
  • At t = 0.75, it’s 3, so a wrong answer costs three times what a right one earns.
  • At t = 0.9, it’s 9.

Under the 0.75 rule, a model that’s only 50% sure loses points on average by guessing. Abstaining becomes the rational move. That’s calibration doing real work: the model’s honesty about its uncertainty finally affects its score.

The elegant part is that the existing benchmarks can stay. You change their rules and state them in the instructions, so model and grader agree up front on when silence beats a guess.

Benchmark reform is the lever nobody owns

This is the hard part. No single company controls the major leaderboards. Changing how they score means the people who maintain those benchmarks and the labs competing on them both accepting something awkward. A model that says “I don’t know” more often might be the better model.

And honesty has a price. Abstention cuts errors by answering fewer questions. You’ll see more “I can’t verify that.” Some people will find it irritating. Some may drift toward a rival that always has an answer ready.

How to tell if an AI is hallucinating

You can’t see inside the model. You can change how you use it. Since AI models hallucinate most on specifics, these habits go a long way:

  • Be most suspicious of specifics. Names, dates, page numbers, statistics, case citations, URLs. This is singleton territory, where the paper’s bound predicts the most trouble.
  • Open every citation. A reference that looks right proves nothing yet. Open the source and find the actual sentence. If you can’t, treat the claim as unverified.
  • Ask twice, differently. Rephrase the question or start a fresh chat. If the details keep shifting between versions, that’s a red flag.
  • Give it permission to abstain. Say it outright: “If you aren’t sure, say you don’t know.” You’re borrowing the paper’s logic and setting your own confidence target.
  • Pin it to your material. When summarizing a document, ask the model to quote the passage behind each claim. That drags faithfulness errors into the open.
  • Match the check to the stakes. A wrong bit of film trivia costs nothing. A wrong drug dose or legal precedent costs a lot. Verify in proportion.

There’s a deeper habit underneath all of these. We’ve asked whether cognitive offloading to AI builds a skill or a crutch. Hallucinations are where that question gets concrete. Hand off the remembering if you like. Keep the judging.

The machine learned to bluff from us

It’s tempting to treat hallucination as a quirk that bigger models will simply outgrow. The research points somewhere less comfortable.

If you use these tools, you have a vote in this too. Each time you pick the chatbot that always has an answer over the one that admits a gap, you send the same message a leaderboard sends. Labs build what wins. Start rewarding the honest gap, and push the benchmarks you follow to do the same.

Why do AI models hallucinate? Largely because we built a world where sounding sure scores better than being careful. The statistics we can’t fully escape. The incentives we wrote ourselves, which means we can rewrite them.

So here’s a test you can run on yourself. The next time a chatbot says “I don’t know,” will you count that against it?

Frequently Asked Questions
Why do AI models make things up?
Language models learn to produce plausible text, not verified facts. When a fact appeared rarely or only once in training data, the model has little to go on, yet it still knows what an answer should look like. Researchers Kalai, Nachum, Vempala and Zhang also show that most benchmarks give 'I don't know' zero credit, so guessing improves scores and models learn to guess.
Can AI hallucinations be fixed or eliminated?
They can be reduced and managed, but probably not eliminated. Xu, Jain and Kankanhalli give a formal proof that a model used as a general problem solver will inevitably hallucinate. Kalai and colleagues argue the practical route is calibrated abstention: answer only above a stated confidence target, and change benchmark scoring so honest uncertainty isn't punished.
Are hallucinations a bug or an inherent feature of LLMs?
Closer to an inherent property than a bug. The 2025 paper 'Why Language Models Hallucinate' shows a statistical floor: on arbitrary facts, a pretrained model's hallucination rate is at least the share of facts it saw exactly once in training. Incentives in training and evaluation then encourage guessing on top of that floor.
Does RAG stop hallucinations?
No. Retrieval-augmented generation narrows the problem but doesn't close it. A Stanford-led study of commercial legal research tools built on retrieval, including Lexis+ AI and Westlaw AI-Assisted Research, found they hallucinated between 17% and 33% of the time, and called vendor 'hallucination-free' claims overstated.
Why do AI models sound so confident when they are wrong?
Because the same process produces right and wrong answers: the model picks likely next tokens, and fluent prose comes out either way. Benchmarks that score a wrong answer the same as an abstention also reward confident guessing over hedging, so models learn that sounding sure pays.
How can I tell if an AI is hallucinating?
Be most suspicious of specifics such as names, dates, statistics, quotes and citations. Open every cited source and find the actual sentence. Rephrase the question or ask in a fresh chat to see if the answer holds, invite the model to say it doesn't know, and verify in proportion to the stakes.
Do newer models hallucinate less?
The research doesn't promise that scale fixes it. Rare, one-off facts set a statistical floor, and as long as benchmarks reward guessing, models have an incentive to bluff. That's why the authors of 'Why Language Models Hallucinate' push to change how the main benchmarks score answers.