Home / Blog / AI Hallucination Statistics 2026
Data · 2026

AI Hallucination Statistics 2026: Rates, Benchmarks & Why They Disagree

Published hallucination rates in 2026 span from 0.7% to 94%. That is not measurement noise — it is four different questions being asked and reported under one word.

By · Published August 17, 2026 · 9 min read

If you have seen "AI hallucinates 3% of the time" and also "AI hallucinates most of the time," both citations were probably accurate. The gap is the benchmark. Grounded summarization — hand the model a document, ask it to summarize only what is in it — produces low single digits. Open recall — ask the model a factual question with no source — produces double digits and worse. Generating SQL against a real enterprise schema produces the worst numbers of all. Below are the figures worth knowing, each with the test that produced it.

1.8%
best hallucination rate on grounded summarization — 98.2% factual consistency
Vectara HHEM leaderboard
24.2%
worst rate among models listed on the same leaderboard
Vectara HHEM leaderboard
22–94%
reported range on open-recall style factual benchmarks
Stanford HAI
52.8%
of examples in a leading text-to-SQL benchmark found to carry annotation errors
Jin et al., UIUC 2026

Grounded summarization: the flattering number

Vectara's Hallucination Leaderboard is the most-cited source in this space and the one behind most low headline figures. Its method is narrow and worth understanding before you quote it: over 7,700 articles across news, technology, science, medicine, legal, sports, business and education are fed to each model, which is asked to summarize each one using only the facts presented in the document. A detector then flags output that introduces unsupported content.

As of the 11 May 2026 update:

This measures intrinsic hallucination — fabrication relative to a source you provided — not general factuality. A 1.8% score means the model is good at not inventing things when the answer is sitting in front of it. It says nothing about what happens when the answer is not.

The same model, 0.7% or 7.6%

The clearest demonstration that the benchmark makes the number: Gemini-2.0-Flash-001 scores 0.7% on grounded summarization. Run the same model against Vectara's harder FaithJudge evaluation and its effective rate rises to 7.6% — more than a tenfold difference, same model, same week.

Widen to open-recall benchmarks, where the model answers from its own parameters with no document to anchor to, and Stanford HAI-style evaluations report rates from roughly 22% to 94% depending on the domain and how strictly "correct" is defined.

So a vendor claiming a 2% hallucination rate and a researcher claiming 60% can both be telling the truth. Always ask which test.

The reasoning-model paradox

An intuition worth discarding: that models which "think longer" hallucinate less. Multiple 2026 evaluations report the opposite on factual recall — reasoning models including OpenAI's o3 and Grok-4-fast show higher hallucination rates than non-reasoning models on those tasks.

More deliberation produces more elaborate output, and elaborate output has more surface area on which to be wrong. Buying a more expensive reasoning tier does not buy factual reliability.

The number that matters for business data: text-to-SQL

Summarization benchmarks tell you little about what happens when you point an assistant at your company's database. For that the relevant literature is text-to-SQL, and it has moved fast enough that this section used to be wrong.

The figure that circulated for two years was the Spider 2.0 cliff: the same model scoring 86.6% on the clean academic schemas of Spider 1.0 and around 10% once the database looked like a real enterprise system. That gap was real, and the general lesson survives — accuracy on a toy schema does not predict accuracy on yours, and the failure is silent, because a wrong query returns a plausible number rather than an error.

The specific numbers, however, no longer hold. Purpose-built agent systems now top the Spider 2.0 leaderboard above 96%, and a 2026 audit found that more than half of the examples in the benchmarks producing all of these figures contain annotation errors in the first place. We pulled the current evidence apart in a dedicated article: how accurate AI actually is at analysing data, covering text-to-SQL and text-to-Python, a sixteen-model comparison, and why the leaderboards are worth less than they look.

Why wrong answers do not look wrong

A hallucinated summary is often detectable by reading it. A hallucinated number is not. If an assistant joins two tables incorrectly and reports net sales of $1.83M instead of $1.94M, nothing about the output signals a problem. It is formatted correctly, it is the right order of magnitude, and it arrives with a confident explanation.

There are two independent sources of variance stacked on top of each other:

Both can be present at once, and neither announces itself. We wrote about the practical version of this in why AI gives a different number every time.

From twenty-five years of this

Running business intelligence teaches you that an error rate is the wrong thing to shop on, and it is the only thing these benchmarks report. What determines whether you can work with a source is the shape of its errors, not their frequency.

I have relied on reports that were wrong far more than 5% of the time and it was fine, because they were wrong in a knowable direction — the regional feed always landed a day late, so the current month always understated, and every person reading it knew to allow for that. I have also had to stop using a source that was accurate the overwhelming majority of the time, because when it was wrong it was wrong unpredictably and by an unremarkable amount. Nobody could tell the good months from the bad ones, so every figure had to be checked, which is the same as having no report at all.

That is what the section above is really about. A 3% error rate distributed at random through your numbers is not a small problem — it is an unusable source dressed as a good one. When you evaluate any of this on your own data, ask what happens when it is wrong, not how often.

What actually reduces the error rate

The pattern across all four benchmark families is consistent: models hallucinate least when they are given the answer and asked to relay it, and most when they are asked to derive it. Grounded summarization scores 1.8%; open recall scores up to 94%. That difference is the entire design principle for reliable AI over business data.

Practically, that means moving work out of the model:

This is the approach Quiriz takes — governed metric definitions that compile to SQL, arithmetic executed by the database, and the model used for interpreting the question and phrasing the answer rather than for computing it. It does not make a language model factual; it removes the steps where a language model was being asked to be. For where each general assistant stops, see our comparisons with ChatGPT, Claude, Copilot and Gemini, and our 2026 AI pricing comparison for what each costs.

If you are quoting these numbers

Cite this page. Quiriz, "AI Hallucination Statistics 2026: Rates, Benchmarks & Why They Disagree," 17 August 2026. https://quiriz.co/blog/ai-hallucination-statistics-2026.html

Sources

Figures compiled from published benchmarks and leaderboards as of 17 August 2026. Hallucination rates are not comparable across benchmarks and should always be quoted together with the evaluation that produced them; leaderboard positions change frequently. Model names reflect those published by the benchmark maintainers at the time of the run and may not correspond to the current release of a given product.