When Verlagshaus24, a Munich publishing house, searched its own company name on Google this past January, the AI Overview that appeared at the top of the results page opened with a confident declaration: the publisher was known for “dubious business practices” and was “frequently perceived as operating a fraud scheme.” None of the underlying search results said anything of the kind. The AI had invented the accusations entirely — cross-wiring information about unrelated, genuinely disreputable companies with the plaintiff’s name, then packaging the result in an authoritative, self-contained answer with its own structure of red flags and user warnings.

This is what hallucination looks like at scale: not a quirky chatbot giving you the wrong capital of a country, but a tool hundreds of millions of people trust serving as the first and often final word on a business’s reputation — and getting it categorically wrong with no indication that anything is amiss.

On May 28, 2026, the Regional Court of Munich issued a preliminary injunction against Google in case 26 O 869/26, making it one of the first courts anywhere to hold an AI company directly liable for false statements its system generated. The court rejected the argument that users should simply verify AI answers themselves — citing Pew Research data showing that fewer than 1% of people who encounter an AI Overview source citation ever click it. Google confirmed on June 12 that it will appeal. With the EU AI Act’s transparency obligations for user-facing AI systems arriving on August 2, 2026, the legal and practical pressure on every AI answer engine — Google, Perplexity, ChatGPT Search, Bing Copilot — is only growing. But a ruling and a regulatory deadline do not tell you what to do the next time you search for something and an AI gives you a confident wrong answer.

Here is what to do.

Why AI Answers Sound Confident Even When They Are Wrong

Before the fix, the mechanism. AI Overviews, ChatGPT Search, and every related tool are built on large language models — systems trained to predict, token by token, which word or phrase most plausibly follows everything that came before it. The training objective is not truth. It is probability. When the model encounters a gap in its knowledge, it does not stop and flag uncertainty. It generates the most statistically likely continuation of whatever text pattern it is following.

Researchers call the result “hallucination,” though a more precise term, used in the technical literature, is confabulation — creative gap-filling that sounds plausible but has no factual grounding. The Munich court’s own analysis captured this precisely: the AI “makes independent, new, and substantive statements” that “do not appear in the search results at all.” The system retrieved a set of sources, then wrote past them.

That behavior is not a bug waiting to be patched. Sourav Banerjee and colleagues at DataLabs published a mathematical proof in September 2024 — drawing on Gödel’s First Incompleteness Theorem — demonstrating that eliminating hallucinations through architectural improvements, dataset enhancements, or fact-checking mechanisms is provably impossible. The core problem is that any sufficiently powerful formal system, including an LLM, will contain true statements that cannot be proven within that system. A language model that has not been exposed to a particular fact during training will guess — and it will do so with the same tone of certainty it uses when it is correct. In 2025, Michał P. Karpowicz at Samsung AI Center Warsaw extended this result using a game-theoretic framework based on the Green-Laffont theorem, proving that no LLM inference mechanism can simultaneously achieve four properties at once: truthful generation, information conservation, relevant knowledge revelation, and constrained optimality.

The practical implication is that hallucination is a permanent structural property of these systems, not a transitional quality-control problem. Rates are improving — Gemini 2.0 Flash achieves a 0.7% hallucination rate on controlled summarization benchmarks — but those benchmarks measure faithfulness to a given document, not factual accuracy on open-ended questions. On more realistic tasks — multi-turn research questions, legal citation retrieval, and medical queries — error rates remain substantially higher. A 2025 study from Stanford’s RegLab and Human-Centered Artificial Intelligence group found that purpose-built legal AI tools still hallucinated between 17% and more than 34% of the time on challenging legal research queries.

A separate confounding factor is training. Most frontier models are refined through a process called reinforcement learning from human feedback, in which human evaluators rate AI outputs and those ratings shape the model’s future behavior. The problem: human raters consistently prefer responses that sound confident and helpful over responses that say “I don’t know.” The model learns to suppress uncertainty and produces what researchers call sycophantic answers — plausible, fluent, and wrong.

There are a few reliable triggers for the worst hallucination:

Gaps in training data. If a fact is rare in the training data, the model fills the gap with whatever text pattern continues most plausibly. Plausible and true are not the same.

Recent, niche, or contested topics. All language models have a training cutoff. They know nothing reliable about events after it, and struggle with subjects that were sparsely covered before it.

Granular specificity. The more precise an answer sounds — exact statistics, full names with titles, specific dates — the more likely it is to contain an invented detail. Precision signals authority to both the model and the reader, which is exactly why the model produces it.

The “I don’t know” suppression effect. Reinforcement learning from human feedback training penalizes hedging. A model that says “I’m uncertain about this” scores poorly on human preference metrics, so models learn not to say it.

An Oumi analysis commissioned by the New York Times found that AI Overviews running on the current Gemini model produce inaccurate answers roughly 9% of the time. More striking: over half of even the correct answers could not be traced back to the sources the AI cited. The model reached accurate conclusions through processes that have no visible relationship to the evidence it pointed to.

Applied to Google’s search volume, a 9% error rate translates to a number of wrong answers per hour that dwarfs the total number of corrections ever issued. The Pew Research Center’s July 2025 study of nearly 70,000 actual Google searches from 900 U.S. adults found that when an AI Overview appeared, users clicked through to any source link — including the ones Google displayed — in just 1% of visits. The click-through rate on traditional search results dropped from 15% to 8% in the presence of an AI summary, and 26% of sessions ended immediately after the AI answer appeared. The Munich court cited the Pew data directly when dismissing Google’s argument that users could simply verify answers themselves. Verification that 99% of users never perform is not a meaningful safeguard.

How to Fact-Check an AI Answer Before You Act on It

None of this makes AI search tools useless. For low-stakes, uncontroversial queries — a recipe, a common phrase’s meaning, a widely-known historical fact — they are genuinely efficient. The risk concentrates in specific categories: anything about a named person or business, any statistic or percentage, any recent event, any medical or legal claim, any subject where being wrong has real consequences for you.

In those cases, the following three steps take between 90 seconds and five minutes and close the gap that courts are only now starting to address.

Step 1: Open at Least One Cited Source and Read Past the Headline

This is the single most valuable check, and it is almost never performed. When an AI Overview or AI-powered answer includes source links, open the one that appears most directly relevant to the specific claim that matters to you.

You are not looking for general agreement. You are looking for the specific claim — the statement, figure, or characterization — actually written by someone other than the AI. This is the gap the Munich court named: in the Verlagshaus24 case, the AI Overview produced accusations that appeared in none of the articles it cited. The court found this disqualifying. The AI had written its own conclusion, attached someone else’s byline by proxy, and served it as a search result.

If you cannot find the specific claim in any of the cited sources, the AI may have synthesized across sources in ways that introduced new meaning — or invented the claim outright. Treat it as unverified.

A useful signal: when an AI summary cites three or more sources but does not quote any of them directly, it is synthesizing. Synthesis is where new claims get introduced. Any synthesis that produces a specific, consequential assertion about a named person or company warrants direct verification against at least one primary source.

Step 2: Rephrase the Question and Ask a Second System

AI answers are inconsistent in ways that reveal underlying uncertainty. Genuine facts are stable: asking “What is the capital of France?” in ten different ways will reliably return “Paris.” But in the zone where hallucination concentrates — niche topics, recent events, precise statistics, claims about specific named entities — answers shift when the framing changes.

Ask the same question differently, or ask it to a different AI system entirely. If you receive a materially different answer, you have direct evidence that at least one response is wrong. The divergence itself is useful information.

You can also prompt the model to surface its own uncertainty directly: “How confident are you in this? If you’re uncertain, say so.” Most current AI systems have some capacity for calibrated hedging — they suppress it by default because confident-sounding answers score better in user feedback, but explicit prompting can reactivate it. An answer that shifts to “I’m not certain, but I believe…” in response to this prompt is signaling genuine uncertainty.

Be especially alert when an AI produces a highly specific answer — a named study, a percentage, a direct quote — but cannot locate the underlying source when you follow up. That pattern is among the most reliable indicators of confabulation.

Step 3: Go Directly to the Source That Would Have Produced the Fact

For anything you are going to repeat, publish, cite, or make a decision based on, skip the AI intermediary entirely and find the primary source.

This means going to the institution, publication, or dataset that originated the claim. If the AI describes “a 2024 study,” find the study. If it attributes a statement to a named individual, find what the individual actually said in their own words. If it makes a claim about a company’s track record, check the company’s official communications or a publication that reported on it directly.

This step matters most for: statistics and percentages, quotes attributed to named people, claims about specific named businesses or individuals, recent events within the past year or two, and medical or legal information.

The Munich case illustrates what happens when this step is skipped at scale. Google’s AI Overview confidently linked two publishers to “scam” operations. Anyone who had searched either publisher directly would have found no such associations. The AI constructed a narrative from patterns in its training data — patterns that happened to fit the wrong company.

What Courts Cannot Fix and Users Must

The Munich ruling is preliminary, and Google is appealing it. Even if upheld, it applies to defamatory claims about identifiable parties under German law — not to the everyday smaller falsehoods that make up the bulk of AI hallucination.

The EU AI Act’s August 2, 2026 deadline will require companies to disclose when users are interacting with AI-generated content and to ensure that outputs are labeled as machine-generated in a machine-readable format. Those rules address transparency. They do not require that AI answers be accurate, and they do not provide individual users with any enforcement mechanism when an AI gives them the wrong medical dosage, invents a legal citation, or mischaracterizes a business they are considering hiring.

The structural condition that produces hallucination — the fundamental gap between predicting probable text and knowing what is true — is not addressable through regulation or through any current technical approach. The mathematical proofs from 2024 and 2025 establish this as a property of the architecture itself, not a quality-control gap to be closed by the next model release.

What that means practically: the confidence of an AI answer and its accuracy are two separate things. The system is optimized for the former. The latter requires a human verification step, and the three steps above are that step. A court confirmed that Google cannot offload this burden onto users by attaching source links. What courts cannot do is perform the verification for you. That part is still yours.

Frequently Asked Questions

Can Google AI Overviews be wrong, and does Google know it?

Yes, on both counts. An analysis by Oumi, an AI startup commissioned by the New York Times, found that AI Overviews using the current Gemini model return inaccurate answers roughly 9% of the time, and that more than half of even the correct answers cannot be traced to the sources Google cites. Google has publicly stated that AI Overviews can “occasionally miss context or misinterpret web content” — the same acknowledgment the Munich court found insufficient when those misinterpretations amounted to fabricated defamatory claims about named companies.

Why does AI give incorrect information even when it cites sources?

Large language models generate text token by token, selecting each word based on its statistical likelihood given everything that preceded it. The training objective is to predict probable text — not to verify truth. When an AI system retrieves source articles and then writes a summary, it is not copying from those sources. It is generating new text that patterns like a summary. That generation process can introduce claims that appear in none of the source articles. The Munich court identified exactly this mechanism in the Verlagshaus24 case, where the AI’s output contained accusations that did not appear in any of the linked sources.

Is Google legally responsible for wrong answers in AI Overviews?

In at least one jurisdiction, the answer is now yes. The Munich Regional Court ruled on May 28, 2026, that Google is directly liable for false statements generated by AI Overviews, because the AI produces “independent, new, and substantive statements” rather than simply pointing to third-party content. Google is appealing this ruling, and it covers only specific defamatory claims about identifiable parties under German law. US law, particularly Section 230 of the Communications Decency Act, provides broader platform protections. The EU AI Act’s August 2, 2026 enforcement deadline will add transparency obligations but not accuracy requirements.

What types of AI answers are most likely to be wrong?

Hallucination concentrates in specific categories: claims about named individuals or businesses, precise statistics and percentages, events that occurred after the model’s training cutoff, niche or specialized topics, and legal or medical information. Paradoxically, highly specific-sounding answers — with exact figures, named studies, and direct quotes — carry higher hallucination risk than general ones, because the model produces specificity to signal authority, not because it has specific knowledge to draw on.