Analysis Finds That Google’s AI Overviews Are Providing Misinformation at a Scale Possibly Unprecedented in the History of Human Civilization

InstitutionalApril 8, 2026·6 min read

Google’s AI Overviews are delivering factually incorrect information to tens of millions of users every hour, according to independent testing, raising urgent questions about the company’s obligation to validate AI outputs before mass deployment. For institutional investors tracking AI governance risk and reputational liability in big tech, this reveals a structural gap between private testing standards and public-facing accuracy thresholds.

  • Google AI Overviews achieve 91% accuracy on factual queries, but process five trillion searches annually, generating tens of millions of wrong answers hourly.
  • Gemini 2, the previous model powering Overviews, achieved only 85% accuracy and provided ungrounded answers 37% of the time before Google upgraded it to Gemini 3.
  • Google’s own internal testing found Gemini 3 produces incorrect information 28% of the time, contradicting the company’s public statements about accuracy improvements.
  • 91% Accuracy rate of current Gemini 3 AI Overviews versus 85% for previous Gemini 2
  • 5 trillion Annual Google search queries generating tens of millions of erroneous answers per hour
  • 37% Rate Gemini 2 provided ungrounded answers citing non-existent or irrelevant sources

Independent testing by AI startup Oumi has exposed a critical gap between Google’s public claims about AI Overviews and the actual error rate of the feature, which appears above standard search results in a prominent position designed to capture user attention before organic results.

The analysis, conducted for The New York Times using SimpleQA, a benchmark developed by OpenAI and widely used across the industry, tested AI Overviews across two major model versions. The current version, powered by Gemini 3, delivers factually accurate responses 91 percent of the time, while its predecessor, Gemini 2, managed only 85 percent accuracy.

At face value, a 91 percent accuracy rate may suggest a functioning system. However, when scaled against Google’s processing of approximately five trillion search queries annually, that 9 percent error rate translates to roughly 450 billion incorrect answers per year, or tens of millions of wrong responses every single hour.

Gemini 2 Operated with Ungrounded Citations 37 Percent of the Time Before Upgrade

The Oumi analysis revealed a more granular and troubling problem with the earlier Gemini 2 model: it generated answers that were “ungrounded” 37 percent of the time, meaning the AI Overviews cited websites, studies, or sources that either did not exist, did not contain the cited information, or were irrelevant to the query.

This distinction matters enormously for institutional investors assessing AI governance risk. An answer can be grammatically coherent and presented with confidence while simultaneously fabricating its evidentiary foundation, a phenomenon researchers call “hallucination.”

Google deployed this less accurate version to its entire user base for months before implementing the upgrade to Gemini 3 in February, following the October round of Oumi’s testing.

The company chose to continue serving a model it knew underperformed, conducting what amounts to a live experiment on hundreds of millions of users without explicit consent or disclosure of the known accuracy shortfall.

This decision reflects a fundamental asymmetry: the company had access to internal testing data showing the model’s limitations, yet chose to optimize for feature launch velocity rather than validation thresholds before public rollout.

The upgrade to Gemini 3 did improve performance, lifting accuracy from 85 percent to 91 percent. However, that improvement may mask a deeper structural issue. Google’s own internal analysis of Gemini 3 found that the model produces incorrect information 28 percent of the time, a figure that diverges sharply from the 9 percent error rate Oumi reported.

This discrepancy suggests either that Google’s internal testing methodology is more lenient, that Oumi’s SimpleQA benchmark captures failure modes Google’s tests miss, or that the model performs differently depending on the distribution of queries it receives.

User Behavior Studies Show 8 Percent Question AI Answers, 80 Percent Trust Errors

The risk created by AI Overviews extends beyond technical accuracy into human behavior. Research cited in the analysis indicates that only 8 percent of users actually fact-check an AI-generated answer before accepting it as true. This baseline trust asymmetry, where AI outputs receive deference by default, compounds the consequences of even moderate error rates.

A 9 percent failure rate becomes far more damaging when 92 percent of users will not independently verify the information.

An additional experiment found that users continued to trust AI systems and act on their guidance nearly 80 percent of the time even when given clearly wrong answers. Researchers labeled this phenomenon “cognitive surrender,” describing a state in which users defer judgment to the system rather than engaging critically with its output.

Large language models naturally adopt an authoritative tone, presenting fabricated information with the same linguistic confidence they use for correct answers.

When that output appears in a featured position at the top of a search results page, a location historically reserved for Google’s own validated information, users have no structural cue that the information is machine-generated and therefore subject to hallucination.

The convenience of receiving a pre-composed answer without clicking further into results amplifies adoption. Users facing time pressure or low-stakes queries are unlikely to invest effort in verification.

Google Disputes Oumi Findings While Internal Tests Confirm Accuracy Problems

Google rejected the Oumi analysis in a statement to the New York Times, with company spokesman Ned Adriance asserting that the study “has serious holes” and “doesn’t reflect what people are actually searching on Google.” The company’s critique centers on the methodological claim that SimpleQA, the benchmark used for testing, does not represent the actual distribution of user queries.

Google argues that because it has access to real search behavior data, its internal testing produces more representative results.

However, Google’s own internal analysis undermines this defense. The company’s testing of Gemini 3 found a 28 percent error rate for factual questions, nearly triple the 9 percent rate implied by Oumi’s 91 percent accuracy figure.

This divergence is not resolved by methodological disagreement; rather, it suggests that Google’s public-facing claims about accuracy may be calibrated to how questions are framed or selected, rather than to a transparent measure of real-world failure rates.

Google has not publicly disclosed how it defines “accuracy” in its internal testing, what query distribution it uses, or how its benchmark differs from SimpleQA in ways that would explain a threefold difference in reported error rates.

Google does claim that AI Overviews are more accurate than the underlying model alone because the feature draws on Google search results before composing an answer. This layering of retrieval-augmented generation should theoretically improve accuracy by grounding responses in indexed content.

Yet Oumi’s finding that Gemini 2 provided ungrounded citations 37 percent of the time suggests this retrieval mechanism itself is failing, the model is either hallucinating sources or misrepresenting what it retrieved.

Institutional Investors Face Reputational and Regulatory Exposure for Google’s AI Governance Gap

For institutional investors holding Google stock or evaluating the company’s risk profile, the AI Overviews situation presents three distinct concerns. First, reputational damage accumulates as users experience or hear about incorrect information presented with unwarranted authority.

Search is core to Google’s brand; users have trusted search results for two decades based on algorithmic ranking of published content, not on AI fabrication. Each viral example of an obviously wrong AI Overview erodes that equity.

Second, regulatory exposure is increasing. Regulators in the EU, US, and elsewhere are developing frameworks for AI governance that may impose liability for false or misleading outputs, particularly when those outputs are presented in positions of prominence.

The US Federal Trade Commission has already raised questions about whether AI-generated summaries require explicit disclosure of their origin. If regulators determine that Google misled users by presenting machine-generated hallucinations in a format designed to suggest editorial validation, enforcement actions or fines could follow.

Third, the gap between Google’s internal testing and public claims raises governance questions for shareholders. Board oversight of AI deployment, sign-off on public statements about accuracy, and disclosure of known limitations before rollout are all materials for institutional

Get this in your inboxThe Crypto Coin Show newsletter covers the policy and market moves institutional crypto investors are pricing in.

Subscribe