AI Answers Are Quietly Deciding Who Looks Credible

On March 13, 2026, a distributed crawler began feeding trending questions into Google Search. Forty days and 55,393 queries later, a large-scale Google AI Overviews preprint reported a stark measure of AI answer credibility: Google’s AI Overviews often cited reputable domains, yet 11% of the 98,020 factual claims the researchers extracted were not supported by the linked pages. The source could look trustworthy while the sentence beside it still outran the evidence.

That gap is becoming consequential at extraordinary scale. In May 2026, Google reported that AI Overviews had more than 2.5 billion monthly active users and AI Mode had passed one billion. OpenAI, Perplexity and other answer engines are also training people to expect a finished response before they inspect a source list. The new gatekeeper is not simply ranking pages. It is writing the paragraph that tells readers which company, expert, publisher or institution deserves belief.

The central risk is easy to miss because the interface feels helpful. Citations, confident prose and a familiar brand name create the appearance of verification. But AI systems make several editorial decisions at once: which sources enter the answer, which are omitted, how their claims are compressed, how uncertainty is phrased and which entities are named positively. Those decisions can manufacture a credibility advantage long before a reader clicks anything.

![](https://unsplash.com/photos/a-group-of-people-looking-at-a-computer-screen-nK3JmYpTPKs)A team reviews information on a shared screen. AI visibility has become a cross-functional concern for communications, search, legal and reputation teams. Photo by Mushvig Niftaliyev, used under the Unsplash License.

The answer box has become a credibility shortcut

Traditional search results made authority visible as a set of competing candidates. A user could scan domains, headlines, dates and snippets, then decide which page deserved attention. AI search collapses that process into a single synthesized answer. The system still retrieves sources, but it performs the first round of comparison on the user’s behalf and presents its conclusion in polished prose.

Google describes AI Mode as an end-to-end search experience that breaks complex questions into multiple searches and reasons across the results. OpenAI says ChatGPT search blends a conversational interface with timely web information and source links. Those product choices are useful: they reduce the work required to explore a complicated topic. They also give the answer engine an editorial power that a conventional results page did not possess—the power to decide which claims become the default account.

OpenAI’s October 31, 2024 announcement framed the change plainly: users would receive fast answers with links to web sources. The official post below is promotional, not independent evidence, but it captures the product promise that has conditioned users to treat a linked answer as a researched answer.

>

🌐 Introducing ChatGPT search 🌐

ChatGPT can now search the web in a much better way than before so you get fast, timely answers with links to relevant web sources.https://t.co/7yilNgqH9T pic.twitter.com/z8mJWS8J9c

— OpenAI (@OpenAI) October 31, 2024

The important word is “links.” A link signals provenance, but it does not by itself show that every sentence is supported, that the best source was selected or that competing evidence was considered.

A 2025 University of Notre Dame experiment found that citations increased participants’ trust in an AI-generated answer even when the citations were random. Trust fell when participants actually checked the links. The finding does not prove that every user behaves this way, but it identifies the psychological mechanism that makes AI answers powerful: citation formatting can function as a credibility cue before it functions as evidence.

Citation is not verification

The most visible evidence of this distinction came from the Tow Center for Digital Journalism’s citation audit. Researchers supplied eight generative search tools with exact quotations from news articles and asked each system to identify the article title, publication date, publisher and URL. Across 1,600 queries, the tools failed to retrieve the correct information more than 60% of the time. Perplexity performed best in that test and was still incorrect 37% of the time; Grok-3 Search was incorrect 94% of the time.

The test was deliberately narrow. It measured source retrieval from quotations, not every kind of factual question, and it captured products at a particular moment in early 2025. Even so, it exposed a pattern that matters for credibility: systems often answered when they should have declined, generated fabricated links, cited copied or syndicated versions instead of original reporting, and expressed wrong answers with confidence.

The Tow Center’s interpretation was especially revealing. News publishers were not only information suppliers; their names were being used to legitimate the chatbot’s answer. When an AI response cites Reuters, the BBC, a university or a government agency, some of that institution’s accumulated trust transfers to the sentence—even when the sentence is a loose paraphrase, an unsupported synthesis or a misattribution.

What users did after Google showed an AI summary

Share of observed Google search visits among 900 U.S. adults, March 2025

Click and session-ending rates with and without Google AI summaries Traditional result clicks occurred on 15 percent of visits without an AI summary and 8 percent with one. Users clicked an AI summary source on 1 percent of visits. Sessions ended after 26 percent of visits with an AI summary and 16 percent without one.

-

0% 10% 20% 30% 40%

Clicked traditional result Clicked AI-summary source Ended session after AI summary Ended session without AI summary

15% without AI summary 8% with AI summary 1% 26% 16%

Takeaway: When an AI summary appeared, users were about half as likely to click a traditional result and rarely opened the summary’s cited sources.

Source: Pew Research Center, published July 22, 2025. Data reflect 68,879 Google searches observed in March 2025; search pages were recollected in April 2025, and AI summaries may change over time. Data as of April 17, 2025.

Pew Research Center’s browsing analysis shows why this matters commercially and epistemically. When an AI summary appeared, users clicked a traditional search result in 8% of visits, compared with 15% when no summary appeared. Only 1% clicked a source inside the summary. Users also ended their browsing session more often after seeing an AI summary. The study covered U.S. adults and Google searches during a defined period, so it should not be generalized to every country, platform or query. It still demonstrates that the synthesized answer often becomes the final stop.

OpenAI’s official product demonstration shows the intended experience: a question, a concise answer, a source panel and conversational follow-ups. Notice how the sources are available but visually secondary to the response. The video explains product functionality; it does not independently validate source selection or accuracy.

OpenAI’s official “Search—12 Days of OpenAI: Day 8” demonstration. Watch how the answer remains central while source inspection is optional.

The design is efficient, but efficiency changes user behavior: the answer engine receives the first opportunity to frame the issue, while the publisher receives a smaller opportunity to correct it or supply missing qualifications.

A credible source can still support an unsupported claim

The 2026 Google AI Overviews preprint complicates any simple claim that AI search merely favors low-quality websites. Researchers Haofei Xu, Umar Iqbal and Jacob M. Montgomery found that AI Overview-cited domains were, on average, more credible than the traditional first-page results displayed beside them. Nearly 30% of cited domains did not appear on that first page, suggesting a source-selection process distinct from conventional ranking.

That is the strongest counterargument to the idea that AI answers automatically degrade information quality. Retrieval can surface reputable institutions and pages that ordinary ranking might miss. Google has also added Preferred Sources, “Highly Cited” labels and more prominent links intended to expose original reporting and user-selected publishers. These changes show that answer engines can be designed to make provenance more visible.

But source quality and claim fidelity were largely independent in the same study. Of 98,020 atomic claims extracted from 7,491 verifiable AI Overviews, 89% were judged consistent with the cited pages and 11% inconsistent. Most failures were omissions: the AI stated something that the cited text did not support. A respected domain can lend its authority to a claim it never made.

This is a subtle failure, harder to detect than a fabricated URL. The citation exists. The page exists. The institution is credible. A hurried reader sees all the expected signals and moves on. Only a claim-by-claim comparison reveals that the answer has crossed from synthesis into invention or overstatement.

Google’s official May 2025 announcement presented AI Mode as Search “transformed,” with Gemini 2.5 at the center. The post is useful evidence of product direction and scale, but its claims about helpfulness should be read alongside independent audits rather than treated as a quality measurement.

>

AI Mode is Search transformed with Gemini 2.5 at the core. It’s our most powerful AI search, with more advanced reasoning and multimodality, and the ability to go deeper through follow-up questions and helpful links to the web.

Here’s a peek at what’s coming soon to AI Mode: 🧵

— Google (@Google) May 20, 2025

The announcement emphasizes reasoning and depth. Independent evaluation must ask a different question: whether each generated claim is traceable to the evidence the interface presents.

Google’s I/O 2025 Search segment gives the clearest official view of that ambition. The demonstration shows AI Mode handling complex, multi-part questions through query fan-out and synthesized responses. What it cannot show is how the system behaves across thousands of ordinary queries, ambiguous sources and repeated runs.

Google’s official Search segment from I/O 2025. The demo illustrates how AI Mode decomposes complex questions and returns one integrated response.

Product demonstrations establish capability; longitudinal audits establish reliability. Credibility requires both.

The source pool is narrower and less stable than it looks

AI answers create another distortion before any sentence is written: source selection. The 2026 “Answer Bubbles” preprint compared 11,000 real queries across conventional Google Search, Google AI Overviews, a search-enabled GPT system and a non-search GPT system. It found systematic differences in the domains selected and in the language used to summarize them. Wikipedia and longer pages were overrepresented, while negatively framed sources and cited social content were underrepresented.

The researchers also reported that adding search reduced hedging by as much as 60% while preserving confident language. That combination matters. Retrieval may improve factual grounding, yet the finished answer can sound more certain than the underlying source set warrants. The system does not only choose evidence; it chooses how much uncertainty survives the rewrite.

Another large preprint covering 24,000 queries across 243 countries found AI search exposed users to fewer long-tail information sources and lower result variety than traditional search. The authors also reported differences in source credibility and political orientation, although those results depend on their classification methods and query set. The broader point is less controversial: when a system produces one answer, source diversity becomes a product-policy decision rather than a visible list the user can compare.

Answer volatility compounds the problem. Identical prompts can return different brands, experts and links across platforms or repeated runs. Small wording changes can alter the retrieval path. A company may look authoritative in one answer, disappear in another and be described inaccurately in a third. That is not a stable “ranking” in the SEO sense; it is a probabilistic reputation surface.

Publishers and brands now compete for machine-authored reputation

The economic dispute over AI search is often described as a traffic fight, but it is also a credibility fight. In July 2025, a coalition of independent publishers filed an EU antitrust complaint alleging that Google’s AI Overviews used publisher content while reducing traffic, readership and revenue. Google said AI Overviews create new discovery opportunities and send billions of clicks to websites. Both positions can be partly true: total outbound traffic may remain large while particular publishers or query categories lose visits.

Penske Media later sued Google, arguing that AI Overviews used reporting from titles including Rolling Stone, Billboard and Variety while weakening the commercial value of the original pages. Encyclopedia Britannica and Merriam-Webster brought a separate case against Perplexity, alleging that its answer engine copied material and diverted traffic. These are allegations, not final judicial findings, but they identify the underlying asset being contested: not only content, but the authority attached to the content.

When an answer engine cites a publisher without sending a visit, it extracts some reputational value while leaving the publisher with less opportunity to build a direct audience. When it names a company as a recommended vendor, it transfers visibility and implied endorsement. When it omits an expert, the expert may lose consideration without knowing a query ever occurred.

![](https://unsplash.com/photos/people-working-at-computers-in-a-modern-office-Ds5FesTkKhk)Teams can no longer infer AI visibility from web traffic alone because many users consume the synthesized answer without opening a cited page. Photo by tommao wang, used under the Unsplash License.

This is why brand teams are beginning to treat AI answers as a separate measurement channel. Conventional analytics can show rankings, impressions, clicks and conversions. They cannot show how a chatbot characterized a company when the user never visited the company’s site. The missing data include mention frequency, recommendation position, cited domains, sentiment, factual errors, competitor inclusion and changes over time.

Why familiar names often win

AI systems have practical reasons to favor sources that are easy to retrieve, parse and reconcile. Well-known publishers have extensive internal linking, clear metadata, strong external references and large archives. Government and university sites often use stable page structures and explicit institutional language. Large brands publish product documentation, support pages, comparison material and third-party coverage that reinforce the same entity relationships.

The GEO-16 observational study, based on 1,702 citations from Brave, Google AI Overviews and Perplexity across English-language B2B software queries, found the strongest associations with metadata and freshness, semantic HTML and structured data. Higher overall page-quality scores were associated with higher citation rates. It does not prove that adding schema or updating a date will cause an AI system to cite a page.

Still, the result explains why credibility can become self-reinforcing. Sources with better technical publishing systems are easier for machines to interpret. Sources already cited gain visibility. That visibility can produce more links, mentions and branded searches, which create more signals for future retrieval. Smaller specialists may possess superior knowledge but weaker machine-readable evidence of authority.

The answer is not to imitate authority with synthetic biographies, mass-produced pages or decorative citations. Those tactics increase the volume of material an AI system might encounter while degrading the information environment. The defensible path is to make real expertise legible: named authors, verifiable credentials, primary documents, dated updates, transparent methods, precise claims and corrections that remain visible.

A serious AI credibility audit measures more than mentions

An organization cannot evaluate this channel by asking one chatbot one flattering question. A useful audit begins with a controlled prompt set tied to real decisions: category discovery, product comparison, executive reputation, safety, pricing, alternatives, implementation, criticism and post-purchase support. Each prompt should be repeated across platforms, locations and dates because outputs vary.

Presence: whether the organization, its competitors and relevant experts appear at all.

  • Position and framing: whether the entity is recommended, merely listed, criticized or treated as a category authority.
  • Citation provenance: which URLs support the answer, whether the original source is used and whether the source actually supports the claim.
  • Factual consistency: whether products, leadership, locations, policies, capabilities and dates are described accurately.
  • Source diversity: whether the answer depends on one domain type, one publisher or a narrow cluster of repeated pages.
  • Volatility: how much the answer changes across runs, models and time.

Platforms built for AI visibility monitoring can help operationalize that work. iSentinel AI is one example of the emerging category. The useful question is not whether a dashboard produces a single “visibility score,” but whether it preserves the underlying answers, citations and dates so a human reviewer can verify what changed and why.

Audit the answer, not just the ranking

Track representative prompts, save the generated wording, open every important citation and compare AI descriptions with your verified source material.

Explore iSentinel AI

The audit should feed a response process. Communications teams need a way to correct outdated third-party pages. Product teams need canonical documentation for features and limitations. Legal and compliance teams need visibility into high-risk misstatements. Editorial teams need to publish evidence in passages that can be quoted without stripping away scope, dates or caveats.

What not to optimize for

The temptation is to treat AI answers as another ranking system and chase whatever appears to produce a citation. That is premature. The major engines use different retrieval stacks, source partnerships, model versions and interface rules. Studies repeatedly find limited overlap between the domains cited by different systems. A tactic that raises visibility in one engine may do nothing in another.

More importantly, citation volume is not the same as credibility. A brand can be cited frequently for criticism. A page can be quoted while the answer misstates its conclusion. A company-controlled source can dominate the narrative without independent corroboration. A “share of voice” metric can rise while factual consistency falls.

Teams should reject three shortcuts: publishing unsupported superlatives for machines to repeat, disguising promotional claims as neutral education, and flooding third-party sites with coordinated messaging. These methods can create temporary visibility, but they increase correction risk and make the organization dependent on opaque systems that may later discount the same patterns.

The more durable objective is evidence density. Important claims should have a clear owner, date, scope, method and primary source. Product pages should distinguish observed results from projections. Case studies should label company-reported outcomes. Comparisons should state the market and criteria. Corrections should be explicit rather than silently overwritten. These practices help readers first; machine extraction is the secondary benefit.

The next credibility crisis will look ordinary

The most dangerous AI answer will not necessarily contain an absurd hallucination. It will look routine: a calm paragraph, three reputable citations and a confident recommendation. One source will not quite support the sentence. A relevant expert will be absent. A caveat will disappear during compression. The user will not click, and the answer will become the remembered version.

Google, OpenAI and other platform companies are adding more links, source controls and quality features. Independent research also shows that AI retrieval can surface credible domains and useful information. The unresolved test is whether these systems can make claim-level support, uncertainty and source choice as visible as the answer itself.

Until then, credibility will be allocated in a layer most organizations do not monitor and most users do not inspect. The institutions that respond well will not be those that merely persuade the machine to mention them. They will be the ones that can prove, repeatedly and publicly, that the machine’s description matches the record.

References

All writing →