Same Question, Same Google.
Three Different Sets of Sources.
A peer-reviewed SIGIR 2026 study built an 11,500-query benchmark and compared the sources Google Search, Google’s AI Overviews, and Google’s Gemini returned for the same query. For the 7,439 queries where all three surfaces returned sources, the overlap is small enough to be the finding: 18 percent, on average, between the AI Overview and the results page it sits on. Run the same query twice under the study’s conditions and the AI Overview’s two source lists overlap at 0.66. Change an apostrophe, an abbreviation, or a question mark — not the meaning — and the AI Overview drops 28.99 percent below its own repeat-run baseline, on the study’s rank-sensitive consistency measure. The authors’ conclusion, in their own words: the results “challenge the effectiveness”
of generative engine optimization.
The experiment reads like a control test Google might have run on itself. Researchers at the New Jersey Institute of Technology, Nanyang Technological University, and Indiana University assembled an 11,500-query benchmark. It spans representative real-user queries, retail product queries and the comparisons and questions built from them, complex informational questions, debate-style queries, localized “near me” searches. Then they put identical queries to three of Google’s own surfaces across the same two December days: classic Search, the AI Overview that now sits above it, and Gemini, the standalone assistant. They compared what each cited as its evidence, across the queries where all three returned sources. One company. Three retrieval surfaces. Three materially different source lists.
The source lists barely touch. “On average, only 18% of the sources returned by either the AIO or traditional SERP will be retrieved by both search engines,” the paper reports; across the three pairings, overlap runs from 0.11 to 0.18 by Jaccard similarity, the standard measure of how much two sets share. And the pairing you would bet on produced the sentence peer review left standing: “Surprisingly, despite AIO being built with a lightweight Gemini model, the source lists retrieved by AIO and Gemini are the least similar.”
Hold the three rooms apart, because the distinction carries the finding. Google Search is the ranked page of links. The AI Overview is the machine-written answer placed above those links. Gemini is the assistant in its own room. Ranking in classic Search establishes visibility in one room. It does not establish source visibility in the AI Overview or in Gemini — the study found that each surface assembled materially different evidence for the same query.
Then the team tested repetition, on a stratified sample of one hundred benchmark queries: the same query, the same device configuration, the same location, run twice. Classic Search’s two source lists overlapped at 0.78. The AI Overview’s overlapped at 0.66 — under the study’s own conditions, identical inputs still produced materially different source sets.
The punctuation tax
The cosmetic-edit test is where instability becomes a tax. The researchers made minor edits to 200 queries: contracting “what is” to “what’s,” abbreviating “United States” to “U.S.,” adding or removing a question mark from queries that are, in the paper’s words, “clearly questions regardless of punctuation.” Their summary finding, verbatim: “even when the query’s underlying intent is unchanged, cosmetic query differences result in different sources retrieved, and subsequently different generated text.” The study scores this on a rank-sensitive overlap measure — RBO — against each engine’s own repeat-run baseline. Classic Search declined 13.95 percent under the edits. The AI Overview fell to 0.49, a 28.99 percent decline. In the study’s test, an apostrophe could change which sources represent you.
Across surfaces, the sources diverge. For identical queries, Google Search, the AI Overview, and Gemini cite source sets that overlap between 0.11 and 0.18 — on average, only 18 percent of the sources returned by either the AI Overview or classic Search appear in both.
Across runs, the sources churn. Run the identical query twice, on a stratified 100-query sample under identical conditions: the AI Overview’s two source lists overlap at 0.66; classic Search’s at 0.78 — both by Jaccard similarity.
Across punctuation, the sources re-roll. An apostrophe, an abbreviation, or a question mark drops the AI Overview 28.99 percent below its own repeat-run baseline on the study’s rank-sensitive measure (RBO 0.69 to 0.49); classic Search declines 13.95 percent.
The principal benchmark responses were collected December 7–8, 2025, using SerpAPI for Search and the AI Overview, and the Gemini API with Google Search grounding for Gemini 2.5 Flash. The decimals are a dated photograph. The instability they document is the finding that lasts.
The conclusion the authors wrote themselves
The discussion section names the industry consequence in the field’s own vocabulary: “Our results challenge the effectiveness of GEO techniques” — generative engine optimization, the emerging playbook for tuning content to be cited by AI answers. The paper challenges it by name, with citations to its founding papers. Then the paper says who pays: “Our results show that generative search will benefit niche content providers at the expense of more popular, established ones.” Read plainly, the demographic on the losing end of that sentence is anyone who spent twenty years becoming the established answer — an editorial inference, not the paper’s own claim.
In Agentics terms, these measurements are evidence of Contextual Ambiguity at the output level: the source set varies by surface, by run, and under minor changes in wording. The paper does not identify Google’s internal selection criteria, establish why the variation occurs, or test any intervention — Agentics’ included. The diagnosis is editorial interpretation, grounded in the observed output pattern — and so is what follows: optimizing one page for one surface is tuning for a target the measurements above show doesn’t hold still. The controllable response is density: reinforced, machine-legible signals distributed across the surfaces a system may sample — a strategy the study never tests, a recognition no one can guarantee. The strategic inference is narrower: when source selection varies, depending on one page, one ranking, or one surface leaves more ways to disappear. Optimization is tactical. Installation is strategic. This paper is the measurement behind the sentence.
Edition No. 046 filed the trust half of this file: the date can be true and the signal still ignored. This edition files the selection half: the surfaces doing the believing don’t agree with one another — or with themselves — about what to cite. The tax has no invoice. It is collected in absences: the run that didn’t sample you, the surface that never saw you, the apostrophe that re-rolled the evidence. What a business controls is not the roll. It is how many ways the record resolves to you when the system rolls again.
Grossman, Liu, Chen, Smith, Borcea, and Chen, “How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews,” Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’26), Melbourne, July 20–24, 2026; DOI 10.1145/3805712.3809667. Read via the open-access record at arXiv:2604.27790 (April 30, 2026). Quoted figures: the key-findings summary and Table 2 (cross-surface similarity), Table 4 (repeat-run similarity), the “Consistency in the Presence of Cosmetic Query Edits” analysis within Section 4, and the Section 6 discussion. Proceedings status is cited from the paper’s own conference header; the ACM Digital Library page blocks automated retrieval.
The study’s 11,500-query benchmark is public at github.com/rag24/AIO. Its acknowledgments disclose National Science Foundation and National Institutes of Health translational-science funding, a university chair, and a foundation grant; they disclose no funding from any search or AI platform.