Research · Research and reference
What replicates in AEO and GEO research
Which AEO and GEO research findings hold up across independent studies, which did not replicate, and which rest on one study. Graded, 2023–2026.
By Paul Maxwell, founder of AEO HQ
Published · Updated
In research on answer engine optimization (AEO) and generative engine optimization (GEO) published from 2023 to 2026, only a few findings hold up across independent studies. A page must be retrieved before its wording can matter (opens in a new tab). Relevance to the question and position among the sources are the levers that replicate best (opens in a new tab), and answers vary substantially from run to run (opens in a new tab). The best-known tactic, adding quotations and statistics to a page, did not hold up in a later benchmark that measured citation rank (opens in a new tab).
This page is AEO HQ's replication record for the field. For each common claim, it sets out what the first study found, which later studies tested it, and whether the result held. It is part of AEO HQ's research. It is a structured review of the studies AEO HQ collected for its research program, not a systematic review. Sources were checked on September 27, 2026.
Scope and method
What "replicates" means on this page
- Retrieval. The step in which an engine fetches candidate pages from a search index while it builds an answer.
- Replication. A later study tests a finding again and reaches the same result. A direct replication repeats the original method. A conceptual replication tests the same claim with a different method, system, or data set.
- Independent. The later study comes from a different research group.
- Stage. A generative engine works in stages, from deciding whether to search to writing and citing the answer. The complete guide to generative engine optimization describes them. Two studies test the same claim only if they measure the same stage and a comparable outcome.
- Laboratory and field. In a laboratory study, researchers control what the model sees, for example by supplying the pages directly. A field study observes live products or live traffic.
- Preprint. A paper posted publicly before, or without, peer review. "Accepted" means accepted for publication but not yet published.
Each claim below gets one of six verdicts. The verdicts are AEO HQ's judgments, and the rules are stated so that readers can check them:
| Verdict | Meaning |
|---|---|
| Replicated | At least two independent studies with different data or systems agree, or platforms document the fact directly |
| Partly replicated | Studies agree in direction, but only in laboratory settings, or one study is supported by evidence of a different kind |
| Did not replicate | A later, stricter test did not reproduce the result |
| Contested | Credible studies disagree |
| Not yet replicated | One study, one vendor's data, or correlations only |
| Untested | We found no study |
Verdicts also depend on the design behind a study. A 2026 critical review of 45 GEO studies proposes a scale for grading a claim by the design that supports it (opens in a new tab) (preprint), paraphrased here:
| Level | Design | What it can show |
|---|---|---|
| A | Randomized field trial, or a strong quasi-experiment with logs and controls | A causal effect on clicks, traffic, or conversions in the setting studied |
| B | Live commercial engines, with repeated runs, paraphrases, and several dates | How visibility is distributed, and how engines differ |
| C | A commercial engine given manually supplied URLs or files | An effect after retrieval, not on crawling or organic retrieval |
| D | A reproducible pipeline with retrieval, reranking, and generation | Stage-by-stage effects inside the testbed |
| E | A fixed context, a synthetic ranker, or a model used as judge | A possible mechanism, given that context |
Sources and conflicts of interest
The studies were published between November 2023 and September 2026, plus a few earlier studies of how language models read their input. Each finding links to its source, labeled as peer-reviewed, preprint, vendor study, journalism, or official documentation. When a study's authors work for a company that sells a related product, we say so. Where we could read only a study's abstract, we say so. AEO HQ's evidence standard describes the source tiers and checks.
AEO HQ sells AEO services, including audits, so it has an interest in which tactics look effective. Readers should weigh our verdicts with that in mind.
The record at a glance
Findings that replicate
1. Pages must be retrievable before anything else matters
Verdict: replicated. Platforms document it, and the research is consistent with it.
- What platforms say (official documentation). To be eligible as a supporting link in Google's AI Overviews or AI Mode, a page "must be indexed and eligible to be shown in Google Search with a snippet" (opens in a new tab). OpenAI's search crawler is OAI-SearchBot, and sites that opt out of it "will not be shown in ChatGPT search answers, though can still appear as navigational links" (opens in a new tab). Bing says that "Bing and Copilot search experiences rely on the same core crawling, indexing, and ranking foundation as traditional search" (opens in a new tab).
- What research shows. In a full retrieve, rerank, and generate pipeline, 5.8% of target documents dropped from rank 10 to rank 11 during reranking and so missed the generator's input cutoff (opens in a new tab) (peer-reviewed, KDD 2026). The 2026 critical review finds that experiments "provide strong evidence that a document that has already been retrieved can alter an answer" (opens in a new tab), but "establish far less often that a page will be retrieved organically" (opens in a new tab).
Our reading: work on a page's wording can matter only for pages that an engine can retrieve. Most GEO experiments start after retrieval, so their results do not show that a page will be found.
2. Position among the sources has large effects
Verdict: replicated in laboratory settings, by independent groups.
- In C-SEO Bench, a NeurIPS 2025 benchmark with four models and six domains, making a document the first one in the model's context produced far greater citation-rank gains than any content-rewriting method (opens in a new tab) (peer-reviewed).
- In 252,000 two-source trials across six models, being listed first rather than second raised the odds of being cited first by a factor of at least 1,795 in every model (opens in a new tab) (peer-reviewed; the authors work for Sprinklr, a software company).
- In a study of how language models use long inputs, performance "is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts" (opens in a new tab) (peer-reviewed).
- The 2026 critical review rates "Query–document relevance and context position are major determinants" at high confidence (opens in a new tab).
Limit: site owners cannot set a page's position directly. It follows from retrieval rank and reranking. We found no test on a live engine that isolates position.
3. Relevance to the question drives which source is used
Verdict: replicated in laboratory settings.
- An ACL 2024 study of real web paragraphs found that models "rely heavily on the relevance of a website to the query, while largely ignoring stylistic features that humans find important such as whether a text contains scientific references or is written with a neutral tone" (opens in a new tab) (peer-reviewed).
- In the six-model trials, an on-topic page beat an off-topic one with odds ratios from 221 to over 10,000 (opens in a new tab). An odds ratio of 221 means 221 times the odds. A "keyword gap," where a page lacks the query's terms, was among the factors that lowered the odds of being cited first in at least four of six models (opens in a new tab).
- In SAGEO Arena, the KDD 2026 pipeline study, rewrites that replaced common words with rarer ones, such as "alimentary routines" for "eating," reduced overlap with the query's words and lowered retrieval (opens in a new tab). Structural fields, which are "dense with query-relevant terms," raised the retrieval hit rate by 22% (opens in a new tab) (peer-reviewed).
Three groups, three designs: counterfactual edits to real paragraphs, a factorial test of synthetic pages, and a retrieval pipeline over real web documents. The review rates relevance as a major determinant at high confidence (see finding 2).
4. Answers vary from run to run
Verdict: replicated on live engines, across groups, engines, and dates.
- ChatGPT and Google's AI returned the same list of brands less than once in 100 runs, and Claude only slightly more often (opens in a new tab) (vendor study; 2,961 runs, November–December 2025; a co-investigator works for a vendor).
- When the same queries were run two months apart, only 18% of the web pages AI Overviews used were the same, against 45% for organic search (opens in a new tab) (preprint; 4,706 queries, 2025).
- Across four engines, the cited sources overlapped by only 34% to 42% from one day to the next (opens in a new tab) (preprint; January–March 2026; the first author is also affiliated with Aurora Intelligence).
- In the 11,500-query benchmark, AI Overviews were "less consistent when processing two runs of the same query" and more sensitive to minor edits of the query (opens in a new tab) (peer-reviewed).
- The products ChatGPT recommended often changed across repeated requests (opens in a new tab) (preprint; abstract read).
The 2026 review summarizes the commercial audits it covers as showing "low source overlap, substantial run-to-run variability, and persistent fidelity gaps" (opens in a new tab). A single answer or screenshot is one sample, not a measurement of AI visibility.
5. AI answers cite beyond the top organic results
Verdict: replicated in direction. The size of the gap varies with the method and the date.
- In a benchmark of 11,500 queries, the sources retrieved by Google Search, AI Overviews, and Gemini were substantially different, with an average Jaccard similarity below 0.2 (opens in a new tab) (peer-reviewed). Jaccard similarity is the number of sources two lists share, divided by the number of distinct sources in either list.
- 53% of the domains AI Overviews consulted were outside the organic top 10 (opens in a new tab) (preprint; 2025).
- Nearly 30% of the domains cited in AI Overviews did not appear anywhere on the first page of results (opens in a new tab) (preprint; March–April 2026).
- 37% of the domains cited by LLM-based search engines were unique to them (opens in a new tab) (preprint; 55,936 queries, July–August 2025).
The size varies. One SEO-tool vendor measured 76.10% of AI Overview-cited pages in Google's top 10 in July 2025 (opens in a new tab) and 37.9% in March 2026, after improving its method for parsing citations (opens in a new tab) (vendor studies). Where AI and traditional results do overlap, the domains AI search engines cited were most often the first-ranked result: 23.27% of the time in Bing and 14.53% in Google (opens in a new tab).
Whether AI answers favor popular sites is contested. In the US benchmark, traditional Google search was "significantly more likely to retrieve information from popular or institutional websites in government or education" (opens in a new tab). A study of queries in 243 countries found instead that AI search surfaced "significantly fewer long tail information sources" (opens in a new tab) than traditional search (preprint).
6. Citations often fail to support the claims attached to them
Verdict: replicated across general, news, and health questions.
- In a 2023 audit of four engines, 51.5% of generated sentences were fully supported by their citations, and 74.5% of citations supported their sentence (opens in a new tab) (peer-reviewed; the engines tested have since changed).
- For medical questions, between 50% and 90% of LLM responses were not fully supported, and sometimes contradicted, by the sources they cited; for GPT-4o with web search, about 30% of individual statements were unsupported (opens in a new tab) (peer-reviewed; abstract read).
- In March and April 2026, 11.0% of 98,020 claims in AI Overviews were not supported by the pages cited for them (opens in a new tab) (preprint).
- In a 2025 evaluation by 22 public service media organizations in 18 countries, almost half of AI assistants' answers about the news had at least one significant issue, and a third showed serious sourcing problems (opens in a new tab) (industry study).
- When eight AI search tools were asked to identify the source of news excerpts, they collectively answered more than 60% of 1,600 queries incorrectly (opens in a new tab) (journalism research).
Users do not reliably catch the gap. In a crowd-sourced comparison, user preferences were influenced by the number of citations "even when the cited content does not directly support the attributed claims" (opens in a new tab) (peer-reviewed, ICLR 2026). We found no study of claims about B2B vendors in particular.
7. Page content can manipulate AI answers
Verdict: replicated in laboratory tests and in some tests on live engines. How well an attack works depends on the model.
- Adding a "strategic text sequence" to a product's page significantly increased its chance of being the top recommendation (opens in a new tab) (preprint).
- A jailbreak-style attack reliably promoted low-ranked products, and the attacks transferred to perplexity.ai (opens in a new tab) (peer-reviewed).
- Crafted website content made production search engines, Bing and Perplexity, promote an attacker's products and discredit competitors (opens in a new tab) (peer-reviewed).
- One polluted page among the retrieved results fooled recommendations up to 27% of the time across 12 models, and polluting the top three results raised that to 73.8% (opens in a new tab) (accepted; abstract read).
- Hidden instructions made ChatGPT search's review of a product "always entirely positive," even when the page carried negative reviews (opens in a new tab) (journalism, December 2024).
Results depend on the model. Across 13 model back ends, attack success ranged from 0.0% on Claude Sonnet 4.6 to 31.4% on Gemini 3 Flash (opens in a new tab) (preprint; abstract read). Defenses are arriving: spotlighting cut attack success from over 50% to under 2% (opens in a new tab) (preprint), a reranker that demotes GEO-rewritten documents cut average attack success from 50.32% to 6.20% (opens in a new tab) (preprint), and a lightweight guard cut attack success by 47.6% in relative terms (opens in a new tab) (accepted). The review rates "Retrieved documents constitute a genuine attack surface" at high confidence (opens in a new tab).
These tactics also break platform rules. Google's spam policies cover attempts to manipulate generative AI responses (opens in a new tab), and Bing warns that content "designed to manipulate or interfere with language models used by Bing or Copilot may result in reduced visibility or removal from search experiences" (opens in a new tab) (official documentation). Generative engine optimization antipatterns describes the variants.
8. Well-known brands start ahead
Verdict: replicated in direction, across different designs.
- In tests of three commercial models, well-known brands were recommended 100% of the time when all products had the same specifications, but the advantage disappeared when a competitor had a rating edge of less than 0.1 stars (opens in a new tab) (preprint).
- In production tracking data, global household-name brands appeared in 73% of unbranded category answers on the first run, established mid-size brands in 44%, and the smallest tier in 11% (opens in a new tab) (preprint; 102 brands, March–May 2026; the author co-founded and holds equity in Ranqo, the platform studied (opens in a new tab)).
- Among 112 recently launched startups, models recognized the products by name 99.4% and 94.3% of the time, but surfaced them in only 3.32% and 8.29% of discovery-style queries (opens in a new tab) (preprint based on a master's thesis).
- LLMs differed significantly in how much they weighted product name, document content, and context position (opens in a new tab), and recommenders were more easily fooled by fake products when they lacked stable prior knowledge of the real ones (opens in a new tab).
Our reading: a new or small company depends more on what engines retrieve than on what models already know about it.
9. Gains shrink when competitors copy a tactic
Verdict: partly replicated. Three groups found it, all in simulations.
- In C-SEO Bench, the average gain per adopter fell as more players adopted the same method, converging toward zero at full adoption (opens in a new tab) (peer-reviewed).
- In tests of manipulative content, attacks led to a prisoner's dilemma: every party had an incentive to attack, but attacks collectively degraded the model's outputs for everyone (opens in a new tab) (peer-reviewed).
- When all brands adopted the same optimization strategy, the individual payoff fell from +0.802 to +0.007 in the authors' payoff measure (opens in a new tab) (preprint).
The review rates this claim at moderate confidence, noting that it holds "in tested multi-actor settings" (opens in a new tab). We found no study that follows a real market over time.
Findings that partly replicate
10. Specific facts, such as prices, raise citation odds
Verdict: partly replicated in laboratory settings.
- In the six-model trials, a stated price raised the odds of being cited first in all six models, with odds ratios from 6.26 to over 10,000 (opens in a new tab), and specifications, comparisons, and evidence were among the factors significant in at least four of six models (opens in a new tab) (peer-reviewed; the pages were synthetic, with anonymized brands).
- In a separate test, a visible difference was enough to flip a model's choice: the well-known brand's advantage disappeared when a competitor had a rating edge of less than 0.1 stars (opens in a new tab) (preprint).
- In a descriptive study of 602 prompts, the pages with the most influence on answers were richer in "extractable evidence such as definitions, numerical facts, comparisons, and procedural steps" (opens in a new tab) (preprint; the authors make no causal claim).
The review rates "Extractable evidence and suitable structure often facilitate use" at moderate confidence (opens in a new tab). We found no field test that isolates specific facts. These results concern facts that answer the question. Numbers added for their own sake did not help (finding 13).
11. An answer early on the page helps
Verdict: partly replicated. One laboratory pipeline, one learned rule, and one suggestive field study point the same way.
- In SAGEO Arena, placing the answer early in the document earned higher reranking scores, while restructuring that moved the answer to later paragraphs caused significant rank drops, even when the answer itself stayed intact (opens in a new tab) (peer-reviewed).
- An automated method that learned engines' preferences from their explanations produced the rule "Conclusion First: State the key conclusion at the beginning of the document" (opens in a new tab) (preprint).
- In the only controlled field study we found, a bundle of changes that included question-form titles and standalone two- to three-sentence answers was followed by a 1.82-fold rise in ChatGPT referrals relative to unchanged pages on the same site (95% CI 1.31 to 2.54) (opens in a new tab). A placebo test gave p = 0.16, so the authors call the effect "suggestive, not conclusive" (opens in a new tab) (preprint; the authors work for Glasp, which runs the site studied). A 95% CI is a confidence interval: the range of values consistent with the data.
The bundle in the field study also changed URLs and added pages, so its effect cannot be credited to the early answers alone.
12. Recent dates help
Verdict: partly replicated in laboratory settings.
- In the six-model trials, a 2026 rather than a 2019 date raised the odds of being cited first in all six models, with odds ratios from 14.4 to over 10,000 (opens in a new tab), but a recently dated page beat an undated one in only two or three of six models (opens in a new tab).
- Seven LLM rerankers promoted passages given newer dates, shifting the top 10 forward by up to 4.78 years and reversing up to 25% of preferences between equally relevant passages; larger models reduced the effect but none eliminated it (opens in a new tab) (peer-reviewed).
We found no study that tests whether a visible, true update date changes citations on a live engine. A date changed without a content change misstates the page.
Findings that did not replicate
13. Adding quotations, statistics, or citations ("GEO +40%")
Verdict: did not replicate as a general effect.
The original result. The 2024 GEO paper found that adding quotations raised a source's share of the generated answer by about 41% and adding statistics by about 31%, in a simulated engine where the page was already retrieved (opens in a new tab) (peer-reviewed, KDD 2024). Its abstract says GEO "can boost visibility by up to 40% in generative engine responses" (opens in a new tab). Three details narrow the result:
- A language model wrote the rewrites, and the paper's own caption says the methods raised visibility "without adding any substantial new information" (opens in a new tab).
- The engine was gpt-3.5-turbo, answering from the full text of the top five Google results (opens in a new tab).
- In the paper's analysis where all sources were optimized, citing sources lowered the visibility of the top-ranked source by 30.3% while raising that of the fifth-ranked source by 115.1% (opens in a new tab).
The later tests.
- C-SEO Bench measured citation rank across four newer models and six domains. It found statistically significant gains for conversational-SEO methods in 3 of 54 cases, and adding statistics lowered rank in 19 of 24 (opens in a new tab). No method was effective for question answering or on Claude 3.5 Haiku, and 26 of 30 product-recommendation cases on Claude 3.5 Haiku were significantly negative (opens in a new tab) (peer-reviewed).
- SAGEO Arena ran the full pipeline over 171,003 web documents. Optimizing body text alone "consistently degrades visibility across all stages" (opens in a new tab), and the result held with dense and hybrid retrievers, a different reranker, and Claude Sonnet 4.6 as the generator (opens in a new tab) (peer-reviewed).
- In an e-commerce testbed, 11 of 15 hand-written rewriting heuristics did worse than a simple rewriting prompt when GPT-4.1 did the rewriting (opens in a new tab) (preprint; product listings ranked by language models).
- Real paragraphs with scientific references did not gain: models largely ignored "whether a text contains scientific references" (opens in a new tab) (peer-reviewed).
Reconciliation. The C-SEO Bench authors say the original study's word-count metric does not measure the model's preference, and conclude that "the results of both papers on LLM preferences do not contradict each other" (opens in a new tab). The 2026 review calls the 40% figure "a relative maximum on one metric under a specific configuration" (opens in a new tab) and rejects it as a general claim. Our reading: the original result describes a page that is already one of five sources, judged by its share of the answer's words, with gpt-3.5-turbo as the engine. It says nothing about retrieval, citation rank, or current models.
14. A persuasive or authoritative tone helps
Verdict: did not replicate. Three studies found no consistent effect.
- The original GEO paper: for a more persuasive and authoritative tone, "we find no significant improvement" (opens in a new tab).
- The ACL 2024 study: models largely ignored stylistic features such as a neutral tone (opens in a new tab).
- The six-model trials: neutral versus overly promotional wording reached significance in only two or three of six models (opens in a new tab).
One distinction held in a single study: confident rather than hedged wording was among the factors significant in at least four of six models (opens in a new tab). Plain, confident statements of fact may help; sales language did not.
15. A rewriting method that wins its own benchmark works in a full pipeline
Verdict: did not replicate in the one full-pipeline test available.
- AutoGEO learns engines' preferences and rewrites pages to match them. Its own paper reports an average improvement of 35.99% on the original GEO metrics while maintaining answer quality (opens in a new tab) (preprint; simulated engines).
- In SAGEO Arena, AutoGEO showed the largest retrieval drop of any method, 22.35 ranks, because its lengthy rewrites diluted keyword density and moved away from the query's vocabulary (opens in a new tab) (peer-reviewed).
- Inside its own testbed, an e-commerce method that searched for better rewriting prompts improved rank in 63 of 75 combinations of prompt and ranking model (opens in a new tab) (preprint).
Rewritten pages can also be detected and demoted. A detector identified GEO-optimized pages with an F1 score of 0.944 and estimated that 8.90% of 10,095 pages in Google and Gemini results were GEO-optimized, rising to 16.36% of pages modified in 2026 (opens in a new tab) (preprint; abstract read; F1 is an accuracy score from 0 to 1). The review rates "Systematically optimized or learned methods often outperform fixed heuristics in controlled benchmarks" at moderate confidence (opens in a new tab), and the claim that "A white-hat GEO intervention durably improves organic discoverability across multiple engines" at low confidence (opens in a new tab). White-hat means within platform rules.
Findings that are contested
16. Formatting and structure raise citations
Verdict: contested. The answer seems to depend on the stage.
- No consistent effect once the model has the full text. In the six-model trials, formatting-only edits, structured versus dense text and organized versus scattered text, had little impact (opens in a new tab) (peer-reviewed).
- Slightly negative, descriptively. Pages formatted as questions and answers had a mean influence of 0.0947 against 0.1005 for other pages, a 5.74% relative difference (opens in a new tab) (preprint; descriptive).
- Small. In the e-commerce testbed, bulleted lists had a small significant effect of +0.23 rank positions, but length and bullets together explained only 0.5% of the variance in rank changes (opens in a new tab) (preprint).
- Positive at retrieval. In SAGEO Arena, optimizing titles, meta descriptions, headings, and schema fields raised the retrieval hit rate by 22% (opens in a new tab) with a keyword-based retriever (peer-reviewed).
- Positive in one test on open-weights models. A Hypertext 2026 paper reports that structured HTML yielded 2.6 times higher citation rates and 4.0 times higher extraction fidelity than equivalent unstructured content (opens in a new tab) (peer-reviewed; abstract read).
Platform guidance points both ways. Microsoft says assistants like Copilot break content into "smaller, structured pieces" (opens in a new tab), while Google says "There's no requirement to break your content into tiny pieces for AI to better understand it" (opens in a new tab) (official documentation). Our reading: structure seems to matter most where pages are split and matched to queries, and least once a model has the full text in front of it.
Findings that rest on one study, one vendor, or correlations
17. Mentions on other sites raise AI visibility
Verdict: not yet replicated as a cause. The correlations are consistent, but we found no controlled test.
- Across 75,000 brands, branded web mentions correlated with AI visibility at 0.66 to 0.71, while link metrics showed "very weak correlations" (opens in a new tab) (vendor study).
- Among 112 startups, a GEO page score showed no correlation with discovery, while referring domains (r = +0.319) and community presence (r = +0.395) predicted visibility on Perplexity (opens in a new tab) (preprint).
- In API tests with ranking-style prompts, ChatGPT drew 93.5% of its cited sources from earned media, such as reviews and editorial coverage, for well-known brands and 95.1% for niche brands (opens in a new tab) (preprint).
- In production tracking data, 2.9% of 149,912 citations pointed to the tracked brand's own domain (opens in a new tab) (preprint; author has a conflict of interest, see finding 8).
These are correlations and descriptive shares. A brand's size can drive both its mentions and its visibility. Brand mentions and AI recommendations covers the practical side.
18. Schema markup raises AI citations
Verdict: not supported.
- In a matched study of 1,885 pages that added JSON-LD markup, compared with about 4,000 control pages, AI Overview citations fell 4.6% (a small, statistically significant decline), while the changes in AI Mode (+2.4%) and ChatGPT (+2.2%) were indistinguishable from zero (opens in a new tab). The pages studied were already cited heavily by AI (opens in a new tab) (vendor study).
- An observational study of B2B SaaS pages found that metadata and freshness, semantic HTML, and structured data showed the strongest associations with citation (opens in a new tab) (preprint; observational).
- Google says you don't need "machine readable files, AI text files, markup, or Markdown" to appear in Google Search, including its generative AI features (opens in a new tab) (official documentation).
We found no controlled study showing that schema markup raises AI citations.
19. llms.txt files help
Verdict: not supported.
- In server logs from 137,210 domains, 97% of published llms.txt files received zero traffic in May 2026 (opens in a new tab) (vendor study).
- In C-SEO Bench, "LLM guidance," a Markdown summary inspired by the llms.txt standard and added to the top of a document, was one of only two methods with any significant gains, and only in the retail and video game domains (opens in a new tab) (peer-reviewed).
- Google says Google Search "doesn't use" AI text files such as llms.txt (opens in a new tab) (official documentation).
20. GEO raises clicks, traffic, or sales
Verdict: not yet replicated.
- The only controlled field study we found reported a suggestive 1.82-fold rise in ChatGPT referrals relative to unchanged pages, with a placebo test at p = 0.16 (opens in a new tab) (see finding 11).
- A study of 973 e-commerce websites found that ChatGPT referrals converted above paid social but below all other traditional channels (opens in a new tab) (peer-reviewed; descriptive). It measured referral quality, not the effect of GEO.
- The review rates the claim that citation scores "predict clicks, conversions, or revenue" at "very low" confidence (opens in a new tab).
Open questions
We found no study that answers these questions:
- How long a page takes to be cited after it is published or changed.
- Whether any technique within platform rules has a lasting effect across engines. The 2026 review found no reviewed technique with "a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior" (opens in a new tab).
- Whether the findings hold for questions about B2B vendors and services. In our reading of the studies above, most use news, health, general-knowledge, or consumer-product queries.
- Whether correcting a source fixes a wrong AI answer about a company.
- How an assistant's memory of a user's past conversations changes its recommendations.
Why studies disagree
- They measure different stages. The 2026 review argues that conflicting findings "can be reconciled once the stages of the pipeline and the causal estimand are distinguished" (opens in a new tab). An estimand is the exact quantity a study sets out to estimate.
- They use different outcomes. A share of the answer's words, a citation rank, a retrieval hit rate, and a referral count can move in different directions. As the C-SEO Bench authors put it, "A higher word count does not necessarily correspond to a better citation ranking" (opens in a new tab).
- Simulations are not live products. Most GEO studies test the stages between context allocation and citation, and far fewer observe crawling, organic retrieval, or user behavior (opens in a new tab). The SAGEO Arena authors observed that "findings from existing benchmarks do not readily generalize to realistic settings" (opens in a new tab).
- Models change. The original GEO study used gpt-3.5-turbo; later studies used models released through 2026, which responded very differently (finding 7). The review rates "Commercial engines differ from one another and vary over time" at high confidence (opens in a new tab).
- APIs are not the apps people use. The ChatGPT and Gemini APIs shared only 12.0% and 14.8% of cited domains with their consumer interfaces (opens in a new tab), and the authors conclude that "neither isolated responses nor API observations can be assumed to represent the commercial advice consumers encounter" (opens in a new tab) (preprint).
- Test pages are often synthetic. In the six-model trials, GPT-4o anonymized the brands and generated the paired page variants (opens in a new tab).
- Many authors sell related products. Authors of studies above work for or with Sprinklr, Glasp, Ranqo, Aurora Intelligence, and Ahrefs, as noted at each finding.
What this means in practice
These points are AEO HQ's recommendations, based on the record above:
- Work in stage order. Make pages crawlable, indexed, and ranked before editing their wording.
- Use the levers with the most support. Relevance to the buyer's question, rank for the question and its sub-questions, an answer early on the page, and specific facts that answer the question.
- Treat other tactics as unproven. That includes rewriting heuristics, formatting alone, schema markup, and llms.txt.
- Ask for the conditions behind any claimed lift. Ask for the stage, the outcome measured, the engine and model, the dates, and the control group. A lift with none of these is an anecdote.
- Measure with repeated runs and untreated pages. The complete guide to generative engine optimization sets out a method.
Limitations of this review
- It is not a systematic review. There was no pre-registered search protocol, and the studies were drawn from the sources collected for AEO HQ's research tracks.
- Many sources are preprints, and several were read in abstract only.
- Engines and models change often, so findings can date quickly.
- The verdicts are AEO HQ's judgments. Others applying the same rules could reach different ones.
- AEO HQ sells AEO services.
Change log
- September 28, 2026: First published.
Next steps
AEO HQ's Instant AEO Audit ($499) is an automated check of a site's crawler access, key pages, and structured data, with a small sample of one AI model's answers. How AEO HQ measures AI visibility explains what the audit measures and what it does not.
Sources
- Martinez, O. (2026). Optimizing visibility in generative engines: A critical survey of generative engine optimization (2023–2026) (arXiv:2607.14035) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.14035 (opens in a new tab)
- Puerto, H., Gubri, M., Green, T., Oh, S. J., & Yun, S. (2025). C-SEO Bench: Does conversational SEO work? [Paper presentation]. 39th Conference on Neural Information Processing Systems (NeurIPS 2025), Datasets and Benchmarks Track. https://arxiv.org/abs/2506.11097 (opens in a new tab)
- Google. (2025, December 10). AI features and your website. Google Search Central. https://developers.google.com/search/docs/appearance/ai-features (opens in a new tab)
- Kim, S., Jeong, W., Kim, S., Lee, S., & Lee, D. (2026). SAGEO Arena: A realistic environment for evaluating search-augmented generative engine optimization. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 2342–2353). ACM. https://doi.org/10.1145/3770855.3818146 (opens in a new tab)
- Vishwakarma, R., Kumar, S., & Jamidar, R. (2026). What gets cited: Competitive GEO in AI answer engines. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 4950–4954). ACM. https://doi.org/10.1145/3805712.3808445 (opens in a new tab)
- Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. https://doi.org/10.1162/tacl_a_00638 (opens in a new tab)
- Wan, A., Wallace, E., & Klein, D. (2024). What evidence do language models find convincing? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 7468–7484). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.403 (opens in a new tab)
- Fishkin, R. (2026, January 28). NEW research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility. SparkToro. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/ (opens in a new tab)
- Kirsten, E., Grosse Perdekamp, J., Wu, Q., Upadhyay, M., Gummadi, K. P., & Zafar, M. B. (2025). Characterizing web search in the age of generative AI (arXiv:2510.11560) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2510.11560 (opens in a new tab)
- Schulte, J., Bleeker, M., & Kaufmann, P. (2026). Don't measure once: Measuring visibility in AI search (GEO) (arXiv:2604.07585) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.07585 (opens in a new tab)
- Grossman, R., Liu, S., Chen, M. K., Smith, M., Borcea, C., & Chen, Y. (2026). How generative AI disrupts search: An empirical study of Google Search, Gemini, and AI Overviews. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 448–459). ACM. https://doi.org/10.1145/3805712.3809667 (opens in a new tab)
- Xu, H., Iqbal, U., & Montgomery, J. M. (2026). Measuring Google AI Overviews: Activation, source quality, claim fidelity, and publisher impact (arXiv:2605.14021) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.14021 (opens in a new tab)
- Zhang, P., Ye, Q., Peng, Z., Garimella, K., & Tyson, G. (2025). Source coverage and citation bias in LLM-based vs. traditional search engines (arXiv:2512.09483) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2512.09483 (opens in a new tab)
- Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating verifiability in generative search engines. In Findings of the Association for Computational Linguistics: EMNLP 2023 (pp. 7001–7025). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-emnlp.467 (opens in a new tab)
- Wu, K., Wu, E., Wei, K., Zhang, A., Casasola, A., Nguyen, T., Riantawan, S., Shi, P., Ho, D., & Zou, J. (2025). An automated framework for assessing how well LLMs cite relevant medical references. Nature Communications, 16, 3615. https://doi.org/10.1038/s41467-025-58551-6 (opens in a new tab)
- Nestaas, F., Debenedetti, E., & Tramèr, F. (2025). Adversarial search engine optimization for large language models. In The Thirteenth International Conference on Learning Representations (ICLR 2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/0f12b3c36a781120c4f60e90e855868d-Abstract-Conference.html (opens in a new tab)
- Pfrommer, S., Bai, Y., Gautam, T., & Sojoudi, S. (2024). Ranking manipulation for conversational search engines. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 9523–9552). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.534 (opens in a new tab)
- Luo, M., & Chen, L. (2026). One polluted page is enough: Evaluating web content pollution in LLM recommenders (arXiv:2606.13610; accepted to Findings of EMNLP 2026) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.13610 (opens in a new tab)
- Chu, X., & Hou, Y. (2026). Incumbent advantage: Brand bias and cognitive manipulation dynamics in LLM recommendation systems (arXiv:2606.17443) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.17443 (opens in a new tab)
- Kumar, P. (2026). Generative engine optimization at scale: Measuring brand visibility across AI search engines (arXiv:2606.20065) [Preprint]. arXiv. https://arxiv.org/abs/2606.20065 (opens in a new tab)
- Sharma, A. P. (2026). The discovery gap: How Product Hunt startups vanish in LLM organic discovery queries (arXiv:2601.00912) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2601.00912 (opens in a new tab)
- Watanabe, K., & Nakayashiki, K. (2026). Disentangling answer engine optimization from platform growth: A log-based natural experiment on ChatGPT referral traffic (arXiv:2606.04362) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.04362 (opens in a new tab)
- Fang, H., Tao, S., Chen, N., Chang, K.-X., & Sakai, T. (2025). Do large language models favor recent content? A study on recency bias in LLM-based reranking. In Proceedings of the 2025 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region (pp. 85–94). ACM. https://doi.org/10.1145/3767695.3769493 (opens in a new tab)
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative engine optimization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5–16). ACM. https://doi.org/10.1145/3637528.3671900 (opens in a new tab)
- Wu, Y., Zhong, S., Kim, Y., & Xiong, C. (2025). What generative search engines like and how to optimize web content cooperatively (arXiv:2510.11438) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2510.11438 (opens in a new tab)
- Mehta, M., Sharma, R., Kalluru, V., & Kotwal, A. (2026). Beyond the blue link: Empirical evaluation of generative engine optimization in stochastic retrieval systems. In Proceedings of the 37th ACM Conference on Hypertext (pp. 276–283). ACM. https://doi.org/10.1145/3800935.3830864 (opens in a new tab)
- Linehan, L. (2025, December 12). Top brand visibility factors in ChatGPT, AI Mode, and AI Overviews (75k brands studied). Ahrefs. https://ahrefs.com/blog/ai-brand-visibility-correlations/ (opens in a new tab)
- Linehan, L. (2026, May 11). We tracked 1,885 pages adding schema. AI citations barely moved. Ahrefs. https://ahrefs.com/blog/schema-ai-citations/ (opens in a new tab)
- Linehan, L. (2026, June 15). We analyzed 137K sites: 97% of llms.txt files never get read. Ahrefs. https://ahrefs.com/blog/llmstxt-study/ (opens in a new tab)
- OpenAI. (n.d.). Overview of OpenAI crawlers. OpenAI Developers. Retrieved September 27, 2026, from https://developers.openai.com/api/docs/bots (opens in a new tab)
- Microsoft Bing. (n.d.). Bing Webmaster Guidelines. Retrieved September 27, 2026, from https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a (opens in a new tab)
- Uberti-Bona Marin, L. G., Bertaglia, T., Astante, G., Rijsbosch, B., van Dijck, G., Hannák, A., Spanakis, G., & Kollnig, K. (2026). "If I had to buy just ONE: Galaxy S26 Ultra": Auditing AI-generated product recommendations (arXiv:2609.18729) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2609.18729 (opens in a new tab)
- Linehan, L. (2025, July 21). 76% of AI Overview citations pull from the top 10. Ahrefs. https://ahrefs.com/blog/search-rankings-ai-citations (opens in a new tab)
- Linehan, L. (2026, March 2). Update: 38% of AI Overview citations pull from the top 10. Ahrefs. https://ahrefs.com/blog/ai-overview-citations-top-10/ (opens in a new tab)
- Aral, S., Li, H., & Zuo, R. (2026). The rise of AI search: Implications for information markets and human judgement at scale (arXiv:2602.13415) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2602.13415 (opens in a new tab)
- Fletcher, J., & Verckist, D. (2025, October). News integrity in AI assistants: An international PSM study. European Broadcasting Union & BBC. https://www.ebu.ch/research/open/report/news-integrity-in-ai-assistants (opens in a new tab)
- Jaźwińska, K., & Chandrasekar, A. (2025, March 6). AI search has a citation problem. Columbia Journalism Review, Tow Center for Digital Journalism. https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php (opens in a new tab)
- Miroyan, M., Wu, T.-H., King, L., Li, T., Pan, J., Hu, X., Chiang, W.-L., Angelopoulos, A. N., Darrell, T., Norouzi, N., & Gonzalez, J. E. (2026). Search Arena: Analyzing search-augmented LLMs [Paper presentation]. Fourteenth International Conference on Learning Representations (ICLR 2026). https://arxiv.org/abs/2506.05334 (opens in a new tab)
- Kumar, A., & Lakkaraju, H. (2024). Manipulating large language models to increase product visibility (arXiv:2404.07981) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2404.07981 (opens in a new tab)
- Evershed, N. (2024, December 24). ChatGPT search tool vulnerable to manipulation and deception, tests show. The Guardian. https://www.theguardian.com/technology/2024/dec/24/chatgpt-search-tool-vulnerable-to-manipulation-and-deception-tests-show (opens in a new tab)
- Chen, Y., Ren, Z., Laakom, F., Li, Y., Guo, D., & Schmidhuber, J. (2026). How much can we trust LLM search agents? Measuring endorsement vulnerability to web content manipulation (arXiv:2606.16821) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.16821 (opens in a new tab)
- Hines, K., Lopez, G., Hall, M., Zarfati, F., Zunger, Y., & Kiciman, E. (2024). Defending against indirect prompt injection attacks with spotlighting (arXiv:2403.14720) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2403.14720 (opens in a new tab)
- Li, H., Shao, Y., Lin, X., Guan, Z., Zhou, M., & Shi, J. (2026). When optimization becomes manipulation: Defending generative search against malicious generative engine optimization (arXiv:2609.02964) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2609.02964 (opens in a new tab)
- Zheng, B., Zhao, Z., & Yang, W. (2026). Counter-GEO-Bench: Evaluating defenses against information-distorting generative engine optimization (arXiv:2609.02316; accepted to EMNLP 2026) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2609.02316 (opens in a new tab)
- Google. (2026, August 28). Spam policies for Google web search. Google Search Central. https://developers.google.com/search/docs/essentials/spam-policies (opens in a new tab)
- Zhang, K., He, X., & Yao, J. (2026). From citation selection to citation absorption: A measurement framework for generative engine optimization across AI search platforms (arXiv:2604.25707) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.25707 (opens in a new tab)
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2023). GEO: Generative engine optimization (arXiv:2311.09735) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2311.09735 (opens in a new tab)
- Bagga, P. S., Farias, V. F., Korkotashvili, T., Peng, T., & Wu, Y. (2025). E-GEO: A testbed for generative engine optimization in e-commerce (arXiv:2511.20867) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2511.20867 (opens in a new tab)
- Chu, J., Leng, Y., Li, M., Shen, Y., Shen, X., & Zhang, Y. (2026). GEO-Flag: Detecting and measuring GEO-optimized web content (arXiv:2608.16824) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.16824 (opens in a new tab)
- Madhavan, K. (2025, October 8). Optimizing your content for inclusion in AI search answers. Microsoft Advertising Blog. https://about.ads.microsoft.com/en/blog/post/october-2025/optimizing-your-content-for-inclusion-in-ai-search-answers (opens in a new tab)
- Google. (2026, July 10). Optimizing your website for generative AI features on Google Search. Google Search Central. https://developers.google.com/search/docs/fundamentals/ai-optimization-guide (opens in a new tab)
- Chen, M., Wang, X., Chen, K., & Koudas, N. (2025). Generative engine optimization: How to dominate AI search (arXiv:2509.08919) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2509.08919 (opens in a new tab)
- Kumar, A., & Palkhouski, L. (2025). AI answer engine citation behavior: An empirical analysis of the GEO-16 framework (arXiv:2509.10762) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2509.10762 (opens in a new tab)
- Kaiser, M., & Schulze, C. (2026). ChatGPT referrals to e-commerce websites: How do LLMs compare against traditional channels? Marketing Science, 45(4), 699–715. https://doi.org/10.1287/mksc.2025.0489 (opens in a new tab)
How to cite this page
Maxwell, P. (2026). What replicates in AEO and GEO research. AEO HQ. Last updated September 28, 2026. https://www.aeohq.ai/articles/what-replicates
More in Research and reference
Reference
AEO glossary: answer engine and AI search terms
Plain definitions of the terms used in answer engine optimization (AEO) and AI search, from AI Overviews to robots.txt, with a source for every fact.
Reference
AI search statistics (2026)
Sourced AI search statistics for 2026: chatbot use, AI Overviews and clicks, AI referral traffic, cited sources, answer accuracy, and B2B buyers.