Complete guide · Generative engine optimization (GEO)
Generative engine optimization (GEO): the complete guide
Generative engine optimization (GEO), stage by stage: how AI engines retrieve and cite pages, which GEO strategies hold up in research, and how to test them.
By Paul Maxwell, founder of AEO HQ
Published · Updated
Generative engine optimization (GEO) is work to get a website's pages retrieved, cited, and described accurately in answers written by AI systems such as ChatGPT, Google's AI Overviews, Perplexity, and Claude. A page must first be indexed: Google's AI features show only pages that are indexed and eligible for a snippet (opens in a new tab). After that, relevance to the question and position among the retrieved sources are the levers that replicate best (opens in a new tab), and rewriting page text for AI engines can make pages harder to retrieve (opens in a new tab).
This guide is the hub for AEO HQ's pages on GEO. It follows a page through the stages of a generative engine, sets out what the research shows at each stage, explains why GEO studies often seem to disagree, and turns the findings into steps. Facts link to their sources, and AEO HQ's own recommendations are labeled as recommendations. The evidence dates from November 2023 to September 2026, and the sources were checked on September 27, 2026.
Guides in this hub
- What is generative engine optimization (GEO)?: the definition, the origin of the term, and what the original study showed.
- GEO vs SEO: what is the difference?: where the two overlap and where they differ.
- Generative engine optimization antipatterns: eleven tactics that fail or break platform rules, each with a test to detect it.
- GEO checklist: each check, how to verify it, and when it passes.
Related pages:
- What replicates in AEO and GEO research: the replication record, claim by claim.
- Answer engine optimization (AEO): the complete guide: the same practice under its other common name.
- AI search optimization: AEO HQ's umbrella model, which treats the work as SEO plus corroboration, specific facts, and measurement.
Scope and definitions
This guide covers unpaid visibility in answers written by AI systems. It does not cover advertising inside those systems or the use of AI tools to write content. The practitioner term answer engine optimization (AEO) describes the same work. Google calls AEO and GEO "both terms you may see used to describe work specifically focused on improving visibility in AI search experiences" (opens in a new tab) (official documentation), and AEO vs GEO compares the labels.
Terms used in this guide:
- Generative engine. The GEO paper's name for a search system that answers queries by "synthesizing information from multiple sources and summarizing them using LLMs" (opens in a new tab). An LLM is a large language model.
- Retrieval. Fetching candidate pages from a search index while an answer is built.
- Reranking. Reordering the retrieved pages, usually with a second model, so that the most relevant ones reach the language model.
- Context. The text the language model reads before it writes, including the retrieved pages. A page's position in the context is its place in that reading order.
- Grounding. Basing an answer on retrieved pages. The general design is called retrieval-augmented generation (RAG). Google says its AI features use RAG, "also known as grounding," to retrieve "relevant, up-to-date web pages from our Search index" (opens in a new tab).
- Query fan-out. Google's term for "a set of concurrent, related queries generated by the model" (opens in a new tab) to find more sources for one question.
- Citation, absorption, and fidelity. A citation is a source that an answer links or names. Absorption is how much a cited page contributes to what the answer says. Fidelity is whether the answer states what the page says. A 2026 study argues that absorption should be measured as an outcome separate from citation (opens in a new tab) (preprint).
- Parametric knowledge. What a model learned in training and can state without searching. This guide is about answers built from retrieved pages.
How a generative engine uses a page
A 2026 review of 45 studies describes GEO as "a stochastic, partially observable pipeline spanning search activation, crawling and indexing, retrieval, reranking and context allocation, citation, prominence, factual absorption, fidelity, and user behavior" (opens in a new tab) (preprint). Stochastic means the output varies by chance from run to run. Partially observable means that site owners do not generally see the retrieved set, the reranking scores, or the model's internal states (opens in a new tab). They see only the answer.
The same review notes that most GEO studies test the stages between context allocation and citation, and far fewer observe crawling, organic retrieval, or user behavior (opens in a new tab). The review's own diagram groups the pipeline into seven stages: activation, crawling and indexing, retrieval, reranking and context, generation and citation, absorption and fidelity, and attention, clicks, and conversions (opens in a new tab). The table follows that grouping. The last column is AEO HQ's reading of what a site owner can influence; the sections after the table give the evidence.
| Stage | What happens | What the research shows | What a site owner can influence |
|---|---|---|---|
| 1. Activation | The engine decides whether to search | Search decisions vary across platforms and models (opens in a new tab) | Little directly |
| 2. Crawling and indexing | A crawler or search partner stores the page | Platforms document indexing as a requirement (opens in a new tab) | Crawler access, indexing, and facts in server-rendered HTML |
| 3. Retrieval | The engine runs one or more searches and pulls candidates | AI answers draw on pages beyond the top organic results (opens in a new tab) | Ranking for the question and its sub-questions; the searcher's words in titles and headings |
| 4. Reranking and context | Candidates are reordered, and the top ones reach the model | Position had larger effects than any rewriting method tested (opens in a new tab) | Retrieval rank; an answer early on the page |
| 5. Generation and citation | The model writes the answer and cites sources | Relevance, a stated price, and a recent date raised citation odds in the lab; formatting alone did not (opens in a new tab) | Specific, consistent facts |
| 6. Absorption and fidelity | The answer uses and restates the page | Many statements are not supported by the sources cited for them (opens in a new tab) | Checking what engines say; clear, quotable facts |
| 7. Attention and action | The reader clicks, or does not | Clicks fall when an AI summary appears (opens in a new tab) | Measurement |
1. Activation: whether the engine searches
An engine can answer from its parametric knowledge or search the web first. Anthropic's documentation says Claude searches when a request depends on information that is "current, changing, or outside its training data," including "information about specific organizations, people, or products that might have changed" (opens in a new tab) (official documentation). In a study of ChatGPT, Claude, Grok, and DeepSeek, web-search decisions "vary substantially across platforms and models" (opens in a new tab) (preprint).
Google's AI features also appear on some searches, not all:
- In a benchmark of 11,500 queries, AI Overviews were generated for 51.5% of representative real-user queries (opens in a new tab) (peer-reviewed).
- In a 40-day crawl of 55,393 trending queries in March and April 2026, AI Overviews appeared on 13.7% of queries overall and on 64.7% of queries phrased as questions (opens in a new tab) (preprint).
- In browsing data from 900 US adults in March 2025, 60% of searches that began with a question word such as "who," "what," "when," or "why" produced an AI summary, against 8% of one- or two-word searches (opens in a new tab) (Pew Research Center).
Our reading: a site cannot make an engine search. It can answer the kinds of questions that the documentation says trigger a search, such as questions about specific companies, current prices, and recent changes.
2. Crawling and indexing: whether the page can be found
Each major engine documents a route into its answers, and each route starts with a crawler or a search index (all official documentation):
- Google. To be eligible as a supporting link in AI Overviews or AI Mode, "a page must be indexed and eligible to be shown in Google Search with a snippet" (opens in a new tab). There are "no additional requirements to appear in AI Overviews or AI Mode" (opens in a new tab).
- Microsoft. "Bing and Copilot search experiences rely on the same core crawling, indexing, and ranking foundation as traditional search" (opens in a new tab).
- OpenAI. Its search crawler is OAI-SearchBot. Sites that opt out of it "will not be shown in ChatGPT search answers, though can still appear as navigational links" (opens in a new tab).
- Anthropic. Its search crawler is Claude-SearchBot. Disabling it "may reduce your site's visibility and accuracy in user search results" (opens in a new tab).
- Perplexity. Its crawler is PerplexityBot, which "is designed to surface and link websites in search results on Perplexity" (opens in a new tab).
Crawlers must also be able to read the text. In Vercel's network data, none of the major AI crawlers rendered JavaScript (opens in a new tab) (hosting-platform measurement, December 2024), so facts that appear only after a script runs may not be seen. How ChatGPT, Gemini, Claude, Perplexity, and Copilot find and cite sources lists each engine's crawlers and controls.
This stage is the gate. A page that is missing from the index an engine searches cannot be retrieved from that index, however well it is written.
3. Retrieval: which pages become candidates
Engines often run several searches for one question. AI Overviews and AI Mode "may use a 'query fan-out' technique — issuing multiple related searches across subtopics and data sources" (opens in a new tab). When ChatGPT search uses outside search providers, it "typically rewrites your query into one or more targeted queries" (opens in a new tab) (both official documentation).
As a result, AI answers draw on pages beyond the top organic results:
- In the 11,500-query benchmark, the sources retrieved by Google Search, AI Overviews, and Gemini were substantially different, with an average Jaccard similarity below 0.2 (opens in a new tab) (peer-reviewed). Jaccard similarity is the number of sources two lists share, divided by the number of distinct sources in either list. It runs from 0 (none shared) to 1 (identical).
- 53% of the domains that AI Overviews consulted were outside the organic top 10 (opens in a new tab) (preprint; 4,706 queries, 2025).
- Nearly 30% of the domains cited in AI Overviews did not appear anywhere in the first page of results (opens in a new tab) (preprint; March–April 2026).
Rank still counts where the lists overlap. When an AI search engine cited a domain that also appeared in traditional results, that domain was most often the first-ranked result: 23.27% of the time in Bing and 14.53% in Google (opens in a new tab) (preprint; 55,936 queries, July–August 2025).
The original GEO study did not test this stage. Its authors wrote that "we didn't evaluate how GEO methods affect search rankings" (opens in a new tab) and expected their text edits to be "less likely to affect search engine rankings" (opens in a new tab). A 2026 study tested the full pipeline: a keyword-based retriever called BM25, a reranker, and GPT-5-mini, over 171,003 web documents (opens in a new tab). It found that optimizing body text alone "consistently degrades visibility across all stages" (opens in a new tab) (peer-reviewed, KDD 2026). The authors traced the drop to word choice: rewrites replaced common words such as "eating" with terms such as "alimentary routines," which reduced the overlap with the query's words (opens in a new tab). With a dense retriever and a hybrid retriever, the trend was the same, which the authors say "confirms that the observed degradation is not specific to BM25" (opens in a new tab).
Structure helped at this stage in the same study. Optimizing titles, meta descriptions, headings, and schema fields raised the retrieval hit rate by 22% (opens in a new tab) with the keyword-based retriever (peer-reviewed; laboratory). Hit rate is the share of queries for which the page appears among the top candidates.
4. Reranking and position in the context
Only the top candidates reach the model. In the 2026 pipeline, 5.8% of target documents dropped from rank 10 to rank 11 during reranking, just missing the cutoff for the generator (opens in a new tab). In the same study, placing the answer early in the document earned higher reranking scores, while restructuring that moved the answer to later paragraphs caused significant rank drops, even when the answer itself stayed intact (opens in a new tab).
In controlled tests, position among the sources had larger effects than any rewriting of the page:
- In a NeurIPS 2025 benchmark with four models, making a document the first one in the model's context produced far greater citation-rank gains than any content-rewriting method (opens in a new tab) (peer-reviewed).
- In 252,000 controlled trials across six models, being listed first rather than second raised the odds of being cited first by a factor of at least 1,795 in every model (opens in a new tab) (peer-reviewed; the authors work for Sprinklr, a software company).
- Language models use the start of their input best: performance "is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades" when it sits in the middle (opens in a new tab) (peer-reviewed).
The NeurIPS benchmark's authors conclude that "traditional SEO strategies, those aiming to improve the ranking of the source in the LLM context, are significantly more effective" (opens in a new tab) than the content methods they tested. Our reading: a page's place among the sources is decided mostly before the model writes, by retrieval rank and reranking. That makes ordinary search optimization, and an answer near the top of the page, the main levers at this stage.
5. Generation and citation: whether the model uses and cites the page
A 2026 controlled study of this step gave six models two anonymized pages at a time, changed one feature between them, and recorded which page was cited first, over 252,000 trials. Topical relevance and list position were the biggest drivers of being cited first; explicit price information and a recent timestamp also helped consistently; completeness and trust cues added smaller gains; and formatting-only edits had little impact (opens in a new tab) (peer-reviewed). The effects were large: an on-topic page beat an off-topic one with odds ratios from 221 to over 10,000, a stated price from 6.26 to over 10,000, and a 2026 rather than a 2019 date from 14.4 to over 10,000 (opens in a new tab). An odds ratio of 10 means ten times the odds. The test supplied both pages directly and never called a search engine (opens in a new tab), so it measures this stage only.
A study with real web paragraphs found the same pattern. Models "rely heavily on the relevance of a website to the query, while largely ignoring stylistic features that humans find important such as whether a text contains scientific references or is written with a neutral tone" (opens in a new tab) (peer-reviewed, ACL 2024).
The original GEO experiments also measured this stage. Adding quotations raised a source's share of the generated answer by about 41% and adding statistics by about 31%, in a simulated engine where the page was already retrieved (opens in a new tab) (peer-reviewed, KDD 2024). A benchmark that measured citation rank instead of share of words found statistically significant gains for conversational-SEO methods in 3 of 54 cases, and adding statistics lowered rank in 19 of 24 (opens in a new tab) (peer-reviewed, NeurIPS 2025). Conversational SEO is that benchmark's name for GEO-style rewriting. The next section explains why the two results differ.
Engines also differ at this stage:
- ChatGPT cited 6.88 sources per prompt on average, Google 12.06, and Perplexity 16.35, but the pages ChatGPT cited had substantially higher average influence on its answers (opens in a new tab) (preprint; 602 prompts; descriptive).
- In tests through the engines' APIs, ChatGPT drew 93.5% of its cited sources from earned media, such as reviews and editorial coverage, for well-known brands and 95.1% for niche brands; Claude drew 87.3% and 86.3% (opens in a new tab) (preprint, 2025).
6. Absorption and fidelity: what the answer takes from the page
Being cited is not the same as being used, or described correctly:
- In the 602-prompt dataset, the pages with the most influence on answers tended to be "longer, more structured, semantically aligned, and richer in extractable evidence such as definitions, numerical facts, comparisons, and procedural steps" (opens in a new tab) (preprint). The authors do not claim that these features cause citation (opens in a new tab).
- In a 2023 audit of four engines, 51.5% of generated sentences were fully supported by their citations, and 74.5% of citations supported their sentence (opens in a new tab) (peer-reviewed; the engines tested have since changed).
- In March and April 2026, 11.0% of 98,020 claims in AI Overviews were not supported by the pages cited for them, with omission the most common failure (opens in a new tab) (preprint).
- In a 2025 study by 22 public service media organizations in 18 countries, almost half of AI assistants' answers about the news had at least one significant issue, and a third showed serious sourcing problems (opens in a new tab) (industry study).
Our reading: check what engines say about your company, not only whether they cite it.
7. Attention, clicks, and conversions
- Google users who saw an AI summary clicked a traditional search result on 8% of visits, against 15% for users who did not see one, and clicked a link inside the summary on 1% of visits (opens in a new tab) (Pew Research Center; 900 US adults, March 2025).
- In a crowd-sourced comparison of search-augmented models, user preferences were influenced by the number of citations "even when the cited content does not directly support the attributed claims" (opens in a new tab) (peer-reviewed, ICLR 2026).
- The 2026 review rates the claim that citation scores "predict clicks, conversions, or revenue" at "very low" confidence (opens in a new tab) (preprint).
How GEO studies measure success, and why they disagree
GEO studies often seem to contradict each other. The main reason is that they measure different stages with different outcomes. The 2026 review argues that the conflicting bodies of evidence "can be reconciled once the stages of the pipeline and the causal estimand are distinguished" (opens in a new tab). An estimand is the exact quantity a study sets out to estimate.
| Study | What was tested | Outcome measured | Main result |
|---|---|---|---|
| Aggarwal et al. (2024), KDD | gpt-3.5-turbo answering from the full text of the top five Google results, with one page rewritten | Share of the answer's words credited to the page, weighted by citation position | Quotations about +41%, statistics about +31% (opens in a new tab) |
| Puerto et al. (2025), NeurIPS | Four models with the documents already in their context, in six domains | Citation rank | Significant gains in 3 of 54 cases (opens in a new tab) |
| Kim et al. (2026), KDD | A retriever, a reranker, and GPT-5-mini over 171,003 web documents | Hit rate and rank at each stage | Body-text rewrites lowered visibility at every stage; structural fields raised the retrieval hit rate by 22% (opens in a new tab) |
| Vishwakarma et al. (2026), SIGIR | Six models given two anonymized pages | Which page is cited first | Topic, position, price, and recency decisive; formatting-only edits had little impact (opens in a new tab) |
| Zhang et al. (2026), preprint | ChatGPT, Google, and Perplexity on 602 prompts | Citations per prompt and influence on the answer | ChatGPT cited fewer sources but relied more on each (opens in a new tab) |
| Watanabe and Nakayashiki (2026), preprint | One live website, comparing changed pages with unchanged pages | ChatGPT referral visits | Referrals to changed pages rose 1.82-fold relative to unchanged pages, but a placebo test gave p = 0.16 (opens in a new tab) |
Three examples show how the choice of outcome changes the answer:
- Share of words versus citation rank. The NeurIPS benchmark's authors note that the original study's main metric does not measure the model's preference, and that "a higher word count does not necessarily correspond to a better citation ranking" (opens in a new tab). They conclude that "the results of both papers on LLM preferences do not contradict each other" (opens in a new tab).
- One stage versus the whole pipeline. AutoGEO is an automated rewriting method. Its own paper reports an average improvement of 35.99% on the original GEO metrics while maintaining answer quality (opens in a new tab) (preprint). In the full retrieve, rerank, and generate pipeline, AutoGEO showed the largest retrieval drop of any method, 22.35 ranks, because its lengthy rewrites diluted keyword density and moved away from the query's vocabulary (opens in a new tab). As the 2026 review puts it, an intervention "may improve one stage while impairing another" (opens in a new tab).
- Laboratory versus live traffic. The only controlled field study we found changed pages on one website and compared them with the site's unchanged pages. Total ChatGPT referrals grew 5.7 times, but untreated pages on the same site grew 3.5 times over the same window (opens in a new tab), so most of the raw growth came from ChatGPT's own growth (preprint; the authors work for Glasp, the company that owns the site (opens in a new tab)).
Most studies in the table are laboratory tests. What replicates in AEO and GEO research grades each claim by the design behind it.
GEO strategies, in order
The steps below are AEO HQ's recommendations, not research findings. Each names the stage it acts on and the strength of the evidence behind it: strong (official documentation, or several independent studies that agree), moderate (consistent evidence that comes mostly from laboratory, correlational, or single studies), or weak (one small or conflicted study).
- List the questions buyers ask, in their words. Stages: retrieval and citation. Collect the questions a buyer might put to an assistant: definitions, comparisons, prices, and "who offers this" questions. Keep the buyer's wording. In the six-model trials, a "keyword gap," in which a page lacks the query's terms, was one of 11 factors that changed the odds of being cited first in at least four of six models (opens in a new tab). Evidence: moderate (laboratory).
- Make every important page crawlable and indexed. Stage: crawling and indexing. Allow each engine's search crawler in robots.txt and in firewall rules, get key pages indexed in Google and Bing, and put the facts you want repeated in the HTML the server sends. Microsoft advises: "Don't hide important answers in tabs or expandable menus: AI systems may not render hidden content" (opens in a new tab) (official guidance). Evidence: strong (official documentation).
- Use the searcher's words in titles, headings, and meta descriptions. Stage: retrieval. This is where structure helped in the 2026 pipeline, and where swapping plain words for technical ones did the most harm (see stage 3). Say each thing once. In the original GEO study, keyword stuffing offered "little to no improvement" (opens in a new tab). Evidence: moderate (one laboratory pipeline, consistent with the six-model trials).
- Rank for the question and its sub-questions, with one page per distinct question. Stages: retrieval and reranking. Because engines fan one question out into several searches, a page can be retrieved for a sub-question the user never typed. Work to rank in Google and Bing for the question and for its likely sub-questions, such as price, scope, and comparisons. Do not build a page for every phrasing. Google says that creating "separate content for every possible variation of how people might search (for example, by focusing on other queries that people have asked, or fan-out queries)" (opens in a new tab) primarily to manipulate rankings or generative AI responses violates its scaled content abuse spam policy (opens in a new tab). Evidence: strong for the policy; moderate for the ranking effect.
- Put the answer in the first paragraph. Stage: reranking. Open each page with a direct answer of two or three sentences that stands on its own. In the 2026 pipeline, early answers earned higher reranking scores (opens in a new tab). An automated method that learned engines' preferences produced the rule "Conclusion First: State the key conclusion at the beginning of the document" (opens in a new tab) (preprint). In the only controlled field study we found, a bundle of changes that included question-form titles and standalone two- to three-sentence answers was followed by a 1.82-fold rise in ChatGPT referrals relative to unchanged pages, an effect the authors call "suggestive, not conclusive" (opens in a new tab). Evidence: moderate.
- State specific, checkable facts in plain text. Stage: citation. Publish prices, what is included, specifications, comparisons, and real dates where a question calls for them. Add numbers because they answer the question, not for their own sake: adding statistics lowered citation rank in 19 of 24 settings (opens in a new tab). Evidence: moderate (laboratory).
- Keep claims consistent, and change dates only when the content changes. Stage: citation. In the six-model trials, internal contradictions were among the factors that lowered the odds of being cited first (opens in a new tab). Dates are a known bias: seven LLM rerankers promoted passages given newer dates, shifting the top 10 forward by up to 4.78 years (opens in a new tab) (peer-reviewed). But in the six-model trials, a recently dated page beat an undated one in only two or three of six models (opens in a new tab). A date changed without a content change misstates the page. Evidence: moderate.
- Earn accurate mentions on other sites. Stages: retrieval and citation. Engines lean on third-party sources (stage 5). Across 75,000 brands, branded web mentions correlated with AI visibility at 0.66 to 0.71, while link metrics showed "very weak correlations" (opens in a new tab) (vendor study; correlation, not cause). Among 112 new startups, a GEO page score showed no correlation with discovery, while referring domains and community presence predicted visibility on Perplexity (opens in a new tab) (preprint based on a master's thesis). Google warns that "seeking inauthentic 'mentions' across the web isn't as helpful as it might seem" (opens in a new tab). Brand mentions and AI recommendations covers this in depth. Evidence: moderate (correlational).
- Check what engines say about your company. Stage: fidelity. Ask each engine your buyers' questions, record what it states about you, and compare that with your own pages. Where an answer is wrong, find the page it cites and correct your own facts, or ask the publisher to correct theirs. Evidence: strong that errors are common (stage 6). We found no study that tests whether correcting a source fixes the answers built on it.
- Measure with repeated runs and a control group. All stages. One answer is one draw. In a daily panel across four engines, cited sources overlapped by only 34% to 42% from one day to the next (opens in a new tab), and the standard error of a brand's detection rate fell below 0.10 at seven runs per prompt (opens in a new tab) (preprint; the first author is also affiliated with Aurora Intelligence). The 2026 review recommends three to five paraphrases per question, repeated runs across several dates, an untreated baseline, and treating answers without search or citations as outcomes rather than discarding them (opens in a new tab). Test the consumer apps as well as APIs: the ChatGPT and Gemini APIs shared only 12.0% and 14.8% of cited domains with their consumer interfaces (opens in a new tab) (preprint). How to measure AI visibility and AEO HQ's methodology describe a full design. Evidence: strong for the method.
What the evidence shows and does not show
The open-weights result on structured HTML comes from a 2026 conference paper that we read in abstract only.
Antipatterns
An antipattern is a practice that looks helpful but fails or backfires. Generative engine optimization antipatterns describes eleven, each with a test. Six of them cost the most:
- Hidden instructions aimed at AI. Google's spam policies cover "attempting to manipulate generative AI responses in Google Search" (opens in a new tab), and Bing warns that content "designed to manipulate or interfere with language models used by Bing or Copilot may result in reduced visibility or removal from search experiences" (opens in a new tab). Instead, state facts in visible text.
- Rewriting the whole site with a language model. Body-text rewrites lowered visibility at every stage of a full pipeline (opens in a new tab). Such rewrites can also be detected: a detector flagged GEO-optimized pages with an F1 score of 0.944 and estimated that 8.90% of 10,095 pages in Google and Gemini results were GEO-optimized (opens in a new tab) (preprint; F1 is an accuracy score from 0 to 1). A defense that demotes GEO-rewritten documents cut attack success from 50.32% to 6.20% (opens in a new tab) (preprint). Instead, edit page by page for readers.
- A page for every variation of a question. Google treats this as a violation of its scaled content abuse spam policy when the aim is to manipulate rankings or generative AI responses (opens in a new tab). Instead, build one page per distinct question.
- Statistics and quotations added as decoration. Adding statistics lowered citation rank in 19 of 24 settings (opens in a new tab). Instead, add a number only when it answers the question, with its source.
- Promising "+40% visibility" or guaranteed citations. The 2026 review calls the 40% figure "a relative maximum on one metric under a specific configuration" (opens in a new tab). Google says third-party tools "can't guarantee performance" (opens in a new tab), and Bing says "GEO does not guarantee grounding or citations in AI experiences" (opens in a new tab). Instead, report measured rates with their conditions.
- Judging by one run. ChatGPT and Google's AI returned the same list of brands less than once in 100 runs, and Claude only slightly more often (opens in a new tab) (vendor study; 2,961 runs, November–December 2025; a co-investigator works for a vendor). Instead, report rates across repeated runs.
Checklist
The GEO checklist gives the full version. The core checks:
| # | Check | How to verify | Pass when | Basis |
|---|---|---|---|---|
| 1 | Search crawlers can fetch key pages | Read /robots.txt; check firewall rules and server logs | OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, and Bingbot receive HTTP 200 on pages you want cited | OpenAI (opens in a new tab); Anthropic (opens in a new tab) |
| 2 | Key pages are indexed | URL Inspection in Google Search Console and Bing Webmaster Tools | Each key page is indexed and eligible for a snippet | Google (opens in a new tab) |
| 3 | Facts are in the server HTML | View the page source without running scripts | Prices, scope, dates, and answers are present | Vercel crawler data (opens in a new tab) |
| 4 | Titles and headings use buyers' words | Compare them with the buyer questions you collected | The main terms appear in plain form | SAGEO Arena (opens in a new tab) |
| 5 | One page per distinct question | Content inventory | No near-duplicate variants | Google (opens in a new tab) |
| 6 | The answer comes first | Read the first paragraph | It answers the page's question on its own | SAGEO Arena (opens in a new tab) |
| 7 | Facts are specific and consistent | Compare the site, profiles, and listings | Prices, scope, and terms match everywhere | Six-model trials (opens in a new tab) |
| 8 | Dates are true | Compare visible dates with the revision history | Dates change only when the content changes | Six-model trials (opens in a new tab) |
| 9 | Engines describe you correctly | Ask each engine buyer questions several times | Statements match your pages | AI Overviews claim audit (opens in a new tab) |
| 10 | Measurement uses repeats and a control | Measurement log | Rates per engine with intervals, compared with untreated pages | Field study (opens in a new tab) |
Frequently asked questions
How does generative engine optimization work?
It works on the steps an AI engine takes before it writes. The engine decides whether to search, retrieves candidate pages from an index, reranks them, and writes an answer from the top few. GEO makes a page easier to retrieve, earlier in the reading order, and easier to use accurately. The 2026 review of 45 studies finds that "already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior" (opens in a new tab).
How do you do generative engine optimization?
Work through the stages in order: make pages crawlable and indexed, use buyers' words, rank for the question and its sub-questions, put the answer first, state specific facts, earn mentions elsewhere, and measure with repeated runs. The ten steps above give the evidence for each.
Why is generative engine optimization important?
Buyers increasingly start with AI assistants. In a March 2026 survey of 1,076 B2B software buyers, 51% said they begin their software research with an AI chatbot more often than with Google, up from 29% in April 2025 (opens in a new tab) (vendor survey). On Google, fewer people click through when an AI summary appears, as stage 7 shows.
How do you measure generative engine optimization?
As rates, not ranks: the share of repeated runs, on each engine, in which answers name or cite the company, with an error range. Two first-party reports help. Search Console's generative AI performance report shows impressions in AI Overviews and AI Mode (opens in a new tab), and Bing Webmaster Tools' AI Performance report shows citations and grounding queries across Microsoft Copilot and Bing's AI summaries (opens in a new tab) (both official documentation).
Is GEO different for ChatGPT, Google, Perplexity, and Claude?
The principles are the same, but the routes differ. Each engine searches a different index, cites a different number and mix of sources (stage 5), and responds differently to page content. In one test, an endorsement-manipulation attack succeeded 0.0% of the time on Claude Sonnet 4.6 and 31.4% of the time on Gemini 3 Flash (opens in a new tab) (preprint). The engine guides cover each one: How to rank in ChatGPT search, How to rank in Google AI Overviews and AI Mode, How to rank in Perplexity, and How to rank in Claude.
What is a generative engine optimization audit?
In AEO HQ's usage, a GEO audit checks a site against the stages in this guide: whether engines' crawlers can fetch and index its pages, whether key facts are in server-rendered text, whether pages answer buyers' questions in their words, and what engines say about the company across repeated runs. There is no standard definition.
Can anyone guarantee GEO results?
No. Google says third-party tools "can't guarantee performance" (opens in a new tab), Bing says "GEO does not guarantee grounding or citations in AI experiences" (opens in a new tab), and OpenAI says of ChatGPT search results that "Placement is not guaranteed" (opens in a new tab).
How long does GEO take to work?
No study has measured how long it takes for a page to start appearing in AI answers. A page must be indexed first, and Google says crawling "can take anywhere from a few days to a few weeks" (opens in a new tab). The 2026 review rates the claim that a white-hat GEO intervention "durably improves organic discoverability across multiple engines" as low confidence (opens in a new tab). White-hat means within platform rules.
Change log
- September 28, 2026: First published.
Next steps
AEO HQ sells this work at fixed, published prices, from a $499 automated audit to $8,995 for an audit, a plan, and technical implementation. The GEO services page lists what each package includes.
Sources
- Google. (2025, December 10). AI features and your website. Google Search Central. https://developers.google.com/search/docs/appearance/ai-features (opens in a new tab)
- Martinez, O. (2026). Optimizing visibility in generative engines: A critical survey of generative engine optimization (2023–2026) (arXiv:2607.14035) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.14035 (opens in a new tab)
- Kim, S., Jeong, W., Kim, S., Lee, S., & Lee, D. (2026). SAGEO Arena: A realistic environment for evaluating search-augmented generative engine optimization. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 2342–2353). ACM. https://doi.org/10.1145/3770855.3818146 (opens in a new tab)
- Google. (2026, July 10). Optimizing your website for generative AI features on Google Search. Google Search Central. https://developers.google.com/search/docs/fundamentals/ai-optimization-guide (opens in a new tab)
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2023). GEO: Generative engine optimization (arXiv:2311.09735) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2311.09735 (opens in a new tab)
- Zhang, K., He, X., & Yao, J. (2026). From citation selection to citation absorption: A measurement framework for generative engine optimization across AI search platforms (arXiv:2604.25707) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.25707 (opens in a new tab)
- Amani, M., Lee, S., Dash, A., El Fraihi, A., Jang, Y., Kirsten, E., Wu, Q., Gummadi, K. P., Gupta, M., Ravichander, A., Zafar, M. B., & Das, S. (2026). Characterizing web search by conversational LLM agents: From search decisions and strategies to results and responses (arXiv:2609.19244) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2609.19244 (opens in a new tab)
- Kirsten, E., Grosse Perdekamp, J., Wu, Q., Upadhyay, M., Gummadi, K. P., & Zafar, M. B. (2025). Characterizing web search in the age of generative AI (arXiv:2510.11560) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2510.11560 (opens in a new tab)
- Puerto, H., Gubri, M., Green, T., Oh, S. J., & Yun, S. (2025). C-SEO Bench: Does conversational SEO work? [Paper presentation]. 39th Conference on Neural Information Processing Systems (NeurIPS 2025), Datasets and Benchmarks Track. https://arxiv.org/abs/2506.11097 (opens in a new tab)
- Vishwakarma, R., Kumar, S., & Jamidar, R. (2026). What gets cited: Competitive GEO in AI answer engines. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 4950–4954). ACM. https://doi.org/10.1145/3805712.3808445 (opens in a new tab)
- Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating verifiability in generative search engines. In Findings of the Association for Computational Linguistics: EMNLP 2023 (pp. 7001–7025). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-emnlp.467 (opens in a new tab)
- Chapekis, A., & Lieb, A. (2025, July 22). Google users are less likely to click on links when an AI summary appears in the results. Pew Research Center. https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/ (opens in a new tab)
- Anthropic. (2026). Web search tool. Claude Platform Docs. Retrieved September 27, 2026, from https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool (opens in a new tab)
- Grossman, R., Liu, S., Chen, M. K., Smith, M., Borcea, C., & Chen, Y. (2026). How generative AI disrupts search: An empirical study of Google Search, Gemini, and AI Overviews. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 448–459). ACM. https://doi.org/10.1145/3805712.3809667 (opens in a new tab)
- Xu, H., Iqbal, U., & Montgomery, J. M. (2026). Measuring Google AI Overviews: Activation, source quality, claim fidelity, and publisher impact (arXiv:2605.14021) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.14021 (opens in a new tab)
- Microsoft Bing. (n.d.). Bing Webmaster Guidelines. Retrieved September 27, 2026, from https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a (opens in a new tab)
- OpenAI. (n.d.). Overview of OpenAI crawlers. OpenAI Developers. Retrieved September 27, 2026, from https://developers.openai.com/api/docs/bots (opens in a new tab)
- Anthropic. (2026, April 7). Does Anthropic crawl data from the web, and how can site owners block the crawler? Claude Help Center. https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler (opens in a new tab)
- Perplexity. (n.d.). Perplexity crawlers. Perplexity Docs. Retrieved September 27, 2026, from https://docs.perplexity.ai/guides/bots (opens in a new tab)
- Zecchini, G., Moore, A. A., Ubl, M., & Siddle, R. (2024, December 17). The rise of the AI crawler. Vercel. https://vercel.com/blog/the-rise-of-the-ai-crawler (opens in a new tab)
- OpenAI. (2026). Searching the web with ChatGPT [Help Center article]. Retrieved September 27, 2026, from https://help.openai.com/en/articles/9237897-chatgpt-search (opens in a new tab)
- Zhang, P., Ye, Q., Peng, Z., Garimella, K., & Tyson, G. (2025). Source coverage and citation bias in LLM-based vs. traditional search engines (arXiv:2512.09483) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2512.09483 (opens in a new tab)
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative engine optimization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5–16). ACM. https://doi.org/10.1145/3637528.3671900 (opens in a new tab)
- Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. https://doi.org/10.1162/tacl_a_00638 (opens in a new tab)
- Wan, A., Wallace, E., & Klein, D. (2024). What evidence do language models find convincing? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 7468–7484). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.403 (opens in a new tab)
- Chen, M., Wang, X., Chen, K., & Koudas, N. (2025). Generative engine optimization: How to dominate AI search (arXiv:2509.08919) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2509.08919 (opens in a new tab)
- Fletcher, J., & Verckist, D. (2025, October). News integrity in AI assistants: An international PSM study. European Broadcasting Union & BBC. https://www.ebu.ch/research/open/report/news-integrity-in-ai-assistants (opens in a new tab)
- Miroyan, M., Wu, T.-H., King, L., Li, T., Pan, J., Hu, X., Chiang, W.-L., Angelopoulos, A. N., Darrell, T., Norouzi, N., & Gonzalez, J. E. (2026). Search Arena: Analyzing search-augmented LLMs [Paper presentation]. Fourteenth International Conference on Learning Representations (ICLR 2026). https://arxiv.org/abs/2506.05334 (opens in a new tab)
- Watanabe, K., & Nakayashiki, K. (2026). Disentangling answer engine optimization from platform growth: A log-based natural experiment on ChatGPT referral traffic (arXiv:2606.04362) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.04362 (opens in a new tab)
- Wu, Y., Zhong, S., Kim, Y., & Xiong, C. (2025). What generative search engines like and how to optimize web content cooperatively (arXiv:2510.11438) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2510.11438 (opens in a new tab)
- Madhavan, K. (2025, October 8). Optimizing your content for inclusion in AI search answers. Microsoft Advertising Blog. https://about.ads.microsoft.com/en/blog/post/october-2025/optimizing-your-content-for-inclusion-in-ai-search-answers (opens in a new tab)
- Fang, H., Tao, S., Chen, N., Chang, K.-X., & Sakai, T. (2025). Do large language models favor recent content? A study on recency bias in LLM-based reranking. In Proceedings of the 2025 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region (pp. 85–94). ACM. https://doi.org/10.1145/3767695.3769493 (opens in a new tab)
- Linehan, L. (2025, December 12). Top brand visibility factors in ChatGPT, AI Mode, and AI Overviews (75k brands studied). Ahrefs. https://ahrefs.com/blog/ai-brand-visibility-correlations/ (opens in a new tab)
- Sharma, A. P. (2026). The discovery gap: How Product Hunt startups vanish in LLM organic discovery queries (arXiv:2601.00912) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2601.00912 (opens in a new tab)
- Schulte, J., Bleeker, M., & Kaufmann, P. (2026). Don't measure once: Measuring visibility in AI search (GEO) (arXiv:2604.07585) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.07585 (opens in a new tab)
- Uberti-Bona Marin, L. G., Bertaglia, T., Astante, G., Rijsbosch, B., van Dijck, G., Hannák, A., Spanakis, G., & Kollnig, K. (2026). "If I had to buy just ONE: Galaxy S26 Ultra": Auditing AI-generated product recommendations (arXiv:2609.18729) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2609.18729 (opens in a new tab)
- Mehta, M., Sharma, R., Kalluru, V., & Kotwal, A. (2026). Beyond the blue link: Empirical evaluation of generative engine optimization in stochastic retrieval systems. In Proceedings of the 37th ACM Conference on Hypertext (pp. 276–283). ACM. https://doi.org/10.1145/3800935.3830864 (opens in a new tab)
- Linehan, L. (2026, May 11). We tracked 1,885 pages adding schema. AI citations barely moved. Ahrefs. https://ahrefs.com/blog/schema-ai-citations/ (opens in a new tab)
- Nestaas, F., Debenedetti, E., & Tramèr, F. (2025). Adversarial search engine optimization for large language models. In The Thirteenth International Conference on Learning Representations (ICLR 2025). https://proceedings.iclr.cc/paper_files/paper/2025/hash/0f12b3c36a781120c4f60e90e855868d-Abstract-Conference.html (opens in a new tab)
- Chen, Y., Ren, Z., Laakom, F., Li, Y., Guo, D., & Schmidhuber, J. (2026). How much can we trust LLM search agents? Measuring endorsement vulnerability to web content manipulation (arXiv:2606.16821) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.16821 (opens in a new tab)
- Google. (2026, August 28). Spam policies for Google web search. Google Search Central. https://developers.google.com/search/docs/essentials/spam-policies (opens in a new tab)
- Chu, J., Leng, Y., Li, M., Shen, Y., Shen, X., & Zhang, Y. (2026). GEO-Flag: Detecting and measuring GEO-optimized web content (arXiv:2608.16824) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2608.16824 (opens in a new tab)
- Li, H., Shao, Y., Lin, X., Guan, Z., Zhou, M., & Shi, J. (2026). When optimization becomes manipulation: Defending generative search against malicious generative engine optimization (arXiv:2609.02964) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2609.02964 (opens in a new tab)
- Google. (2026, June 5). Google Search's guidance on using third-party SEO tools, services, and advice. Google Search Central. https://developers.google.com/search/docs/fundamentals/third-party-seo (opens in a new tab)
- Fishkin, R. (2026, January 28). NEW research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility. SparkToro. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/ (opens in a new tab)
- G2. (2026, April 15). New G2 research: Half of B2B software buyers now start their research with AI chatbots [Press release]. PR Newswire. https://www.prnewswire.com/news-releases/new-g2-research-half-of-b2b-software-buyers-now-start-their-research-with-ai-chatbots-302742807.html (opens in a new tab)
- Google. (2026). Generative AI performance report [Search Console Help]. Retrieved September 27, 2026, from https://support.google.com/webmasters/answer/16984139 (opens in a new tab)
- Madhavan, K., Merchant, M., Canel, F., & Nigam, S. (2026, February 10). Introducing AI Performance in Bing Webmaster Tools (public preview). Bing Webmaster Blog. https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview (opens in a new tab)
- Google. (2025, December 10). Ask Google to recrawl your URLs. Google Search Central. https://developers.google.com/search/docs/crawling-indexing/ask-google-to-recrawl (opens in a new tab)
How to cite this page
Maxwell, P. (2026). Generative engine optimization (GEO): the complete guide. AEO HQ. Last updated September 28, 2026. https://www.aeohq.ai/generative-engine-optimization
Pages in Generative engine optimization (GEO)
Definition
What is generative engine optimization (GEO)?
Generative engine optimization (GEO) is work to get content cited in AI-generated answers. Where the term came from and what the evidence supports.
Comparison
GEO vs SEO: what is the difference?
GEO vs SEO: goals, where results appear, what is measured, and what drives inclusion. AI answers draw on search indexes, so the two overlap heavily.
Checklist
GEO checklist
A GEO checklist of 18 checks for retrieval, reranking, citation, accuracy, and testing, with the evidence for each and the tactics that did not replicate.
Antipatterns
Generative engine optimization antipatterns
Eleven generative engine optimization tactics that fail or break platform rules, what the research shows about each, and a test to detect it.