Skip to content
AEO HQ

Complete guide · Ranking in AI assistants

How ChatGPT, Gemini, Claude, Perplexity, and Copilot find and cite sources

The search index, crawlers, inclusion controls, and site-owner reports behind ChatGPT, Gemini, Claude, Perplexity, and Copilot, from official docs.

By , founder of AEO HQ

Published · Updated

When an AI assistant cites web pages, it has usually found them in a search index while writing its answer. Google's AI Overviews and AI Mode use Google's index (opens in a new tab). Copilot uses Bing (opens in a new tab). ChatGPT uses OpenAI's own crawler (opens in a new tab) plus search partners that include Microsoft (opens in a new tab). For Claude, Anthropic lists Brave Search as a web-search subprocessor (opens in a new tab) and runs its own search crawler (opens in a new tab). Perplexity runs its own crawler (opens in a new tab). A page missing from the relevant index cannot be cited through search.

This page is a reference. For each assistant, a table lists the index or search provider behind its answers, its crawlers and what each one is for, the controls that decide whether a page can be included, and the reports each platform gives site owners. The platform facts come from each company's own documentation, checked on September 27, 2026. Research on how the assistants differ when they cite follows the tables. How to act on these facts is covered in AI search optimization.

Guides in this hub

Each assistant has its own guide:

Scope and definitions

This page covers the organic route into AI answers: how assistants find, fetch, and cite public web pages. It does not cover ads or shopping feeds. The work of improving that route goes by several names, including answer engine optimization (AEO) and generative engine optimization (GEO).

  • Retrieval. The step in which an assistant runs searches and pulls pages into its working context before it writes. Answers based on those pages are said to be grounded; grounding is the general term.
  • Parametric knowledge. What a model learned in training. An assistant that does not search answers from this alone.
  • Search index. The store of crawled pages that a search system retrieves from.
  • Crawler. A program that fetches web pages. Each crawler announces a name, its user agent, which site owners use in robots.txt rules.
  • Training crawler, search crawler, user-triggered fetcher. The three purposes the documentation distinguishes. A training crawler collects content that may be used to train models. A search crawler builds the index used for answers. A user-triggered fetcher opens a page because a user's request needs it.
  • robots.txt. A file at the root of a site that tells crawlers what they may fetch. The rules are standardized in RFC 9309 (opens in a new tab).
  • Query fan-out. Turning one question into several searches. Google describes it as "issuing multiple related searches across subtopics and data sources (opens in a new tab)."
  • First-party reporting. Data that the platform itself gives site owners about their pages.

The five assistants at a glance

"None documented" means no such report appears in the company's documentation as of September 27, 2026.

Two routes into an answer

Assistants answer from training data or from retrieved pages, and they do not search for every question. In a US clickstream panel, ChatGPT ran a web search on 34.5% of queries in February 2026, down from 46% in late 2024 (opens in a new tab) (vendor study). Claude's documentation says it searches "when the request depends on information that is current, changing, or outside its training data" (opens in a new tab), including information about specific organizations, people, or products.

Training data has a cutoff: Anthropic lists June 2026 as the reliable-knowledge cutoff of its newest Claude models (opens in a new tab). It follows that a company that is new, or that recently changed its prices or services, reaches these assistants mainly through retrieval. The tables below describe that route.

ChatGPT (OpenAI)

ItemWhat OpenAI's documentation says
Where answers get web pagesOAI-SearchBot "is used to surface websites in search results in ChatGPT's search features" (opens in a new tab). ChatGPT search "sometimes partners with other search providers"; the help center names Microsoft and Shopify (opens in a new tab). For Enterprise and Edu accounts, ChatGPT "may share disassociated search queries with the Bing search engine" (opens in a new tab).
Search crawlerOAI-SearchBot. Sites that opt out "will not be shown in ChatGPT search answers, though can still appear as navigational links." (opens in a new tab)
Training crawlerGPTBot, which crawls content "that may be used in training our generative AI foundation models" (opens in a new tab). Each setting is independent of the others (opens in a new tab), so a site can allow OAI-SearchBot and block GPTBot.
User-triggered fetcherChatGPT-User, which may visit a page when a user asks a question. Because these visits are user-initiated, "robots.txt rules may not apply," and ChatGPT-User "is not used to determine whether content may appear in Search." (opens in a new tab)
Other crawlerOAI-AdsBot, which validates the safety of pages submitted as ChatGPT ads (opens in a new tab).
What controls inclusionrobots.txt rules for OAI-SearchBot, which take about 24 hours to apply (opens in a new tab), and a host or content delivery network that allows traffic from OpenAI's published searchbot IP addresses (opens in a new tab). If OpenAI obtains a disallowed page's URL from a third-party search provider, it "may surface just the link and page title in ChatGPT Atlas." (opens in a new tab)
First-party reportingNo citation report for site owners is documented. ChatGPT adds utm_source=chatgpt.com to referral URLs (opens in a new tab).
How results are chosenChatGPT "rewrites your query into one or more targeted queries," ranks results "using multiple factors," and "Placement is not guaranteed." (opens in a new tab)

What research adds. ChatGPT's domain concentration differed significantly from Google's but not from Bing's, which the authors read as reliance on Bing (opens in a new tab) (preprint; July–August 2025). ChatGPT drew 93.5–95.1% of the domains it cited from earned media, such as reviews and editorial coverage (opens in a new tab) (preprint; API versions, mid-2025). Pages ChatGPT cited averaged 958 days old, the newest of the assistants measured (opens in a new tab) (vendor study).

Guides: ranking in ChatGPT's search results and getting recommended by ChatGPT.

Google: AI Overviews, AI Mode, and Gemini

ItemWhat Google's documentation says
Where answers get web pagesAI Overviews and AI Mode use Google Search: to appear as a supporting link, "a page must be indexed and eligible to be shown in Google Search with a snippet," and no special optimization is required (opens in a new tab). Both may use query fan-out (opens in a new tab). The Gemini app can ground answers in content Google crawls (opens in a new tab).
Search crawlerGooglebot, the crawler for Google Search (opens in a new tab). Google's generative AI features are "rooted in our core Search ranking and quality systems." (opens in a new tab)
Training controlGoogle-Extended, a robots.txt token with no user agent of its own: "Crawling is done with existing Google user agent strings." (opens in a new tab) The Google-Extended robots.txt token controls whether content is used for Gemini training and for grounding in Gemini Apps and Vertex AI; it does not affect inclusion in Google Search (opens in a new tab).
User-triggered fetchersGoogle-Agent, used by "agents hosted on Google infrastructure" to act on a user's request, and Google-GeminiNotebook. These fetchers "generally ignore robots.txt rules." (opens in a new tab)
What controls inclusionIndexing and snippet eligibility. Page-level controls include nosnippet, data-nosnippet, max-snippet, and noindex (opens in a new tab). A site must also be included in Search Console's "Search generative AI" setting, which is on by default and was rolled out to all websites on August 31, 2026 (opens in a new tab). For the Gemini app, the Google-Extended token.
First-party reportingThe Search Console Generative AI performance report shows impressions in AI Overviews and AI Mode, but not clicks (opens in a new tab). Traffic from AI features is included in the Performance report's "Web" search type (opens in a new tab). Google Analytics 4 counts AI Overviews and AI Mode visits as Organic Search, and Gemini visits in its AI Assistant channel (opens in a new tab).
Files and markup"You don't need to create new machine readable files, AI text files, or markup to appear in these features," (opens in a new tab) and markup should match the visible text on the page (opens in a new tab).

What research adds. In a December 2025 benchmark of more than 11,000 queries, AI Overviews and organic results shared 18% of their sources, and AI Overviews and Gemini shared 11% (opens in a new tab) (peer-reviewed). The same study found that 21 publishers blocking Google-Extended were never cited by Gemini and were less likely to be retrieved by AI Overviews (opens in a new tab), although Google says the token does not affect Search. The link is an association, not a tested cause. 38% of Gemini responses cited no website at all (opens in a new tab) (preprint; July–August 2025). AI Overviews showed no preference for newer pages than organic results (opens in a new tab) (vendor study).

Guides: ranking in Google's AI answers and ranking in Gemini.

Microsoft Copilot

ItemWhat Microsoft's documentation says
Where answers get web pagesBing. Copilot "generates a search query that it sends to the Bing search service" (opens in a new tab), and Bing and Copilot "rely on the same core crawling, indexing, and ranking foundation as traditional search." (opens in a new tab)
Search crawlerBingbot, Bing's crawler (opens in a new tab).
Training crawlerNo separate training crawler is described in the Bing and Copilot documents reviewed for this page.
User-triggered fetcherNone described in those documents.
What controls inclusionBeing indexed by Bing, which asks site owners to use IndexNow and XML sitemaps with accurate lastmod values (opens in a new tab). NOARCHIVE "prevents content from being used in Copilot responses and grounding results," and NOCACHE limits Copilot to the URL, title, and snippet (opens in a new tab). In the work versions of Copilot, IT admins can turn web search on or off for their users (opens in a new tab).
First-party reportingBing Webmaster Tools' AI Performance report shows total citations, average cited pages, grounding queries, and page-level citations across Microsoft Copilot, Bing's AI summaries, and "select partner integrations," with no click data (opens in a new tab). A June 2026 update added Citation Share, "the percentage of citations attributed to your site out of all citations shown across all sites for that same grounding query." (opens in a new tab)
How citations appearCopilot shows users the search queries it generated and sent to Bing, as "web search query citations." (opens in a new tab)
How content is readMicrosoft says its systems parse pages into "smaller, structured pieces," and advises against hiding answers in tabs, expandable menus, PDFs, or images (opens in a new tab).

Indexing notes. IndexNow shares each submission with Bing, Yandex, Naver, Seznam.cz, Yep, and Amazon, but not Google, and "Submitting a URL does not guarantee immediate indexing." (opens in a new tab) In 2023, Bing said 12% of new URLs clicked in its results were first discovered through IndexNow (opens in a new tab) (platform figure). Because Copilot searches Bing, a page Bing has not indexed cannot be cited in Copilot's web answers.

Guide: ranking in Microsoft Copilot.

Claude (Anthropic)

ItemWhat Anthropic's documentation says
Where answers get web pagesAnthropic lists Brave Search ("All products") and TurboPuffer (all products except Claude for Government) as web-search subprocessors (opens in a new tab). It also runs Claude-SearchBot, which analyzes web content "to enhance the relevance and accuracy of search responses." (opens in a new tab) The Claude API's web search documentation does not name a provider (opens in a new tab).
Search crawlerClaude-SearchBot. Blocking it "prevents our system from indexing your content for search optimization, which may reduce your site's visibility and accuracy in user search results." (opens in a new tab)
Training crawlerClaudeBot, which collects web content that could contribute to model training; blocking it signals that future content should be excluded from training (opens in a new tab).
User-triggered fetcherClaude-User. Blocking it "prevents our system from retrieving your content in response to a user query, which may reduce your site's visibility for user-directed web search." (opens in a new tab)
What controls inclusionAnthropic's bots honor robots.txt and the Crawl-delay extension, and their IP addresses are published at claude.com/crawling/bots.json (opens in a new tab). For Brave, its crawler has no distinct user agent and follows Googlebot's robots.txt rules, and its submit-URL form is documented only for re-fetching pages to delist them (opens in a new tab).
First-party reportingNone documented. Traffic referred by Claude's native app carries no Referer header (opens in a new tab), which hides its source in analytics.
When Claude searchesWhen a request depends on information that is current, changing, or outside its training data, including about specific organizations, people, or products; results carry a URL, title, and page age (opens in a new tab).

What research adds. Claude drew 86–87% of the domains it cited from earned media, and its set of cited domains was the most stable across languages (opens in a new tab) (preprint; API version, mid-2025). How much of Claude's web search comes from Brave's index and how much from Anthropic's own crawl is not documented.

Guide: ranking in Claude.

Perplexity

ItemWhat Perplexity's documentation says
Where answers get web pagesIts own index, built by PerplexityBot, which is "designed to surface and link websites in search results on Perplexity." (opens in a new tab) The crawler documentation describes no third-party index.
Search crawlerPerplexityBot, which site owners should allow in robots.txt to appear in results (opens in a new tab).
Training crawlerNone documented. PerplexityBot "is not used to crawl content for AI foundation models." (opens in a new tab)
User-triggered fetcherPerplexity-User, which fetches pages when users ask questions. Because users initiate these requests, "this fetcher generally ignores robots.txt rules." (opens in a new tab)
What controls inclusionrobots.txt rules for PerplexityBot; its IP addresses are published at perplexity.com/perplexitybot.json (opens in a new tab).
First-party reportingNone documented.

What research adds. Perplexity's Sonar model consulted about 14 pages per query, against about 9 for AI Overviews and Gemini and about 4 for OpenAI's GPT-4o search model (opens in a new tab) (preprint; 4,706 queries, 2025). Perplexity drew 17.5–23.8% of the domains it cited from social platforms, the most of the engines tested (opens in a new tab) (preprint). Among 112 recently launched startups, those with more referring domains and more Reddit presence were more likely to be found by Perplexity (opens in a new tab) (preprint, thesis-based). Perplexity's documentation does not mention Bing or any other third-party index.

Guide: ranking in Perplexity.

How the assistants differ when they cite

These findings come from independent studies, mostly run through each product's API. They describe tendencies, not fixed rules.

FindingResultEvidence
Share of cited domains that were earned media (reviews, comparisons, editorial)ChatGPT 93.5–95.1%, Claude 86–87%, Perplexity 67–73%, Gemini 63–66% (opens in a new tab)Preprint; ranking-style prompts, mid-2025
Pages consulted per queryPerplexity Sonar about 14, AI Overviews and Gemini about 9, OpenAI's GPT-4o search model about 4 (opens in a new tab)Preprint; 4,706 queries, 2025
Responses with no cited websiteGemini 38% (opens in a new tab)Preprint; 55,936 queries, July–August 2025
News answers with significant sourcing problemsGemini 72%, ChatGPT 24%, Copilot 15%, Perplexity 15% (opens in a new tab)Public broadcasters' study; 2,709 answers, May–June 2025
Cited domains shared by the ChatGPT and Gemini apps for the same query5.4% (opens in a new tab)Preprint; September 2026
Age of cited pages compared with organic resultsAI assistants about 25% "fresher"; AI Overviews no fresher (opens in a new tab)Vendor study; about 17 million citations

Answers also vary from run to run within one assistant. When the same prompt was repeated, ChatGPT and Google's AI returned the same list of brands less than once in 100 runs, and Claude only slightly more often (2,961 runs, November–December 2025) (opens in a new tab) (vendor study; a co-investigator works for a visibility-tracking vendor), and about 65% of cited sources changed from one day to the next (opens in a new tab) (preprint). Measuring AI visibility therefore takes repeated runs on each assistant.

Rules that apply to every assistant

Access checklist by assistant

Each check follows from the documentation cited in the tables above. The checks are our recommendations.

AssistantWhat to checkPass when
ChatGPTrobots.txt and the host or CDN firewallOAI-SearchBot is allowed, and OpenAI's published searchbot IP addresses are not blocked
Google AI Overviews and AI ModeSearch ConsoleKey pages are indexed with snippets allowed, and "Search generative AI" includes the site
Gemini approbots.txtGoogle-Extended is not disallowed
Microsoft CopilotBing Webmaster Tools and robots meta tagsKey pages are indexed in Bing, with no NOARCHIVE on pages you want cited
Clauderobots.txt and Brave SearchClaude-SearchBot and Claude-User are allowed, and key pages appear in Brave Search
Perplexityrobots.txtPerplexityBot is allowed

Measuring across assistants

Only Google and Microsoft give site owners data about AI answers, as the tables show. For the rest, three sources fill part of the gap:

Other AI crawlers in the same documentation

What the documentation does not settle

  • ChatGPT and Bing. OpenAI names Microsoft as one search provider among others and runs its own crawler. How much of ChatGPT search comes from each source is not published.
  • Claude and Brave. Anthropic lists Brave Search and TurboPuffer and runs Claude-SearchBot, but does not say how results are split between them.
  • Perplexity and other indexes. Perplexity's documentation describes only its own crawlers.
  • Google-Extended and AI Overviews. Google says the token does not affect Search, while one study found an association between blocking it and fewer AI Overview citations.
  • Speed. No controlled study has measured how long a new page takes to appear in each assistant's answers.

Frequently asked questions

If I block GPTBot, will my site disappear from ChatGPT?

No. GPTBot collects training data, and OAI-SearchBot controls search. OpenAI says each setting is independent (opens in a new tab), so a site can block GPTBot and still appear in ChatGPT search answers.

Does blocking Google-Extended remove my pages from AI Overviews?

According to Google, no: Google-Extended "does not impact a site's inclusion in Google Search" (opens in a new tab), and AI Overviews use the Search index. It does control grounding in the Gemini app, and one study found that publishers blocking it were less likely to be retrieved by AI Overviews (opens in a new tab) (association only).

Does ChatGPT use Bing?

Partly. OpenAI names Microsoft among ChatGPT's search providers (opens in a new tab), and Enterprise and Edu accounts may send queries to Bing (opens in a new tab). ChatGPT also uses OpenAI's own index, built by OAI-SearchBot (opens in a new tab). Being indexed in Bing is one route into ChatGPT search, not the only one.

Can I see which of my pages AI assistants cite?

For Copilot, yes: Bing Webmaster Tools reports page-level citations (opens in a new tab). For AI Overviews and AI Mode, Search Console reports impressions (opens in a new tab). ChatGPT, Claude, and Perplexity document no citation report, so server logs, analytics, and repeated prompt tests are the options.

Why do assistants give different answers to the same question?

They search different indexes, and each answer is a sample. The ChatGPT and Gemini apps shared only 5.4% of cited domains for the same query (opens in a new tab) (preprint), and repeated prompts rarely return the same list of brands (opens in a new tab).

Next steps

To check whether these crawlers can reach your key pages, AEO HQ sells a fixed-price Instant AEO Audit for $499. It includes a live crawl of your key pages, AI crawler access checks, and a side-by-side analysis of up to three competitors.

Change log

  • September 27, 2026: First published.

Sources

  1. Google. (2025, December 10). AI features and your website. Google Search Central. https://developers.google.com/search/docs/appearance/ai-features (opens in a new tab)
  2. Microsoft. (2026, August 18). Data, privacy, and security for web search in Microsoft Copilot and Microsoft Copilot Chat. Microsoft Learn. https://learn.microsoft.com/en-us/copilot/microsoft-365/manage-public-web-access (opens in a new tab)
  3. OpenAI. (n.d.). Overview of OpenAI crawlers. OpenAI Developers. Retrieved September 27, 2026, from https://developers.openai.com/api/docs/bots (opens in a new tab)
  4. OpenAI. (2026). Searching the web with ChatGPT. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/9237897-chatgpt-search (opens in a new tab)
  5. Anthropic. (n.d.). Subprocessors. Anthropic Trust Center. Retrieved September 27, 2026, from https://trust.anthropic.com/subprocessors (opens in a new tab)
  6. Anthropic. (2026, April 7). Does Anthropic crawl data from the web, and how can site owners block the crawler? Claude Help Center. https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler (opens in a new tab)
  7. Perplexity. (n.d.). Perplexity crawlers. Perplexity Docs. Retrieved September 27, 2026, from https://docs.perplexity.ai/guides/bots (opens in a new tab)
  8. Koster, M., Illyes, G., Zeller, H., & Sassman, L. (2022). Robots Exclusion Protocol (RFC 9309). RFC Editor. https://doi.org/10.17487/RFC9309 (opens in a new tab)
  9. OpenAI. (2026). Publishers and developers – FAQ. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/12627856-publishers-and-developers-faq (opens in a new tab)
  10. Google. (2026, July 14). Google's common crawlers. Google Crawling Infrastructure documentation. https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers (opens in a new tab)
  11. Google. (2026). Generative AI performance report (Search). Search Console Help. Retrieved September 27, 2026, from https://support.google.com/webmasters/answer/16984139 (opens in a new tab)
  12. Microsoft Bing. (n.d.). Bing Webmaster Guidelines. Retrieved September 27, 2026, from https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a (opens in a new tab)
  13. Madhavan, K., Merchant, M., Canel, F., & Nigam, S. (2026, February 10). Introducing AI Performance in Bing Webmaster Tools (public preview). Bing Webmaster Blog. https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview (opens in a new tab)
  14. Harsel, L. (2026, April 7). ChatGPT traffic analysis: Insights from 17 months of clickstream data. Semrush. https://www.semrush.com/blog/chatgpt-search-insights/ (opens in a new tab)
  15. Anthropic. (2026). Web search tool. Claude Platform Docs. Retrieved September 27, 2026, from https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool (opens in a new tab)
  16. Anthropic. (2026). Models overview. Claude Platform documentation. Retrieved September 27, 2026, from https://platform.claude.com/docs/en/about-claude/models/overview (opens in a new tab)
  17. OpenAI. (2026). ChatGPT search for Enterprise and Edu. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/10093903-chatgpt-search-for-enterprise-and-edu (opens in a new tab)
  18. Zhang, P., Ye, Q., Peng, Z., Garimella, K., & Tyson, G. (2025). Source coverage and citation bias in LLM-based vs. traditional search engines (arXiv:2512.09483). arXiv. https://doi.org/10.48550/arXiv.2512.09483 (opens in a new tab)
  19. Chen, M., Wang, X., Chen, K., & Koudas, N. (2025). Generative engine optimization: How to dominate AI search (arXiv:2509.08919). arXiv. https://doi.org/10.48550/arXiv.2509.08919 (opens in a new tab)
  20. Law, R. (2025, July 28). New study: AI assistants prefer to cite "fresher" content (17 million citations analyzed). Ahrefs. https://ahrefs.com/blog/do-ai-assistants-prefer-to-cite-fresh-content/ (opens in a new tab)
  21. Google. (2026, July 10). Optimizing your website for generative AI features on Google Search. Google Search Central. https://developers.google.com/search/docs/fundamentals/ai-optimization-guide (opens in a new tab)
  22. Google. (2026, August 19). Google's user-triggered fetchers. Google Crawling Infrastructure documentation. https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers (opens in a new tab)
  23. Google. (2026). Search generative AI control. Search Console Help. Retrieved September 27, 2026, from https://support.google.com/webmasters/answer/16908024 (opens in a new tab)
  24. Google. (2026). Default channel group. Google Analytics Help. Retrieved September 27, 2026, from https://support.google.com/analytics/answer/9756891 (opens in a new tab)
  25. Grossman, R., Liu, S., Chen, M. K., Smith, M., Borcea, C., & Chen, Y. (2026). How generative AI disrupts search: An empirical study of Google Search, Gemini, and AI Overviews. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 448–459). ACM. https://doi.org/10.1145/3805712.3809667 (opens in a new tab)
  26. Madhavan, K., Merchant, M., Nigam, S., & Shah, T. (2026, June 16). New AI visibility insights in Bing Webmaster Tools: Intents, topics, citation share, compare. Bing Search Blog. https://blogs.bing.com/search/June-2026/New-AI-Visibility-Insights-in-Bing-Webmaster-Tools-Intents-Topics-Citation-Share-Compare (opens in a new tab)
  27. Madhavan, K. (2025, October 8). Optimizing your content for inclusion in AI search answers. Microsoft Advertising Blog. https://about.ads.microsoft.com/en/blog/post/october-2025/optimizing-your-content-for-inclusion-in-ai-search-answers (opens in a new tab)
  28. IndexNow. (n.d.). FAQ. Retrieved September 27, 2026, from https://www.indexnow.org/faq (opens in a new tab)
  29. Canel, F. (2023, September 21). Wix now supports IndexNow for faster content indexing. Bing Webmaster Blog. https://blogs.bing.com/webmaster/september-2023/Wix-now-supports-IndexNow-for-Faster-Content-Indexing (opens in a new tab)
  30. Brave. (n.d.). Brave Search crawler. Retrieved September 27, 2026, from https://search.brave.com/help/brave-search-crawler (opens in a new tab)
  31. Belson, D., & Rhea, S. (2025, July 1). The crawl before the fall... of referrals: Understanding AI's impact on content providers. Cloudflare Blog. https://blog.cloudflare.com/ai-search-crawl-refer-ratio-on-radar/ (opens in a new tab)
  32. Kirsten, E., Grosse Perdekamp, J., Wu, Q., Upadhyay, M., Gummadi, K. P., & Zafar, M. B. (2025). Characterizing web search in the age of generative AI (arXiv:2510.11560). arXiv. https://doi.org/10.48550/arXiv.2510.11560 (opens in a new tab)
  33. Sharma, A. P. (2026). The discovery gap: How Product Hunt startups vanish in LLM organic discovery queries (arXiv:2601.00912). arXiv. https://arxiv.org/abs/2601.00912 (opens in a new tab)
  34. Fletcher, J., & Verckist, D. (2025, October). News integrity in AI assistants: An international PSM study. European Broadcasting Union & BBC. https://www.ebu.ch/research/open/report/news-integrity-in-ai-assistants (opens in a new tab)
  35. Uberti-Bona Marin, L. G., Bertaglia, T., Astante, G., Rijsbosch, B., van Dijck, G., Hannák, A., Spanakis, G., & Kollnig, K. (2026). "If I had to buy just ONE: Galaxy S26 Ultra": Auditing AI-generated product recommendations (arXiv:2609.18729). arXiv. https://doi.org/10.48550/arXiv.2609.18729 (opens in a new tab)
  36. Fishkin, R. (2026, January 28). NEW research: AIs are highly inconsistent when recommending brands or products. SparkToro. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/ (opens in a new tab)
  37. Schulte, J., Bleeker, M., & Kaufmann, P. (2026). Don't measure once: Measuring visibility in AI search (GEO) (arXiv:2604.07585). arXiv. https://doi.org/10.48550/arXiv.2604.07585 (opens in a new tab)
  38. Zecchini, G., Moore, A. A., Ubl, M., & Siddle, R. (2024, December 17). The rise of the AI crawler. Vercel. https://vercel.com/blog/the-rise-of-the-ai-crawler (opens in a new tab)
  39. Google. (2026, August 28). Spam policies for Google web search. Google Search Central. https://developers.google.com/search/docs/essentials/spam-policies (opens in a new tab)
  40. Microsoft. (2025). AIPlatform and PaidAIPlatform. Microsoft Clarity documentation, Microsoft Learn. https://learn.microsoft.com/en-us/clarity/insights/ai-channel-group (opens in a new tab)
  41. Apple. (2026, September 4). About Applebot. Apple Support. https://support.apple.com/en-us/119829 (opens in a new tab)
  42. Meta. (n.d.). Meta web crawlers. Meta for Developers. Retrieved September 27, 2026, from https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/ (opens in a new tab)

How to cite this page

Maxwell, P. (2026). How ChatGPT, Gemini, Claude, Perplexity, and Copilot find and cite sources. AEO HQ. Last updated September 27, 2026. https://www.aeohq.ai/articles/how-ai-assistants-find-sources

Pages in Ranking in AI assistants

Next step

Find out who AI recommends.

Book the audit to see where you rank, where AI cites you, and where competitors win. The full audit price credits toward the Blueprint within 30 days.