Reference · Measuring AI visibility
How AEO HQ measures AI visibility, and how our research is done
How AEO HQ grades evidence, pulls search data, and measures AI visibility, and what its $499 automated audit does and does not measure today.
By Paul Maxwell, founder of AEO HQ
Published · Updated
AEO HQ recommends measuring AI visibility as a rate: the share of repeated runs, on each AI engine, in which answers to buyer questions name or cite a company, reported with an error range. Our $499 automated audit does not yet measure it this way. Our research uses sources published from September 2023 to September 2026, labeled by source type and strength of evidence, and every fact our pages state links to a source that was retrieved and checked.
This page describes those methods for AEO HQ's work on answer engine optimization (AEO), which is also called generative engine optimization (GEO). It also states what our $499 automated audit measures today, and what it does not. Statements about other organizations link to their sources. Statements about AEO HQ's own methods describe our research records and code as of 27 September 2026.
Evidence standard for our research
AEO HQ's research is organized in seven tracks. Each track is written up as a research dossier dated 27 September 2026.
| Track | Topic |
|---|---|
| 01 | How answer engines select and cite sources, and which tactics replicate |
| 02 | How AI assistants decide which companies to recommend, and how to measure it |
| 03 | Changes in search: Google's AI features, crawlers and indexes, and spam policies |
| 04 | How B2B buyers use AI assistants, and what makes self-serve buying work |
| 05 | Whether AI agents can find and buy services |
| 06 | Competing AEO providers, and the sources that feed AI answers about them |
| 07 | Search demand and rankings for AEO topics, from Ahrefs data |
The first five dossiers list 385 sources in their evidence tables, and the competitor track logs 206 retrieved web pages. All seven follow the rules below.
Evidence window. Sources published from September 2023 to September 2026. Older work appears only as labeled background.
Source tiers. Each source carries one of four tiers. A tier describes the kind of source, not whether a finding is right. Official documentation, for example, is the best evidence of what a platform says it does, but not of how large an effect is.
| Tier | What it covers |
|---|---|
| T1 | Peer-reviewed journal articles and scholarly books |
| T2 | Peer-reviewed conference papers |
| T3 | Preprints and working papers, including arXiv papers and theses |
| T4 | Grey literature: official platform documentation, industry and vendor studies, surveys, and journalism |
Evidence strength. Where it matters, a finding carries one of four labels:
- Strong: official documentation that states the fact directly, or several independent studies, including peer-reviewed work, that agree.
- Moderate: consistent evidence that comes mostly from one study, from laboratory settings, from vendor data, or from correlations.
- Weak: one small, observational, or conflicted study.
- Contested: credible studies disagree.
Verification. Every source was retrieved and read during the research. When only an abstract could be read, the dossier says so. Leads that could not be verified are listed as unverified and are not used. Numbers are copied from the source with their sample, method, and date. When a study's authors work for a company that sells a related product, we say so, because much of the research on AI search comes from such companies. When findings conflict, we report both.
Source copies and the claims registry. Each cited page is saved as text with the date it was fetched, and facts are checked against that copy. Every fact a page states is recorded in a claims registry with its source, tier, strength, and the date it was last checked. On 27 September 2026 the registry held 657 claims. When a source changes, we update the claim and every page that uses it.
What we never cite. A source that was not read, an AI-generated summary, a secondary article in place of the primary source, or a number without its sample and date.
Errata. Corrections found after a dossier was written go into an errata list, which overrides the dossiers. Each entry records the date, what the dossier said, the correction, and the source that settles it. On 27 September 2026 the list had 21 entries. They include a documentation page that no longer stated a figure a dossier had used, a survey figure quoted without its qualifier, a percentage that did not match the dossier's own table, and co-authors listed for articles that have a single byline. When a correction affects a published page, we correct the page and record the change in its change log.
Search data
The search volumes and keyword difficulty scores on AEO HQ's pages come from one pull of Ahrefs data through the Ahrefs API (version 3). Keyword and results-page data were pulled on 27 September 2026, and site data is dated 26 September 2026. All keyword, results-page, and organic-traffic data are for the United States. The reports we pulled describe Google search, not AI assistants, so this pull measured no AI answers.
What we pulled
| Ahrefs report (API endpoint) | What it returns | What we pulled |
|---|---|---|
Keyword overview (keywords-explorer/overview) | Estimated monthly searches, keyword difficulty, traffic potential, parent topic, results-page features such as an AI Overview, search intent, and cost per click | 122 queries: buyer questions collected from search autocomplete, FAQ headings on ranking pages, and our own searches (Reddit thread titles were left out), plus 11 broad terms. Ahrefs returned data for 84. |
Matching terms (keywords-explorer/matching-terms) | Other queries that contain a seed's words, or question forms of the seed | First pass: 6 seeds ("aeo", "answer engine optimization", "generative engine optimization", "llm seo", "ai search optimization", and "ai visibility") and question forms of 3 of them, top 50 each with at least 10 searches a month. Second pass: "answer engine" and question forms of "aeo" (top 50 each), and 11 narrower seeds such as "aeo checklist" (up to 30 each). |
Related terms (keywords-explorer/related-terms) | Queries related to a seed | "answer engine optimization" (77 rows) |
Results-page overview (serp-overview/serp-overview) | The top 10 organic results and page features, including the sources an AI Overview cites | 25 queries, of which 20 returned data |
Domain rating and metrics (site-explorer/domain-rating, site-explorer/metrics) | Domain Rating, Ahrefs' 0–100 score for the strength of a site's backlink profile (opens in a new tab), and estimated organic keywords and traffic | 15 competing firms, plus aeohq.ai as a baseline |
Organic keywords (site-explorer/organic-keywords) | The queries a site ranks for, with estimated traffic | 5 competitors, filtered to AEO topics, top 50 by traffic |
Referring domains (site-explorer/refdomains) | The sites that link to a page | 4 "best agencies" list pages |
In all, the pull made 93 paid calls and 13 free pricing calls and used 29,394 API units, against an approved cap of 30,000. The cap check skipped two planned calls.
Volumes are estimates
Ahrefs says "Search volume in any SEO tool, including Ahrefs, is always an estimation. Only Google has access to the exact data." (opens in a new tab) It builds its estimates from "a combination of Google Keyword Planner (GKP), Google trends data and other third party data sources" (opens in a new tab) (vendor documentation). Our own pull shows the gap between methods: "aeo agency" had 1,900 estimated monthly searches in the keyword overview and 2,600 in the organic keywords report. We therefore use volumes to compare queries with one another, not to forecast traffic.
Keyword difficulty is also an estimate. Ahrefs bases it on the number of referring domains that the top 10 organic results have, and it does not take on-page SEO factors into account (opens in a new tab) (vendor documentation).
Queries we excluded
- Suspected synthetic queries. Some rows were long, prompt-like questions or odd name combinations with high estimated volumes, often with no difficulty score. An example is "which home sauna has the highest aeo and seo?", at 2,600 estimated monthly searches. We flagged 25 such rows in the first expansion pull. After the second pull, our keyword registry excluded 34 suspected synthetic queries from every total and priority. This is AEO HQ's judgment, not an Ahrefs label.
- Low-reliability volumes. Prompt-like questions of six or more words, or ending in a question mark, with 150 or more searches and no difficulty score are kept as question ideas but left out of demand totals. On 27 September 2026 this rule covered 119 keywords.
- Another meaning of "aeo". Some rows used "aeo" to mean the retailer American Eagle Outfitters. For example, "aeo stock" had about 25,000 estimated monthly searches. These rows were removed.
Two problems came up during the pull. First, a keyword filter on the matching-terms report had no effect, so we filtered those rows ourselves. Second, five low-volume queries returned no results-page data but still cost the 50-unit minimum each, 250 units in total.
Cost controls
Ahrefs charges API units for each request. The cost is the larger of 50 units and the per-row cost times the number of rows returned, and requests served from cache cost nothing (opens in a new tab) (vendor documentation). We controlled spending in five ways:
- Price first. Before each paid call, we ran the same kind of call against one of Ahrefs' free test targets, which do not consume units (opens in a new tab), and read the per-row cost from the response headers.
- Cap. We agreed a cap of 30,000 units in advance, and the script skipped any call whose estimated cost would pass it.
- Cache. Every response was saved and never pulled twice.
- Log. Every call, free or paid, was logged with its rows and units. During the first two stages, the workspace's usage counter rose by exactly the 19,560 units in our log.
- Keys. The API key was read from the environment and never stored in files.
The raw responses, the call log, and the scripts are kept in AEO HQ's research files and are not published. Pages that use these figures cite them as an unpublished AEO HQ data set pulled on 27 September 2026, with a link to this section.
The measurement design we recommend
This is the design we recommend for measuring a company's visibility in AI answers. Our automated audit does not yet follow it: What the automated $499 audit does today describes what the audit measures instead. The design follows the research reviewed in track 02 and treats each answer as one sample, because the same question can get a different answer each time: when the same prompt was repeated, ChatGPT and Google's AI returned the same list of brands less than once in 100 runs, and Claude only slightly more often (opens in a new tab) (2,961 runs, November–December 2025; vendor study).
What the design reports
Each measure is reported separately for each engine, over rolling four-week windows.
| Measure | Definition |
|---|---|
| Mention rate (primary) | The share of runs on unbranded buyer prompts in which the answer names the company as an option |
| Citation rate | The share of runs in which the answer cites or links any of the company's pages |
| Share of voice | The company's mentions divided by all mentions of a named set of 5 to 10 competitors |
| Accuracy | The share of branded prompts, such as "What is [company]?", answered correctly about services, prices, and people, graded by a person |
| Sentiment | Reported as description only: in one panel, whether a brand was framed positively or negatively flipped about 6.7 times more often than whether it was mentioned (opens in a new tab) (preprint; the author co-founded the tracking platform studied) |
The prompt panel
- Fixed. A fixed set of buyer questions, kept the same so that trends can be compared. By default it has 40 buyer intents, each written three ways (120 prompts), plus 20 branded prompts. A smaller panel uses 20 intents.
- Varied. The intents cover asking for a provider, asking about a specific service, describing a problem, comparing options, adding constraints such as price, and different buyer roles.
- Written the way people write. Wording comes from sales calls, forums, and the questions buyers type into search, because people phrase the same need very differently: 142 human-written prompts for the same intent had a mean semantic similarity of 0.081, yet produced similar brand sets (opens in a new tab) (vendor study).
- Market. U.S. English, unless the company sells in other markets.
Engines and runs
- Engines. Seven consumer interfaces, used logged out or with a clean account and default settings: ChatGPT, Google AI Mode, Google AI Overviews (recording whether one appears), Gemini, Perplexity, Microsoft Copilot, and Claude with web search on. APIs serve only as low-cost proxies, checked monthly against the interfaces, because the API and the consumer interface of the same assistant shared only 12.0% (ChatGPT) and 14.8% (Gemini) of cited domains (opens in a new tab) (preprint).
- What each run records. The product, mode, model label, date and time, location, account state, whether search was on, and the citations returned. A 2026 critical survey of 45 studies asks studies to record these details (opens in a new tab) (preprint).
- Runs. Three runs per prompt per engine each week, and eight for the ten most important prompts. A standard error is the typical size of the error in an estimate. In one daily panel, the standard error of a brand's detection rate fell from 0.370 with one run per prompt to 0.081 with seven and 0.062 with eight, and to 0.033 when runs were pooled over 28 days (opens in a new tab) (preprint; the lead author is affiliated with a tracking vendor). The same survey of 45 studies recommends 7 to 8 repetitions per prompt, 3 to 5 paraphrases, several named engines, and multiple time windows (opens in a new tab) as a starting design.
- More prompts before more runs. Repeating a prompt reduces only the variance within that prompt, and clustered standard errors can be more than three times larger than naive ones (opens in a new tab) (preprint). Answers to one prompt cluster: across 102,025 API responses, 77.5% of brand, prompt, and engine combinations were either always or never mentioned (opens in a new tab) (preprint). So the design adds prompts before it adds runs.
Statistics
- The main number is a share of runs: the runs whose answer names the company, divided by all runs.
- Intervals suited to small samples. The design uses Wilson score intervals or Bayesian intervals, not normal-approximation intervals, because normal-approximation intervals are too narrow below a few hundred data points: nominal 95% intervals covered the true value only 92.5% of the time at 100 data points (opens in a new tab) (peer-reviewed). A Wilson score interval is a range for a proportion; the same authors recommend it for small samples.
- Uncertainty across prompts. Intervals come from resampling buyer intents, then runs within them (a cluster bootstrap), because answers to the same prompt are correlated.
- Engines separately. Each engine is reported on its own, never only as a pooled number.
- Changes over time. A change between two windows is reported only when the interval for the difference excludes zero. Model releases and product changes are marked on the timeline.
- Matching. Matching rules, such as exact company names and domains, are fixed in advance. A person checks every match of a name that other companies or people share, and a random 10% of other matches.
The table shows how much different numbers of runs can tell you. The intervals are AEO HQ's calculation. They treat runs as independent, and because answers to one prompt are correlated, real intervals are wider. One week of the default panel on one engine is 360 runs.
| Runs that named the company | Observed rate | 95% Wilson interval |
|---|---|---|
| 0 of 5 | 0% | 0.0% to 43.4% |
| 1 of 5 | 20% | 3.6% to 62.4% |
| 3 of 30 | 10% | 3.5% to 25.6% |
| 12 of 120 | 10% | 5.8% to 16.7% |
| 36 of 360 | 10% | 7.3% to 13.5% |
Outcome data
Platform reports show outcomes, not mention rates, so the design uses them alongside the panel (all official documentation):
- Google Search Console's generative AI performance report counts impressions of a site's links in AI Overviews and AI Mode (opens in a new tab), not clicks.
- Bing Webmaster Tools' AI Performance report shows citations and cited pages for Copilot and Bing's AI summaries (opens in a new tab), along with the grounding queries behind them.
- In Google Analytics 4, the default AI Assistant channel excludes AI Overviews and AI Mode, which count as Organic Search (opens in a new tab), so AI referral traffic needs a channel definition of its own.
- The design also asks buyers where they first heard of the company, because referrer data misses influence that does not end in a click.
What the automated $499 audit does today
This section describes the Instant AEO Audit, AEO HQ's $499 automated audit, as its code worked on 27 September 2026. It is a quick check of a site's technical setup plus a small sample of one AI model's answers. It is not the measurement design described above, and its AI results are directional only.
What it does
- Pages. It fetches the home page and up to five other key pages through a web-scraping service (Firecrawl). It picks them by path from a map of up to 200 of the site's URLs, for example /pricing, /product, /about, /faq, /customers, /articles, /docs, and /compare. From the home page's HTML it records the schema.org types in the page's JSON-LD structured data, the number of H1 headings, the canonical URL, the meta robots tag, whether a social preview image (og:image) is set, and the page language.
- Crawler access and files. It requests three files at the site's root: robots.txt, llms.txt, and sitemap.xml. It records whether robots.txt exists; which of seven crawler names robots.txt blocks from the whole site (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, and CCBot); whether robots.txt lists a sitemap; whether sitemap.xml exists; and whether llms.txt exists as a text file.
- Authority. From Ahrefs, it pulls Domain Rating, which Ahrefs describes as showing "the strength of a website's backlink profile compared to the others in our database on a 100-point scale" (opens in a new tab) (vendor documentation). It also pulls estimated organic keywords and organic traffic, the 25 U.S. organic keywords that bring the most estimated traffic, and up to eight U.S. organic competitors.
- Competitors. It compares up to three competitor domains given by the buyer. If none are given, it takes the first three from the Ahrefs competitor list. For each one, it pulls Domain Rating and organic estimates and runs the same home-page checks.
- AI answers. One model,
openai/gpt-5.4-mini, called through an API, reads the opening text of up to three crawled pages. It writes the company's name, its category, and five questions a buyer might ask an AI assistant before choosing a vendor, without naming the company. The same model then answers each question once, with the instruction "Name specific vendors or products and briefly say why." It has no web search, so its answers come from its parametric knowledge. An answer counts as a mention if its text contains the company's name or the first part of its domain, such as "acme" for acme.com. Competitors are matched by the first part of their domains. The report gives the share of the five answers that mention the company. If no page could be crawled, this step is skipped. - Report. A second model,
anthropic/claude-sonnet-5, writes the report from the collected data: a short summary, an overall readiness score from 0 to 100, four sub-scores (technical, content, authority, and AI visibility), 12 to 20 prioritized tasks with steps, and notes on each competitor. It is instructed to justify each task with the collected evidence, not to invent metrics, and to say when data is missing. If an earlier step fails, the audit records the error and continues without that data.
Limits
- One run per question. Five answers are one small sample, and answers vary from run to run. If none of the five answers names the company, the 95% Wilson interval for its true mention rate still runs from 0% to 43% (AEO HQ's calculation, from the table above).
- One model, through its API. The audit does not test the ChatGPT, Google, Gemini, Perplexity, Claude, or Copilot interfaces that buyers use. A 2026 audit of AI product recommendations concluded that "neither isolated responses nor API observations can be assumed to represent the commercial advice consumers encounter" (opens in a new tab) (preprint).
- No live search. The audit shows what one model recalls from training, not what an assistant finds when it searches. This understates newer companies most: among 112 startups from the 2025 Product Hunt leaderboard, a model without web access named 5.4% of them at least once in discovery-style questions, while a search-based model named 27.7% (opens in a new tab) (preprint; a master's thesis).
- Questions written by the model. The five questions come from the company's own pages, not from research into what its buyers ask.
- Text matching without review. A company whose name is a common word can be over-counted, and a name written another way can be missed. No person checks the matches.
- A narrow robots.txt check. The audit flags a crawler only when a group of rules naming that crawler, or the "*" group, disallows the whole site ("Disallow: /"). It does not apply the standard's rule that a crawler obeys the group that names it and uses the "*" group only when no group matches (opens in a new tab) (Internet standard), so it can report a block that a crawler would not obey. It does not detect partial blocks. It does not check Claude-SearchBot, which Anthropic uses to index content for Claude's search results (opens in a new tab), or Claude-User, Perplexity-User, Googlebot, or Bingbot. It cannot see blocks set in a firewall or content delivery network, although Google lists making sure crawling is allowed "in robots.txt, and by any CDN or hosting infrastructure" (opens in a new tab) among the basics for its AI features (official documentation).
- Structured data on the home page only. The audit lists the schema.org types found in the home page's JSON-LD. It does not validate the markup or check other pages.
- Files found, contents unchecked. The llms.txt and sitemap checks confirm that the files respond, not what they contain. The sitemap check looks only at /sitemap.xml and at a Sitemap line in robots.txt.
- A score without a formula. The report model writes the readiness score and its sub-scores, and the code sets no formula or weights for them. Read them as the model's summary judgment of the evidence, not as a measurement. We have not tested how much they vary when the same site is audited twice.
- Estimates from Ahrefs. Domain Rating, keyword, and traffic figures are Ahrefs estimates, as described in Search data.
Use the audit as a quick technical check and a first look at one model's built-in associations. To measure AI visibility, use the design in the previous section. When the audit changes, we will update this section and note the change in the change log.
What we do not claim
- No guaranteed rankings, citations, mentions, or traffic. No outside provider controls these results. Google says third-party tools "can't guarantee performance," and lists services that promise improvements for AI experiences, "also known as 'AEO' or 'GEO' tools," among those to evaluate critically (opens in a new tab). Bing says "GEO does not guarantee grounding or citations in AI experiences" (opens in a new tab), and OpenAI says of ChatGPT search that "Placement is not guaranteed" (opens in a new tab) (all official documentation).
- No "#1 in ChatGPT." Assistants do not return a stable ranked list: it took about 1 in 1,000 runs to see two brand lists in the same order (opens in a new tab) (2,961 runs; vendor study). The design we recommend reports shares of runs with intervals, for each engine.
- No general uplift figures such as "+40% visibility". That figure comes from an experiment in a simulated engine where the page was already retrieved (opens in a new tab) (peer-reviewed), and a 2026 critical survey of 45 studies rejects the 40% figure as a general claim (opens in a new tab) (preprint).
- No timelines. No study we found has measured how long a company takes to first appear in AI answers, and the same survey rates the claim that a white-hat (policy-compliant) GEO intervention durably improves discoverability across multiple engines as low-confidence (opens in a new tab), so we do not promise when a company will appear in AI answers.
- No single screenshot or single run as evidence. One answer is one draw from a distribution, as the limits of our own automated audit show.
- No inside knowledge of any platform. Our statements about ChatGPT, Google, Gemini, Perplexity, Claude, and Copilot come from their public documentation and from published studies. Google says it "doesn't evaluate third-party services" and that third-party tools "don't have access to our internal ranking data" (opens in a new tab) (official documentation). That applies to AEO HQ too.
Change log
- September 27, 2026: First published. The section on the automated audit describes its code as of 27 September 2026.
Sources
- Ahrefs. (n.d.). What is Domain Rating (DR)? Ahrefs Help Center. Retrieved September 27, 2026, from https://help.ahrefs.com/en/articles/1409408-what-is-domain-rating-dr (opens in a new tab)
- Ahrefs. (n.d.). How accurate is keyword search volume in Ahrefs? Ahrefs Help Center. Retrieved September 27, 2026, from https://help.ahrefs.com/en/articles/72571-how-accurate-is-keyword-search-volume-in-ahrefs (opens in a new tab)
- Ahrefs. (n.d.). What does KD stand for in Keywords Explorer? Ahrefs Help Center. Retrieved September 27, 2026, from https://help.ahrefs.com/en/articles/72265-what-does-kd-stand-for-in-keywords-explorer (opens in a new tab)
- Ahrefs. (n.d.). Limits consumption. Ahrefs for Developers. Retrieved September 27, 2026, from https://docs.ahrefs.com/en/api/docs/limits-consumption (opens in a new tab)
- Ahrefs. (n.d.). Free test queries. Ahrefs for Developers. Retrieved September 27, 2026, from https://docs.ahrefs.com/en/api/docs/free-test-queries (opens in a new tab)
- Fishkin, R. (2026, January 28). NEW research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility. SparkToro. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/ (opens in a new tab)
- Kumar, P. (2026). Generative engine optimization at scale: Measuring brand visibility across AI search engines (arXiv:2606.20065) [Preprint]. arXiv. https://arxiv.org/abs/2606.20065 (opens in a new tab)
- Uberti-Bona Marin, L. G., Bertaglia, T., Astante, G., Rijsbosch, B., van Dijck, G., Hannák, A., Spanakis, G., & Kollnig, K. (2026). "If I had to buy just ONE: Galaxy S26 Ultra": Auditing AI-generated product recommendations (arXiv:2609.18729) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2609.18729 (opens in a new tab)
- Martinez, O. (2026). Optimizing visibility in generative engines: A critical survey of generative engine optimization (2023–2026) (arXiv:2607.14035) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.14035 (opens in a new tab)
- Schulte, J., Bleeker, M., & Kaufmann, P. (2026). Don't measure once: Measuring visibility in AI search (GEO) (arXiv:2604.07585) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.07585 (opens in a new tab)
- Miller, E. (2024). Adding error bars to evals: A statistical approach to language model evaluations (arXiv:2411.00640) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2411.00640 (opens in a new tab)
- Bowyer, S., Aitchison, L., & Ivanova, D. R. (2025). Position: Don't use the CLT in LLM evals with fewer than a few hundred datapoints. In Proceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267, pp. 81143–81184). PMLR. https://proceedings.mlr.press/v267/bowyer25a.html (opens in a new tab)
- Google. (2026). Generative AI performance report (Search) [Search Console Help]. Retrieved September 27, 2026, from https://support.google.com/webmasters/answer/16984139 (opens in a new tab)
- Madhavan, K., Merchant, M., Canel, F., & Nigam, S. (2026, February 10). Introducing AI Performance in Bing Webmaster Tools public preview. Bing Webmaster Blog. https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview (opens in a new tab)
- Google. (2026). Default channel group [Analytics Help]. Retrieved September 27, 2026, from https://support.google.com/analytics/answer/9756891 (opens in a new tab)
- Sharma, A. P. (2026). The discovery gap: How Product Hunt startups vanish in LLM organic discovery queries (arXiv:2601.00912) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2601.00912 (opens in a new tab)
- Koster, M., Illyes, G., Zeller, H., & Sassman, L. (2022). Robots Exclusion Protocol (RFC 9309). RFC Editor. https://doi.org/10.17487/RFC9309 (opens in a new tab)
- Anthropic. (2026, April 7). Does Anthropic crawl data from the web, and how can site owners block the crawler? Claude Help Center. https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler (opens in a new tab)
- Google. (2025, December 10). AI features and your website. Google Search Central. https://developers.google.com/search/docs/appearance/ai-features (opens in a new tab)
- Google. (2026, June 5). Google Search's guidance on using third-party SEO tools, services, and advice. Google Search Central. https://developers.google.com/search/docs/fundamentals/third-party-seo (opens in a new tab)
- Microsoft Bing. (n.d.). Bing Webmaster Guidelines. Retrieved September 27, 2026, from https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a (opens in a new tab)
- OpenAI. (n.d.). Searching the web with ChatGPT [Help Center article]. Retrieved September 27, 2026, from https://help.openai.com/en/articles/9237897-chatgpt-search (opens in a new tab)
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative engine optimization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 5–16). ACM. https://doi.org/10.1145/3637528.3671900 (opens in a new tab)
How to cite this page
Maxwell, P. (2026). How AEO HQ measures AI visibility, and how our research is done. AEO HQ. Last updated September 27, 2026. https://www.aeohq.ai/methodology
More in Measuring AI visibility
Complete guide
How to measure AI visibility
How to measure AI visibility: what to count, how many prompts and runs, error ranges, and the Google, Bing, and GA4 reports that fill the gaps.
Guide
How to track AI referral traffic in GA4
Track AI referral traffic in GA4: what the AI Assistant channel counts, a custom channel for ChatGPT, Claude, Perplexity, and others, and what GA4 cannot see.
Reference
AEO and AI visibility tools compared
AEO and AI visibility tools by kind: what search engine reports, analytics, prompt trackers, graders, and log tools can and cannot measure, from vendors' pages.
Reference
AEO metrics and KPIs: definitions
AEO and GEO metrics and KPIs defined: mention rate, citation rate, share of voice, accuracy, AI impressions, and AI referrals, with formulas and limits.