Complete guide · Measuring AI visibility
How to measure AI visibility
How to measure AI visibility: what to count, how many prompts and runs, error ranges, and the Google, Bing, and GA4 reports that fill the gaps.
By Paul Maxwell, founder of AEO HQ
Published · Updated
AI visibility is how often AI assistants mention your company, cite your pages, and describe you correctly when people ask them questions. To measure it, run a fixed set of buyer questions many times on each assistant, count the share of answers that name or cite you, and report that share for each assistant with an error range. Then add the reports Google and Microsoft give site owners, your analytics, your server logs, and what buyers tell you.
The method follows from one fact: the same question gets different answers from one run to the next. When the same prompt was repeated, ChatGPT and Google's AI returned the same list of brands less than once in 100 runs, and Claude only slightly more often (opens in a new tab) (vendor study; 2,961 runs, November–December 2025). A rate across many runs is steadier than any single answer. This page explains what can be measured, which instrument observes each part, the steps, and the statistics. AEO HQ's recommended panel sizes are on the methodology page, and four common measurement mistakes are covered in AEO antipatterns. Sources were checked on September 27, 2026.
Pages in this section
- AEO metrics and KPIs: definitions
- How to track AI referral traffic in GA4
- AEO and AI visibility tools compared
- How AEO HQ measures AI visibility
Scope and definitions
This page covers assistants that answer questions in full sentences and sometimes cite web pages: ChatGPT, Google's AI Overviews and AI Mode, Gemini, Claude, Perplexity, and Microsoft Copilot. Work aimed at appearing in their answers is called answer engine optimization (AEO) or generative engine optimization (GEO). The measurement is the same under either name.
- AI visibility is the extent to which AI assistants mention, cite, and correctly describe a company, its products, or its pages. It is also called AI search visibility or AI brand visibility. Microsoft uses the idea for its own site-owner reports and says that in AI answers, visibility "is also about whether your content is cited and referenced when AI systems generate answers." (opens in a new tab)
- A prompt is the question typed into an assistant. A run is one answer to one prompt, collected once. The run is the unit you count.
- A mention is the company's name in an answer. A citation is a link to one of its pages shown with the answer.
- The mention rate is the number of runs that mention the company divided by all runs. The citation rate is the same calculation for citations.
- Share of voice is the company's mentions divided by all mentions of a named set of competitors.
- Accuracy is the share of answers about the company that state its facts correctly, as graded by a person.
- A 95% confidence interval is a range built by a method that, over many repeated samples, contains the true rate 95% of the time. It is the error range around a measured rate.
- A prompt panel is a fixed set of prompts run on a schedule.
What there is to measure
AI visibility is not one number. A 2026 critical survey of 45 studies proposes treating it as a "visibility vector" of seven separate parts rather than a single rank (opens in a new tab) (preprint). The first two columns below follow the survey. The third is our mapping of each part to the instruments a site owner can use.
| Part | What it means | How a site owner can observe it |
|---|---|---|
| Discoverability | The chance that the assistant retrieves your page | Crawler and fetcher requests in server logs; pages listed in Bing's AI Performance report |
| Exposure | How much of the assistant's working context your page gets, such as its rank or length | Not visible from outside the assistant |
| Mention or citation | The chance that the answer names or links you | A prompt panel; Bing's citation counts; Search Console's AI impressions |
| Prominence | Where and how often you appear in the answer | Position can be recorded in a panel, but order is unstable between runs |
| Absorption | How much your content shaped the facts and wording of the answer | Only by reading answers against your pages |
| Fidelity | Whether claims attributed to you are supported and accurate | Grading answers against a written fact sheet |
| Behavior | What people did next: clicks, visits, sign-ups, purchases | Analytics, a "How did you hear about us?" question, your CRM |
The survey adds two warnings. A single score that combines these parts is "defensible only when the weights" match an explicit objective, and combining a mention, an accurate citation, and a conversion without one "merely obscures normative choices" (opens in a new tab). And a high chance of being cited once retrieved does not make up for a low chance of being retrieved at all (opens in a new tab), so the two should be reported separately. Undisclosed visibility scores are covered as an antipattern in AEO antipatterns.
How AI answers vary, and what that means for measurement
Some answers come from memory and some from search
An assistant answers some questions from what its model learned in training, called parametric knowledge, and others by searching the web first, called retrieval. Search is not always on. In a U.S. clickstream panel, ChatGPT ran a web search on 34.5% of queries in February 2026, down from 46% in late 2024 (opens in a new tab) (vendor data). In a panel of German-language commercial prompts, 57.8% of ChatGPT runs returned no citations (opens in a new tab) (preprint; January–March 2026).
This changes what a citation rate means. The survey of 45 studies writes the chance that a page is cited as the product of three chances: that the assistant searches, that the page is retrieved when it does, and that it is cited once retrieved (opens in a new tab). It says that "outputs without search, without citations, or with errors are outcomes, not data to be discarded." (opens in a new tab) A rate calculated only among answers that contain citations leaves out the first factor and overstates visibility.
The same prompt gets different answers
- In an October 2023 audit that entered 28 prompts four times each into Perplexity, Google Bard, and Bing Chat, answers from different instances of the same chatbot differed for up to 53% of outputs (opens in a new tab) (peer-reviewed).
- In the SparkToro study, it took about 1 in 1,000 runs to see two brand lists in the same order (opens in a new tab) (vendor study; 600 volunteers ran 12 prompts; a co-investigator works for a visibility-tracking vendor).
- In a daily panel of prompts run on ChatGPT, Gemini, Google AI Mode, and Perplexity, about 65% of cited sources changed from one day to the next (opens in a new tab) (preprint; the first author is affiliated with a tracking vendor).
Rates across many runs are steadier. Across 102,025 API responses, 77.5% of brand, prompt, and engine combinations were either always or never mentioned (opens in a new tab), and whether an answer framed a brand positively or negatively flipped about 6.7 times more often than whether it mentioned the brand (opens in a new tab) (preprint; March–May 2026; the author co-founded the tracking platform studied). Whether a brand is mentioned is far steadier than how it is described.
Wording, interface, user, and time change the answer
- Wording. Across 6.5 million test instances, 20 models, and 39 tasks, evaluations that relied on a single prompt wording gave brittle results (opens in a new tab) (peer-reviewed). Buyers also word the same need very differently: 142 human-written prompts for the same intent had a mean semantic similarity of 0.081, yet produced similar brand sets (opens in a new tab) (vendor study). The type of question matters too: on broad discovery questions the average tracked brand was named about 23% of the time, and on specific problem and use-case questions about 11% (opens in a new tab) (preprint).
- Interface. For the same queries, the ChatGPT and Gemini consumer interfaces shared only 5.4% of cited domains, and each assistant's API shared only 12.0% (ChatGPT) and 14.8% (Gemini) with its own interface (opens in a new tab) (preprint; September 2026). The authors concluded that audits "should therefore account for repeated responses, consumer-facing conditions, and the source layer being observed." (opens in a new tab)
- User. A user's revealed identity significantly changed chatbot recommendations (p < 0.001) (opens in a new tab) (peer-reviewed). Account history and settings are part of what you are measuring unless you control them.
- Time. GPT-4's accuracy on one task fell from 84% in March 2023 to 51% in June 2023, and the authors call for continuous monitoring (opens in a new tab) (peer-reviewed). A trend line can move because the model changed, not because anything about your company did.
The author of the SparkToro study began as a skeptic of AI tracking and ended with two conclusions: tracking "ranking position" in AI tools is "foolhardy," while "visibility % across dozens to hundreds of prompts run multiple times is a reasonable metric." (opens in a new tab) The steps below put that into practice.
Steps
These steps are AEO HQ's recommendations, based on the studies above. AEO HQ's default panel sizes, engines, and runs are on the methodology page, so they are not repeated here.
1. Decide what the numbers are for
Write down the decision the measurement serves, such as which assistant to prioritize or whether a change worked. Then pick the measures: the mention rate on unbranded buyer prompts as the main measure, plus the citation rate, share of voice, and accuracy on branded prompts. Report each one separately rather than as a combined score.
2. Write the prompt panel in your buyers' words
Collect questions from sales calls, emails, forum threads, and the search queries in Google Search Console. Group them into intents, and write three to five phrasings of each: a review of 45 studies proposes 7–8 repetitions per prompt and 3–5 paraphrases per question as a starting point (opens in a new tab) (preprint). Mix broad questions ("best tools for X") with specific ones ("which X works with Y at under $Z"), because their rates differ. Add a smaller set of branded prompts, such as "What does [company] charge?", for the accuracy check. Then freeze the panel. If you must change it, run the old and new versions side by side for one window so the trend line can be joined.
3. Choose the assistants and control how you run them
Measure each assistant on its own, in the consumer interface that buyers use. Use APIs only as low-cost proxies, checked against the interface. Start a new chat for each prompt, as the 2023 audit above did to limit the influence of earlier chat history (opens in a new tab), and use a logged-out session or a clean account, default settings, and a fixed location. Record the same fields for every run. The review of 45 studies asks studies to record the product, mode, model, date, locale, account, and whether search was enabled (opens in a new tab). A run log can look like this:
| Field | What to record |
|---|---|
| Run ID | A unique ID per answer |
| Prompt | The prompt ID and exact text |
| Assistant and mode | For example, ChatGPT on the web with search on |
| Model | The model label shown, if any |
| Date and time | With time zone |
| Location and language | For example, United States, English |
| Account state | Logged out, clean account, or signed in |
| Answer | The full text, saved |
| Mentions | Your company and each competitor, with the matched text |
| Citations | Every cited URL, saved as shown |
4. Set the number of runs and the time window
In one daily panel, the standard error of a brand's detection rate was 0.370 with one run per prompt, 0.081 with seven, and 0.062 with eight (opens in a new tab) (preprint). A standard error is the typical size of the error in an estimate. The authors call a single run "essentially uninformative" (opens in a new tab) and recommend a two-to-four-week rolling window, which brought the standard error to 0.080 at 14 days, 0.053 at 21 days, and 0.033 at 28 days (opens in a new tab). The review of 45 studies warns that seven to eight runs is not a universal standard; it recommends repeating the measurement "until the interval around the estimand is sufficiently narrow for the decision at hand." (opens in a new tab) Another panel found that mention rates need at least three runs to settle for combinations that are neither always nor never mentioned (opens in a new tab) (preprint).
5. Add prompts before you add runs
Answers to the same prompt are correlated, so repeating a prompt helps less than adding a new one. Re-running a prompt shrinks only the variance within that prompt, and clustered standard errors can be over three times larger than naive ones (opens in a new tab) (preprint). In the same paper, the clustered standard error works as a "sliding scale": if answers within a group are perfectly correlated, each group counts as a single independent observation (opens in a new tab). The table below shows the effect (AEO HQ calculation, standard design-effect formula, assuming a true mention rate of 10% and a correlation of 0.6 between runs of the same prompt). The research behind this page found no published estimate of that correlation for brand mentions, so treat 0.6 as an assumption and estimate it from your own first weeks of data.
| Prompts × runs per prompt | Total runs | Approximate 95% error range |
|---|---|---|
| 120 × 3 | 360 | ±4.6 points |
| 120 × 8 | 960 | ±4.3 points |
| 240 × 3 | 720 | ±3.3 points |
| 60 × 8 | 480 | ±6.1 points |
Running each of 120 prompts eight times instead of three nearly triples the work and narrows the range by 0.3 points. Doubling the prompts to 240, at three runs each, narrows it by 1.3 points.
6. Fix the matching rules before you look at the results
List the exact strings that count as a mention: the company name and its common variants, the domain, product names, and named founders. Have a person review matches for any name that is also a common word or is shared with another company. For citations, decide how URLs are grouped, because, as the review notes, parameters, redirects, anchors, translated pages, and aggregators can "artificially fragment a domain" (opens in a new tab). Keep every run in the denominator, including answers with no search and no citations.
7. Grade accuracy against a fact sheet
Write down the facts buyers ask about most: what you sell, prices, who it is for, locations, and founders. Grade each branded answer against that sheet, and when an answer is wrong, record the page it cited, because that page is the one to fix. The evidence that assistants misattribute facts is covered under "Counting mentions without checking what the answer says" in AEO antipatterns.
8. Compute each rate with a small-sample interval
The mention rate is the number of runs that mention you divided by all runs. Do not use the usual normal-approximation interval on small samples: nominal 95% intervals of that kind covered the true value only 92.5% of the time at 100 data points, and the authors recommend Wilson score or Bayesian intervals (opens in a new tab) (peer-reviewed, ICML 2025). The same paper found that with clustered data in small samples, neither the simple nor the clustered normal-approximation interval gave correct coverage (opens in a new tab). Two worked examples (AEO HQ calculation):
- 18 of 120 runs mention you. The rate is 15.0%, and the 95% Wilson interval is 9.7% to 22.5%.
- 0 of 60 runs mention you. The true rate is not proven to be zero. The one-sided 95% upper bound is about 4.9% (1 − 0.05^(1/60)), close to the "rule of three" shortcut of 3 ÷ 60 = 5%.
Both examples treat runs as independent. Because runs of the same prompt are correlated, the real intervals are wider. For an interval that allows for this, resample whole prompts, then runs within them (a cluster bootstrap), as described in AEO HQ's statistics section.
9. Compare periods with an interval for the difference
Suppose you were named in 18 of 120 runs in one four-week window (15.0%) and 30 of 120 in the next (25.0%). The 95% interval for the difference, by Newcombe's score method, runs from −0.2 to +20.0 points (AEO HQ calculation). It includes zero, so a 10-point rise on this sample is not yet a demonstrated change, and correlation between runs would widen the interval further. Report a change only when the interval for the difference excludes zero, and mark model releases and product changes on the timeline.
10. Add the first-party reports and outcome data
These show outcomes that a prompt panel cannot, and each has its own gaps:
- Google Search Console. Its generative AI performance report counts impressions of your links in AI Overviews and AI Mode, not clicks, may not appear until a site has enough impressions, and has been available worldwide since August 31, 2026 (opens in a new tab) (official documentation). Clicks from AI Overviews and AI Mode are counted in the Performance report under the Web search type and are not shown separately (opens in a new tab).
- Bing Webmaster Tools. Its AI Performance report shows total citations, average cited pages, grounding queries, and page-level citations across Microsoft Copilot, Bing's AI summaries, and "select partner integrations," with no click data (opens in a new tab). Grounding queries are the key phrases the AI used when it retrieved the content it cited, shown as a sample (opens in a new tab). In June 2026 Microsoft added Intents, Topics, Compare, and Citation Share, "the percentage of citations attributed to your site out of all citations shown across all sites for that same grounding query" (opens in a new tab). Citation Share does not show competitor domains or traffic share (opens in a new tab) (official documentation).
- Google Analytics 4. The default AI Assistant channel counts visits "from sources like ChatGPT, Gemini, Deepseek, Copilot, or Grok" and excludes AI Overviews and AI Mode, which count as Organic Search (opens in a new tab) (official documentation). Setup, and what GA4 cannot see, are covered in how to track AI referral traffic in GA4. The numbers are small: 0.17% of the average site's visitors came from AI chatbots (opens in a new tab) in a February 2025 study of 3,000 sites, which its author calls the "visible" AI traffic, "the minimum amount," (opens in a new tab) because some AI visits are recorded as direct (vendor study).
- Server logs. Two kinds of agent that OpenAI, Anthropic, and Perplexity run matter here. Search crawlers build the index their assistants search: OpenAI says OAI-SearchBot "is used to surface websites in search results in ChatGPT's search features" (opens in a new tab), and Anthropic and Perplexity describe Claude-SearchBot (opens in a new tab) and PerplexityBot (opens in a new tab) the same way. User-initiated fetchers visit a page while answering a person's question: "When users ask ChatGPT or a CustomGPT a question, it may visit a web page with a ChatGPT-User agent" (opens in a new tab), and Claude-User (opens in a new tab) and Perplexity-User (opens in a new tab) do the same for their assistants. The glossary explains the difference between a crawler and a user-initiated fetcher and lists each company's agents under OpenAI crawlers, Anthropic crawlers, and Perplexity crawlers. A fetch shows that an answer retrieved your page, not whether it cited or recommended you. These fetches do not appear in analytics: in Vercel's December 2024 data, none of the major AI crawlers rendered JavaScript, including OpenAI's ChatGPT-User (opens in a new tab).
- What buyers tell you. Ask "Where did you first hear about us?" at sign-up or checkout. In one agency's records, 189 of 213 leads named an AI tool, but first-touch attribution credited AI with only 28 of those 189 (15%) (opens in a new tab) (single firm; the agency sells AEO services). Ahrefs reported more than 14,000 new users who said ChatGPT sent them (opens in a new tab), collected through a feedback question at registration (vendor data).
11. Report each assistant separately, with the method attached
A report that others can check has these columns for each assistant and window:
| Column | What goes in it |
|---|---|
| Assistant and mode | For example, Perplexity, web, default settings |
| Window | Start and end dates |
| Prompts and runs | How many of each |
| Mention rate | The rate, the count behind it, and the 95% interval |
| Citation rate | Same format |
| Share of voice | With the competitor set named |
| Accuracy | Share of branded answers graded correct |
| Changes and events | Model releases, panel changes, site changes |
Instruments compared
| Instrument | What it shows | What it misses |
|---|---|---|
| Prompt panel, run by hand or by a tool | Mention rate, citation rate, share of voice, and accuracy for each assistant | Questions you did not include; sampling error |
| Search Console generative AI performance report | Impressions of your links in AI Overviews and AI Mode (opens in a new tab) | Clicks, which sit inside the Web totals; other assistants; answers that mention you without a link |
| Bing Webmaster Tools AI Performance | Citations, cited pages, and a sample of grounding queries for Copilot and Bing's AI summaries (opens in a new tab) | Clicks; competitor domains (opens in a new tab); answers from assistants outside Microsoft's products and its unnamed partner integrations |
| Google Analytics 4 | Visits and key events from links in assistant answers | Answers without clicks; AI Overviews and AI Mode, which count as Organic Search (opens in a new tab); app traffic that carries no referrer (opens in a new tab) |
| Microsoft Clarity AI channel | Sessions from the standalone sites of Claude, Copilot, Gemini, ChatGPT, and Perplexity (opens in a new tab) | Hidden sources, which "might appear as Direct," and AI features embedded in search engines or productivity tools (opens in a new tab) |
| Server logs | Fetches by index crawlers and user-initiated fetchers, by URL | Whether the answer used or cited the page |
| "Where did you hear about us?" | Influence that left no click | Recall errors; people who skip the question |
Choosing a tracking tool
Tracking tools automate the prompt panel. AEO and AI visibility tools compared reviews them. Whatever the tool, these questions separate a measurement from a guess. Each follows from the evidence above.
- Does it query the consumer interfaces or only APIs? APIs shared only 12.0% to 14.8% of cited domains with their own interfaces (opens in a new tab).
- How many runs per prompt, and does it show intervals? One run per prompt left a standard error of 0.370 (opens in a new tab).
- How many prompts and phrasings, and who wrote them? People word the same need very differently: 142 human-written prompts for one intent had a mean semantic similarity of 0.081 (opens in a new tab), although in that study they still produced similar brand sets. Prompts drawn from real buyer language are easier to defend than prompts a model invented.
- Which denominator does it use? A share of citations counted only among answers that cite something leaves out the answers where the assistant did not search (opens in a new tab).
- Does it report a rank or position? Two lists in the same order appeared about once in 1,000 runs (opens in a new tab).
- How is any combined score weighted? A combined score needs weights tied to a stated objective (opens in a new tab).
- Can you export the raw answers? Without them, you cannot check the matching or the accuracy grading.
Much of the published research on measurement comes from companies that sell tracking. The disclosures are given with each study on this page.
What the evidence shows and does not show
Antipatterns
An antipattern is a practice that looks helpful but fails or backfires. AEO antipatterns covers four measurement mistakes in depth: treating one screenshot as a measurement, reporting a rank or an undisclosed score, counting mentions without checking accuracy, and judging by referral traffic alone. Six more come up when running a panel:
- Dropping answers without citations from the denominator. It overstates visibility, because the chance of being cited includes the chance that the assistant searched at all (opens in a new tab). Instead, count every run.
- Pooling all assistants into one number. Two consumer interfaces shared only 5.4% of cited domains (opens in a new tab), so a pooled rate describes none of them. Instead, report each assistant separately.
- Adding runs instead of prompts. Repeated runs reduce only the variance within a prompt (opens in a new tab). Instead, add prompts and phrasings first.
- Changing the prompt set mid-series. Rates differ by the kind of question asked (opens in a new tab), so a new panel can move the number on its own. Instead, freeze the panel, and run old and new versions side by side for one window when you change it.
- Reading a change from two small windows. A 10-point rise on 120 runs can sit inside the error range, as in step 9. Instead, calculate the interval for the difference.
- Using normal-approximation intervals on small samples. They cover the true value less often than stated below a few hundred data points (opens in a new tab). Instead, use Wilson or Bayesian intervals.
Checklist
| # | Check | How to verify | Pass when | Basis |
|---|---|---|---|---|
| 1 | Measures are defined separately | Read the report | Mention rate, citation rate, share of voice, and accuracy are separate columns | Visibility vector (opens in a new tab) |
| 2 | The panel uses buyer wording | Compare prompts with sales calls and search queries | Each intent has 3–5 phrasings | Review of 45 studies (opens in a new tab) |
| 3 | The panel is frozen | Compare prompt lists across windows | Unchanged, or run in parallel when changed | Question type changes rates (opens in a new tab) |
| 4 | Consumer interfaces are tested | Read the run log | Interface runs exist for each assistant reported | API and interface overlap (opens in a new tab) |
| 5 | Sessions are controlled | Read the run log | New chat per prompt; account state recorded | Identity effects (opens in a new tab) |
| 6 | Each run records its conditions | Read the run log | Product, mode, model, date, locale, account, and search status are filled | Minimum checklist (opens in a new tab) |
| 7 | Enough runs | Count runs per prompt per window | At least 3 per prompt; more for key prompts | Run counts (opens in a new tab) |
| 8 | Every run is in the denominator | Recount from raw answers | Answers without search or citations are included | Decomposition (opens in a new tab) |
| 9 | Intervals suit small samples | Read the method | Wilson, Bayesian, or cluster bootstrap | Coverage study (opens in a new tab) |
| 10 | Assistants are reported separately | Read the report | One row per assistant | Interface overlap (opens in a new tab) |
| 11 | First-party reports are connected | Open Search Console and Bing Webmaster Tools | Both properties are verified and reviewed each window | Search Console (opens in a new tab); Bing (opens in a new tab) |
| 12 | Self-reported source is collected | Test the sign-up or checkout form | The question is asked and stored | Attribution gap (opens in a new tab) |
FAQ
What is AI visibility?
It is how often AI assistants mention your company, cite your pages, and describe you correctly in their answers. Microsoft's Bing team uses the term in the same sense for its "AI Visibility Insights" in Bing Webmaster Tools (opens in a new tab).
How do I check my brand's AI visibility without buying a tool?
Write 20 to 40 buyer questions with three phrasings each, run each one at least three times on each assistant in a clean session, and record the answers in a spreadsheet using the run log above. Count the share of runs that name you, and calculate a Wilson interval. Add Search Console's generative AI performance report and Bing Webmaster Tools' AI Performance report. This is slow, but it is the same method a sound tool uses.
What is an AI visibility score?
There is no standard one. Each vendor defines its own. A survey of 45 studies says a combined score is defensible only when its weights match an explicit objective (opens in a new tab). Ask how the score is calculated and which runs it counts, or use the separate rates on this page.
How many times should I run each prompt?
At least three times, and seven or eight for the prompts that matter most. Seven runs brought the standard error of a brand's detection rate to 0.081 in one panel (opens in a new tab) (preprint). If you have a fixed budget, spend it on more prompts before more runs.
How do I compare my AI visibility with competitors'?
Name five to ten competitors before you start, count their mentions in the same runs, and report share of voice alongside your own mention rate. Bing's Citation Share compares your citations with all citations for a grounding query, but it does not show competitor domains (opens in a new tab).
Do brand mentions affect AI visibility?
They are strongly associated with it. Across 75,000 brands, branded web mentions correlated with AI visibility at about 0.66 to 0.71 and YouTube mentions at about 0.74, while link metrics such as the number of backlinks showed very weak correlations (opens in a new tab) (vendor study; correlation, not cause). How to earn them is covered in brand mentions and AI recommendations.
Why doesn't my AI visibility show up in Google Analytics?
Analytics sees visits, not answers. An answer that names you without a click leaves no trace: Google users who saw an AI summary clicked a link in the summary itself in just 1% of visits (opens in a new tab) (Pew Research Center; 900 U.S. adults, March 2025). Some clicks arrive without a referrer: traffic referred by Claude's native app does not include a Referer header (opens in a new tab) (network measurement). The full list is in what GA4 cannot see.
How long does it take for AI visibility to change?
No study has measured how long a change takes to show up. In a large observational panel, the fastest-moving small brands gained 10 to 20 points over a March to May 2026 window (opens in a new tab) (preprint), and a review of 45 studies found no technique with "a stable, longitudinal, cross-platform causal effect on organic discoverability" (opens in a new tab) (preprint). For how assistants choose sources in the first place, see how AI assistants find and cite sources.
Next steps
AEO HQ's Instant AEO Audit ($499) is an automated check of a site's crawler access, key pages, and structured data, with a five-question sample from one AI model. It is not the repeated, multi-assistant measurement described on this page. The methodology page lists what it measures and what it does not.
Change log
- September 28, 2026: First published.
Sources
- Fishkin, R. (2026, January 28). NEW research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility. SparkToro. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/ (opens in a new tab)
- Madhavan, K., Merchant, M., Canel, F., & Nigam, S. (2026, February 10). Introducing AI Performance in Bing Webmaster Tools public preview. Bing Webmaster Blog. https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview (opens in a new tab)
- Martinez, O. (2026). Optimizing visibility in generative engines: A critical survey of generative engine optimization (2023–2026) (arXiv:2607.14035). arXiv. https://doi.org/10.48550/arXiv.2607.14035 (opens in a new tab)
- Harsel, L. (2026, April 7). ChatGPT traffic analysis: Insights from 17 months of clickstream data. Semrush. https://www.semrush.com/blog/chatgpt-search-insights/ (opens in a new tab)
- Schulte, J., Bleeker, M., & Kaufmann, P. (2026). Don't measure once: Measuring visibility in AI search (GEO) (arXiv:2604.07585). arXiv. https://doi.org/10.48550/arXiv.2604.07585 (opens in a new tab)
- Makhortykh, M., Sydorova, M., Baghumyan, A., Vziatysheva, V., & Kuznetsova, E. (2024). Stochastic lies: How LLM-powered chatbots deal with Russian disinformation about the war in Ukraine. Harvard Kennedy School (HKS) Misinformation Review. https://doi.org/10.37016/mr-2020-154 (opens in a new tab)
- Kumar, P. (2026). Generative engine optimization at scale: Measuring brand visibility across AI search engines (arXiv:2606.20065). arXiv. https://arxiv.org/abs/2606.20065 (opens in a new tab)
- Mizrahi, M., Kaplan, G., Malkin, D., Dror, R., Shahaf, D., & Stanovsky, G. (2024). State of what art? A call for multi-prompt LLM evaluation. Transactions of the Association for Computational Linguistics, 12, 933–949. https://doi.org/10.1162/tacl_a_00681 (opens in a new tab)
- Uberti-Bona Marin, L. G., Bertaglia, T., Astante, G., Rijsbosch, B., van Dijck, G., Hannák, A., Spanakis, G., & Kollnig, K. (2026). "If I had to buy just ONE: Galaxy S26 Ultra": Auditing AI-generated product recommendations (arXiv:2609.18729). arXiv. https://doi.org/10.48550/arXiv.2609.18729 (opens in a new tab)
- Kantharuban, A., Milbauer, J., Sap, M., Strubell, E., & Neubig, G. (2025). Stereotype or personalization? User identity biases chatbot recommendations. In Findings of the Association for Computational Linguistics: ACL 2025 (pp. 24418–24436). Association for Computational Linguistics. https://aclanthology.org/2025.findings-acl.1254/ (opens in a new tab)
- Chen, L., Zaharia, M., & Zou, J. (2024). How is ChatGPT's behavior changing over time? Harvard Data Science Review, 6(2). https://doi.org/10.1162/99608f92.5317da47 (opens in a new tab)
- Miller, E. (2024). Adding error bars to evals: A statistical approach to language model evaluations (arXiv:2411.00640). arXiv. https://doi.org/10.48550/arXiv.2411.00640 (opens in a new tab)
- Bowyer, S., Aitchison, L., & Ivanova, D. R. (2025). Position: Don't use the CLT in LLM evals with fewer than a few hundred datapoints. In Proceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267, pp. 81143–81184). PMLR. https://proceedings.mlr.press/v267/bowyer25a.html (opens in a new tab)
- Google. (2026). Generative AI performance report (Search) [Search Console Help]. Retrieved September 27, 2026, from https://support.google.com/webmasters/answer/16984139 (opens in a new tab)
- Google. (2025, December 10). AI features and your website. Google Search Central. https://developers.google.com/search/docs/appearance/ai-features (opens in a new tab)
- Madhavan, K., Merchant, M., Nigam, S., & Shah, T. (2026, June 16). New AI visibility insights in Bing Webmaster Tools: Intents, topics, citation share, compare. Bing Search Blog. https://blogs.bing.com/search/June-2026/New-AI-Visibility-Insights-in-Bing-Webmaster-Tools-Intents-Topics-Citation-Share-Compare (opens in a new tab)
- Google. (2026). Default channel group [Analytics Help]. Retrieved September 27, 2026, from https://support.google.com/analytics/answer/9756891 (opens in a new tab)
- Linehan, L. (2025, February 6). 63% of websites receive AI traffic (new study of 3,000 sites). Ahrefs. https://ahrefs.com/blog/ai-traffic-study/ (opens in a new tab)
- OpenAI. (n.d.). Overview of OpenAI crawlers. OpenAI API documentation. Retrieved September 27, 2026, from https://developers.openai.com/api/docs/bots (opens in a new tab)
- Anthropic. (2026, April 7). Does Anthropic crawl data from the web, and how can site owners block the crawler? Claude Help Center. https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler (opens in a new tab)
- Perplexity. (n.d.). Perplexity crawlers. Perplexity Docs. Retrieved September 27, 2026, from https://docs.perplexity.ai/guides/bots (opens in a new tab)
- Zecchini, G., Moore, A. A., Ubl, M., & Siddle, R. (2024, December 17). The rise of the AI crawler. Vercel. https://vercel.com/blog/the-rise-of-the-ai-crawler (opens in a new tab)
- Birkett, A. (2026, August 28). First-touch attribution captures 15% of our AI-sourced leads [Research]. Omniscient Digital. https://beomniscient.com/blog/first-touch-vs-self-reported-attribution-aeo/ (opens in a new tab)
- Belson, D., & Rhea, S. (2025, July 1). The crawl before the fall... of referrals: Understanding AI's impact on content providers. Cloudflare Blog. https://blog.cloudflare.com/ai-search-crawl-refer-ratio-on-radar/ (opens in a new tab)
- Microsoft. (2025, September 23). AIPlatform and PaidAIPlatform. Microsoft Learn (Clarity). https://learn.microsoft.com/en-us/clarity/insights/ai-channel-group (opens in a new tab)
- Linehan, L. (2025, December 12). Top brand visibility factors in ChatGPT, AI Mode, and AI Overviews (75k brands studied). Ahrefs. https://ahrefs.com/blog/ai-brand-visibility-correlations/ (opens in a new tab)
- Chapekis, A., & Lieb, A. (2025, July 22). Google users are less likely to click on links when an AI summary appears in the results. Pew Research Center. https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/ (opens in a new tab)
How to cite this page
Maxwell, P. (2026). How to measure AI visibility. AEO HQ. Last updated September 28, 2026. https://www.aeohq.ai/articles/how-to-measure-ai-visibility
Pages in Measuring AI visibility
Guide
How to track AI referral traffic in GA4
Track AI referral traffic in GA4: what the AI Assistant channel counts, a custom channel for ChatGPT, Claude, Perplexity, and others, and what GA4 cannot see.
Reference
AEO and AI visibility tools compared
AEO and AI visibility tools by kind: what search engine reports, analytics, prompt trackers, graders, and log tools can and cannot measure, from vendors' pages.
Reference
AEO metrics and KPIs: definitions
AEO and GEO metrics and KPIs defined: mention rate, citation rate, share of voice, accuracy, AI impressions, and AI referrals, with formulas and limits.
Reference
How AEO HQ measures AI visibility, and how our research is done
How AEO HQ grades evidence, pulls search data, and measures AI visibility, and what its $499 automated audit does and does not measure today.