Antipatterns · Answer engine optimization (AEO)
Answer engine optimization antipatterns
Thirteen common answer engine optimization mistakes, why each one fails according to research and platform documentation, and how to detect it.
By Paul Maxwell, founder of AEO HQ
Published · Updated
An answer engine optimization antipattern is a common practice that is meant to get a brand cited or recommended in AI answers but does not work, cannot be measured, or breaks a platform's rules. This page describes 13 of them, covering measurement, claims, third-party corroboration, and page content. Each entry gives the evidence for why the practice fails and a check you can run to detect it.
Scope and definitions
Answer engine optimization (AEO) is the work of making a brand's pages and facts easy for AI assistants to find, trust, and cite. An answer engine is a system that replies to a question with a generated answer, usually with links to its sources. Examples are ChatGPT, Perplexity, Claude, Microsoft Copilot, Gemini, and Google's AI Overviews and AI Mode. The method itself is explained in the complete guide. This page covers what to avoid.
Four terms recur below:
- Retrieval is the step in which an answer engine finds candidate pages in a search index.
- Grounding is the step in which it bases its answer on the pages it retrieved.
- A citation is a link to a source shown with an answer.
- A mention is the brand's name appearing in an answer, with or without a link.
AEO overlaps with generative engine optimization (GEO), and the two terms describe much of the same work. Two companion pages cover the neighboring mistakes. Generative engine optimization antipatterns covers tactics aimed at how models read and rewrite pages, such as hidden prompts and mass rewriting. SEO antipatterns covers technical search mistakes that keep pages out of AI answers, such as blocked crawlers and facts that load only with JavaScript.
How to read this page
Each entry has six parts: what the antipattern looks like, why people do it, why it fails, what to do instead, how to detect it, and its sources. "Why it fails" reports evidence. "What to do instead" is AEO HQ's recommendation.
Every fact links to its primary source. The label after a fact gives the type of source:
- (peer-reviewed): a journal article or conference paper that passed peer review. "Accepted" means accepted for publication but not yet published.
- (preprint): a research paper, working paper, or thesis that has not been peer reviewed.
- (vendor study) or (vendor blog): research or guidance published by a company that sells a related product. Read it as descriptive.
- (official documentation): what a platform states about its own systems. It is authoritative about policy, not about the size of effects.
- Other labels, such as (industry survey), (journalism), (network measurement), and (practitioner test), mark published sources that were not peer reviewed and often use small or self-selected samples.
Each entry also rates the overall evidence as strong (official documentation, or several independent studies that agree), moderate (consistent evidence limited to lab settings, a few studies, or correlations), or weak (one small or conflicted source).
The 13 antipatterns at a glance
| # | Antipattern | Do this instead | Evidence |
|---|---|---|---|
| 1 | Treating one screenshot as a measurement | Sample each question many times, in clean sessions | Strong |
| 2 | Reporting a "rank in ChatGPT" or an undisclosed score | Report mention rates with intervals | Strong |
| 3 | Counting mentions without checking what the answer says | Grade answers against a fact sheet | Strong |
| 4 | Judging AEO by AI referrals alone | Combine referrals, self-reported source, and citation reports | Strong |
| 5 | Promising guaranteed citations, rankings, or timelines | Promise work and measurement, not outcomes | Strong |
| 6 | Self-ranking "best agencies" lists | Publish comparisons with a method, without ranking yourself first | Weak to moderate |
| 7 | Paid or fake reviews, testimonials, and community posts | Collect real reviews and disclose incentives | Strong |
| 8 | Optimizing only your own site | Earn third-party coverage and reviews | Moderate |
| 9 | Inconsistent facts across your site and profiles | Keep one fact sheet and match every source to it | Moderate |
| 10 | Hiding prices and scope behind "contact us" | Publish price, scope, and turnaround in plain text | Moderate |
| 11 | Burying the answer | Answer in the first one to three sentences | Moderate |
| 12 | Relying on FAQ markup after May 7, 2026 | Keep visible FAQs for readers; treat markup as optional | Strong |
| 13 | Over-investing in schema as a citation lever | Keep markup as hygiene that matches the visible page | Moderate |
1. Treating one screenshot as a measurement
What it looks like. Someone asks ChatGPT "Who are the best [category] consultants?" once, takes a screenshot, and reports that the brand does or does not appear. Sales decks use the same kind of evidence. A variant runs each prompt once through an API and presents the output as what buyers see.
Why people do it. A single answer looks like a fact. It is fast, free, and easy to share.
Why it fails. AI answers vary between runs, between sessions, and between the API and the app, so one answer is a sample of one.
- When the same prompt was repeated, ChatGPT and Google's AI returned the same list of brands less than once in 100 runs, and Claude only slightly more often (opens in a new tab), in a study of 2,961 runs (vendor study; one co-investigator works for a visibility-tracking vendor).
- In a daily panel of German-language prompts across four assistants (preprint; the lead author is affiliated with a tracking vendor), about 65% of cited sources turned over from one day to the next, and the standard error of a brand's detection rate fell from 0.370 with one run to 0.081 with seven runs (opens in a new tab).
- In an audit of product recommendations (preprint), the API and the consumer app of the same assistant shared only 12.0% (ChatGPT) and 14.8% (Gemini) of cited domains (opens in a new tab). The authors concluded that "neither isolated responses nor API observations can be assumed to represent the commercial advice consumers encounter" (opens in a new tab).
- Account history changes answers too. In a six-run test by an agency (practitioner test), a logged-in ChatGPT account that had discussed one agency named that agency first, while the same prompt in a temporary chat did not name it (opens in a new tab).
Evidence: strong that answers vary. The exact figures differ by assistant and date.
What to do instead. Treat each buyer question as a distribution to sample. Use several phrasings of each question, repeated runs, clean sessions, and a separate result for each assistant, pooled over a rolling window. A survey of 45 studies (preprint) recommends 7 to 8 repetitions, 3 to 5 paraphrases, several engines, and several time windows (opens in a new tab) as a starting design. The guide to measuring AI visibility sets out a full method.
How to detect it. Ask for the raw log behind any visibility claim. It passes if, for each prompt, the log records the assistant and mode, the date and time, the session state, the number of runs (at least seven), and an interval around the result. It fails if the claim rests on one run per prompt, or on API output presented as what buyers see.
Sources: 1, 2, 3, 4, 5.
2. Reporting a "rank in ChatGPT" or an undisclosed visibility score
What it looks like. Reports that say "you rank #3 in ChatGPT for 'best CRM consultant'," dashboards that plot a brand's position day by day, or a single proprietary visibility score with no published method.
Why people do it. Rank is the familiar SEO metric, and one number is easy to put in a report.
Why it fails. The order of brands in AI answers is too unstable to rank. Whether a brand is mentioned at all is much steadier.
- In the same repeated-prompt study, it took about 1 in 1,000 runs to see two lists in the same order (opens in a new tab) (vendor study).
- ChatGPT's recommended products overlapped by a mean Jaccard similarity of 0.178 across three repeats of the same query (opens in a new tab) (preprint). Jaccard similarity is the share of items two lists have in common, so 0.178 means most items changed between repeats.
- Across 102,025 API responses (preprint; the author co-founded the tracking platform studied), 77.5% of brand, prompt, and engine combinations were either always or never mentioned, and mention flips were 6.7 times rarer than sentiment flips (opens in a new tab).
- A software company that sells its own AEO product warns buyers that "a proprietary 'AI visibility score' that only ever climbs, with no disclosed methodology, is easy to sell and hard to verify" (opens in a new tab) (vendor blog).
Evidence: strong that rank is unstable. Moderate that mention rates are the better measure.
What to do instead. Report a mention rate: the share of runs, on a fixed set of buyer prompts, in which the brand is named, for each assistant, with a confidence interval. Use a Wilson or Bayesian interval, because the usual normal-approximation interval is too narrow when there are fewer than a few hundred data points (opens in a new tab) (peer-reviewed). To compare against competitors, report share of voice: your mentions divided by all mentions of a fixed list of competitors.
How to detect it. Search the report for "rank," "position," "#1," or a single score. It passes if every visibility number is a rate with its number of runs, an interval, a date range, and a published method. It fails if a position or a score appears without them.
Sources: 1, 3, 6, 7, 8.
3. Counting mentions without checking what the answer says
What it looks like. Tracking whether the brand is named or linked, but never reading what the assistant says about it: its services, prices, founder, or the claims it attributes to the brand.
Why people do it. Mentions can be counted automatically. Accuracy needs a person to read and grade answers.
Why it fails. Being cited is not the same as being described correctly.
- In a study of 2,709 assistant answers about the news, run by 22 public service media organizations (industry study), 45% of answers had at least one significant issue, and 31% had significant sourcing problems (opens in a new tab).
- In an audit of Google AI Overviews (preprint), 11.0% of 98,020 claims were not supported by the pages cited for them (opens in a new tab).
- When eight AI search tools were asked to identify the source of news excerpts (journalism), they answered more than 60% of 1,600 queries incorrectly (opens in a new tab).
- An audit of four generative search engines (peer-reviewed) measured citation recall of 51.5% and citation precision of 74.5% (opens in a new tab). Recall is the share of statements fully supported by their citations. Precision is the share of citations that support their statement.
These studies used news and general questions. No study has measured misattribution for B2B vendor questions.
Evidence: strong that misattribution is common. Untested for B2B vendor questions.
What to do instead. Add an accuracy check. Run a fixed set of branded prompts, such as "What does [brand] do?", "What does [brand] cost?", and "Who founded [brand]?", and grade each answer against a written fact sheet. When an answer is wrong, find and fix its source: your page, your profile, or the third-party page the assistant cited.
How to detect it. Look for an accuracy measure in the reporting. It passes if at least one metric counts answers that state the brand's facts correctly, graded by a person. It fails if every metric is a count of mentions or citations.
Sources: 9, 10, 11, 12.
4. Judging AEO by AI referral traffic alone
What it looks like. Declaring AEO a success or a failure from AI referral traffic alone, such as the "AI Assistant" channel in Google Analytics 4 (GA4) or visits from chatgpt.com.
Why people do it. Referrals are the one AI signal that most analytics tools report by default.
Why it fails. Referral counts are small, and they miss much of the influence.
- Across 3,000 sites in early 2025 (vendor study), 0.17% of the average site's traffic came from AI chatbots (opens in a new tab).
- Traffic referred by the Claude app carries no Referer header, and Cloudflare believes the same holds for other native apps (opens in a new tab) (network measurement). Analytics tools cannot attribute such visits to an assistant.
- GA4's AI Assistant channel "excludes Google's AI Overviews and AI Mode" (opens in a new tab), which GA4 counts as organic search (official documentation).
- Google Search Console counts traffic from AI features inside the "Web" search type of its Performance report, together with other search traffic (opens in a new tab) (official documentation).
- In one agency's own records, first-touch attribution credited AI with only 28 of the 189 leads (15%) who named an AI tool when asked how they found the agency (opens in a new tab) (single-firm data; the agency sells AEO services).
Evidence: strong that referrals undercount AI influence (tool documentation). Weak on the size of the gap.
What to do instead. Combine sources and treat referrals as a floor. Count referrals, including the utm_source=chatgpt.com parameter that ChatGPT adds to referral URLs (opens in a new tab). Add a "Where did you hear about us?" question at signup or checkout, the AI Performance report in Bing Webmaster Tools (opens in a new tab), the generative AI performance report in Search Console (opens in a new tab), and a repeated prompt panel.
How to detect it. Read the measurement plan. It passes if it combines at least three of these: referral data, self-reported source, a first-party citation or impression report, and a prompt panel. It fails if referral traffic is the only success metric.
Sources: 13, 14, 15, 16, 17, 18, 19, 20.
5. Promising guaranteed citations, rankings, or timelines
What it looks like. Offers such as "guaranteed ChatGPT citations in 30 days" or "rank #1 in AI Overviews," or a contract that promises appearance in named assistants by a date.
Why people do it. Guarantees close deals, and buyers want certainty.
Why it fails. No platform sells or promises placement, and no study supports a timeline.
- Google warns about services "promising improvements for AI experiences and search formats (also known as 'AEO' or 'GEO' tools)" (opens in a new tab) and states that third-party tools "can't guarantee performance" (opens in a new tab) (official documentation).
- OpenAI says of ChatGPT search that "placement is not guaranteed" (opens in a new tab) (official documentation).
- Bing's guidelines state that "GEO does not guarantee grounding or citations" (opens in a new tab) (official documentation).
- Google says it "doesn't guarantee that it will crawl, index, or serve your page" (opens in a new tab) at all (official documentation).
- On timing, the fastest-moving small brands in a 102-brand panel gained only 10 to 20 percentage points of visibility over a March to May 2026 tracking window (opens in a new tab) (preprint). The 45-study survey found no reviewed technique with a stable, longitudinal, cross-platform causal effect on organic discoverability (opens in a new tab) (preprint). No rigorous study has measured how long a new brand takes to be recommended.
Evidence: strong (each platform's own documentation).
What to do instead. Promise work, not outcomes: named deliverables, a stated method, a measurement plan with a baseline, and dates for the work itself. State the limit in writing. The guide to choosing an agency lists questions to ask a vendor.
How to detect it. Search the proposal, contract, and sales pages for "guarantee," "#1," "will appear," "within [number] days," and percentage lifts. It passes if outcome language is limited to what will be measured and how. It fails if any placement in an assistant is promised. For percentage lifts, see the "+40%" entry in Generative engine optimization antipatterns.
Sources: 5, 6, 21, 22, 23, 24.
6. Self-ranking "best agencies" lists
What it looks like. A vendor publishes "The 10 best AEO agencies in 2026" on its own site and ranks itself first. Related tactics are buying a sponsored slot in someone else's list and paying for a guest post that ranks the buyer.
Why people do it. Lists are what AI assistants read for these questions, and the tactic appears to work.
- In ChatGPT answers to 750 prompts, including prompts that asked for agencies, "best X" lists made up 43.8% of cited page types (opens in a new tab) (vendor study).
- In a published test of 48 AI answers collected between August 20 and September 2, 2026, the most-named agency appeared in 58% of answers, the second-most-named in 40%, and the third in 31% (opens in a new tab). The tester reported that "almost every high-frequency source is an agency's own best-AEO-or-GEO listicle that ranks itself first" (opens in a new tab), and that one agency's site was the source for 14 of the 48 answers (opens in a new tab) (practitioner test by an agency).
Why it fails. The advantage is visible, fragile, and easy to discount.
- In the same test, "ChatGPT flagged the self-ranking in several answers" (opens in a new tab).
- A practitioner analysis linked self-promotional listicles to search visibility drops at several brands in January 2026 (opens in a new tab). It is one analyst's reading, not a statement from Google.
- 35% of the lists that ChatGPT cited sat on low-authority domains (opens in a new tab), and the study's author expects citations of such lists to decline (vendor study).
- For bought placement, in a simulated shopping study (preprint), AI agents chose a product less often when it carried a "Sponsored" tag: 7.9% to 8.9% of the time, against a 10% baseline (opens in a new tab).
Evidence: weak to moderate (small published tests and practitioner reports; no controlled study).
What to do instead. If you publish a comparison, publish the scoring method, date each price and claim, disclose authorship, and leave yourself out of the ranking or score yourself separately. Seek inclusion in lists that publish their method and state that they take no payment.
How to detect it. Review every "best" list or comparison page you publish or pay for. It passes if the method is published, authorship and any payment are disclosed, and the publisher does not rank itself first. It fails if the publisher ranks itself first, or if placement was bought without disclosure.
Sources: 25, 26, 27, 28.
7. Paid or fake reviews, testimonials, and community posts
What it looks like. Buying reviews, offering undisclosed incentives for them, publishing testimonials from clients who do not exist, or posting as a satisfied customer on Reddit or Quora from accounts the company controls.
Why people do it. Ratings and endorsements move AI choices.
- In a study of four widely used AI models (preprint), ratings were the only promotional cue that consistently raised the chance of a product being chosen across models (opens in a new tab).
- In lab tests of three models (preprint), authority-style language, including fabricated claims, broke a well-known brand's hold on recommendations 55% to 99% of the time, depending on the model (opens in a new tab).
Why it fails.
- Google says that "seeking inauthentic 'mentions' across the web isn't as helpful as it might seem" (opens in a new tab) (official documentation).
- On July 24, 2026, Google added a guideline about fake and undisclosed incentivized reviews to its review snippet documentation (opens in a new tab). Since April 2026, Google may use spam report submissions to take manual action (opens in a new tab) (official documentation). Competitors can report you.
- Buyers check what they read: 94% of technology buyers who used AI said they fact-check its responses at least some of the time (opens in a new tab) (industry survey by a review platform).
Evidence: strong on platform policy (official documentation). Moderate on why the tactic tempts (lab studies).
What to do instead. Ask real clients for reviews on third-party platforms, follow each platform's rules on incentives, and publish no review counts or testimonials until they exist. Take part in communities under your own name and within their rules.
How to detect it. Audit every review, testimonial, and community post attributed to a customer. It passes if each one traces to a real, identifiable customer and any incentive is disclosed. It fails if any was written, bought, or posted by the company or its vendors without disclosure.
Sources: 29, 30, 31, 32, 33.
8. Optimizing only your own site
What it looks like. Every AEO task is on the brand's own pages: rewrites, markup, and FAQs. Nothing is done about the reviews, comparisons, lists, and discussions that third parties publish.
Why people do it. Your own site is the part you control, and on-site work produces visible deliverables.
Why it fails. Answer engines draw mostly on third-party pages, and third-party presence tracks visibility.
- In ranking prompts sent to AI search APIs (preprint), 93.5% to 95.1% of the domains that ChatGPT cited were earned media (opens in a new tab): reviews, comparisons, and editorial coverage rather than brand-owned pages.
- In a panel of 102 brands (preprint), only 2.9% of 149,912 citations pointed to the tracked brand's own domain, and ranked "best-of" lists made up about 21% of all citations (opens in a new tab).
- Among 112 startups from the 2025 Product Hunt leaderboard (preprint; a master's thesis), an on-page optimization score did not predict whether the startups were surfaced, while on Perplexity their discovery correlated with referring domains (r = 0.319) and with Reddit mentions (r = 0.395) (opens in a new tab).
- Across 75,000 brands (vendor study), branded web mentions correlated with AI visibility at about 0.66 to 0.71 and YouTube mentions at about 0.74, while link metrics such as the number of backlinks showed "very weak correlations" (opens in a new tab). These are correlations, and larger brands have more of everything.
- Established brands start far ahead: on a first answer to a generic prompt, household brands appeared 73% of the time, mid-market brands 44%, and small brands 11% (opens in a new tab) (preprint).
Evidence: moderate (the direction is consistent; the studies are mostly correlational, vendor, or preprint work).
What to do instead. Put part of the effort off-site: coverage in publications your buyers read, inclusion in lists and directories that publish their method, genuine participation in communities, and reviews from real clients. The guide to brand mentions covers how to earn them.
How to detect it. From your prompt panel, list the third-party pages cited for your top buyer prompts, and count how many mention your brand with correct facts. It passes if that count is tracked over time and the plan includes work aimed at raising it. It fails if the count is unknown or the plan has no off-site work.
Sources: 6, 34, 35, 36.
9. Inconsistent facts across your site and profiles
What it looks like. The pricing page shows one price and an old blog post another. The LinkedIn page describes a service the company no longer sells. The founder's bio differs between sites. The structured data still carries last year's offer.
Why people do it. No one owns the facts. Pages and profiles are edited at different times by different people.
Why it fails.
- In a controlled test of six AI models (peer-reviewed), making consistent rather than contradictory claims was one of seven factors that changed which source a model cited first, each significant in at least four of the six models (opens in a new tab).
- An entity is a thing a system identifies as one unit, such as a company, a person, or a product. In a benchmark built to test entities that share a name (peer-reviewed), models "often yield ambiguous answers or incorrectly merge information belonging to different entities" (opens in a new tab).
- Buyers check facts at the source. In an observational study of nine participants (usability study), people validated AI answers on vendors' own sites, and several did not fully trust prices quoted by AI (opens in a new tab).
Evidence: moderate.
What to do instead. Keep one written fact sheet with the legal name, product and service names, prices, founder, and locations, and make every page, profile, and markup file match it. Give the company and each named person one canonical page, and point profiles to it with sameAs links in the markup. The brand consistency checklist lists the places to check.
How to detect it. Write down the ten facts buyers ask about most. Compare each one against the homepage, pricing page, product pages, structured data, the llms.txt file if you have one, LinkedIn, and directory profiles. It passes if every source agrees. It fails on any mismatch, including an outdated price in an old post.
Sources: 37, 38, 39.
10. Hiding prices and scope behind "contact us"
What it looks like. Service pages that describe benefits but give no price, price range, scope, or turnaround, and send every question to a sales call.
Why people do it. Custom pricing keeps options open, and sales teams prefer to discuss price on a call.
Why it fails. Assistants and buyers look for exactly these facts, and models favor pages that state them.
- In the six-model test (peer-reviewed), a source that mentioned a price was significantly more likely to be cited first in all six models (opens in a new tab), with per-model odds ratios from about 6 to over 10,000. An odds ratio compares the odds of an outcome with and without a factor, so 6 means six times the odds.
- In lab tests of AI recommendations (preprint), a fictional brand beat a well-known one half the time when it showed a 7.3% lower price, a 0.075-star rating advantage, or 1.6 times the reviews (opens in a new tab). Specific, visible facts outweighed brand familiarity.
- In TrustRadius surveys, transparent pricing has been buyers' top wish-list item for four years running (opens in a new tab) (industry survey by a review platform), and 27% of US business professionals who use AI at work said AI vendor recommendations don't reflect real pricing or contract structures (opens in a new tab) (vendor survey).
Evidence: moderate (lab evidence plus surveys).
What to do instead. Publish a price, a range, or the basis for pricing, plus scope, deliverables, and turnaround, in plain text on the page. If the price varies, state what it depends on. For published market prices, see what AEO costs.
How to detect it. Open the page with JavaScript turned off. It passes if a reader can find a price (or a range or basis), the scope, and the turnaround in the text. It fails if the only answer to "What does it cost?" is a form or a call.
Sources: 30, 33, 37, 40.
11. Burying the answer
What it looks like. Pages that open with a story, a history of the industry, or a pitch, and state the answer to the page's question several screens down.
Why people do it. Older content-marketing habits favor a hook and a slow build.
Why it fails.
- In a full pipeline from retrieval to answer (peer-reviewed), placing the answer early in the document produced higher reranking scores, and restructuring that pushed the answer later caused significant rank drops (opens in a new tab). Reranking is the step that orders retrieved pages before the model reads them.
- In experiments on what evidence language models find convincing (peer-reviewed), relevance to the question drove which text won, and even prefixing a paragraph with "The following text is about the question:" followed by the question raised its win rate (opens in a new tab).
- The one controlled field study (preprint; the authors work for the site they studied) added answer-first question titles and standalone summaries of two to three sentences. It found a 1.82-times rise in ChatGPT referrals (95% CI 1.31 to 2.54), but a placebo test gave p = 0.16, so the result is suggestive only (opens in a new tab).
Evidence: moderate.
What to do instead. Put a direct answer of one to three sentences at the top of the page, under a heading that uses the buyer's words, and give the detail after it.
How to detect it. Read the first 100 words of each priority page. It passes if they answer the question in the page title. It fails if the answer first appears later.
Sources: 41, 42, 43.
12. Relying on FAQ markup after May 7, 2026
What it looks like. Adding FAQPage markup to every page as a core AEO step, often for questions that are not visible on the page.
Why people do it. FAQ markup once produced a visible rich result in Google, a search result with extra display elements, so adding it became a habit.
Why it fails.
- Google stopped showing FAQ rich results on May 7, 2026 (opens in a new tab) (official documentation).
- Question-and-answer packaging by itself did not help. In a descriptive study of 21,143 citations (preprint), pages formatted as questions and answers had slightly less influence on answers than other pages (−5.74%) (opens in a new tab).
- In a matched study of pages that added markup of several types, including FAQ markup, citations did not rise (opens in a new tab) (vendor study; details in the next entry).
- Google asks site owners to make sure structured data matches the visible text on the page (opens in a new tab) (official documentation).
Evidence: strong for the removal (official documentation). Moderate for the lack of a citation effect.
What to do instead. Keep visible FAQ sections where they answer real buyer questions, because readers use them. Treat the markup as optional, and mark up only questions and answers that appear on the page.
How to detect it. Crawl the site and list the pages with FAQPage markup. It passes if each marked-up question and answer appears in the visible text and came from a real buyer question. It fails if FAQ markup is presented as an AI visibility tactic, or if it marks up text that readers cannot see.
Sources: 16, 32, 44, 45.
13. Over-investing in schema markup as a citation lever
What it looks like. Treating schema markup as the main way to be cited, and spending most of an AEO budget on markup types.
Why people do it. Markup is concrete and easy to check, and a 2025 preprint reported a correlation of 0.63 between structured data and the likelihood of citation (opens in a new tab).
Why it fails.
- In a matched difference-in-differences study of 1,885 pages that added JSON-LD markup, compared with about 4,000 control pages (vendor study), AI Overview citations fell 4.6%, and the changes for AI Mode (+2.4%) and ChatGPT (+2.2%) were statistically indistinguishable from zero (opens in a new tab). A difference-in-differences study compares the change in treated pages with the change in similar untreated pages. All treated pages were already heavily cited, so the study says nothing about new sites.
- Google states that "structured data isn't required for generative AI search, and there's no special schema.org markup you need to add" (opens in a new tab) (official documentation).
- When five assistants fetched a live test page (practitioner test, one page), none extracted facts that appeared only in JSON-LD (opens in a new tab).
- Bing says structured data "may support clearer grounding but does not guarantee visibility or grounding traffic" (opens in a new tab) (official documentation).
- The correlational study scored only pages that had already been cited, without a clear comparison set of uncited pages, and its authors co-founded a GEO vendor (opens in a new tab) (preprint).
Evidence: moderate (one quasi-experiment, official statements, and one small test).
What to do instead. Keep markup as hygiene: Organization, Person, and Service or Product markup that matches the visible page, with sameAs links to real profiles. Spend the rest of the budget on facts, pages, and third-party corroboration.
How to detect it. Read the plan or proposal. It passes if markup is described as hygiene and every marked-up fact is visible on the page. It fails if the plan states or implies that adding markup will increase AI citations.
Sources: 23, 31, 45, 46, 47.
Next steps
AEO HQ sells answer engine optimization services at fixed, published prices: an automated audit for $499, an audit of search and AI visibility for $2,500, the AEO Blueprint for $4,995, and the Blueprint + Implementation for $8,995.
Sources
- Fishkin, R. (2026, January 28). NEW research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility. SparkToro. https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/ (opens in a new tab)
- Schulte, J., Bleeker, M., & Kaufmann, P. (2026). Don't measure once: Measuring visibility in AI search (GEO) (arXiv:2604.07585). arXiv. https://arxiv.org/abs/2604.07585 (opens in a new tab)
- Uberti-Bona Marin, L. G., Bertaglia, T., Astante, G., Rijsbosch, B., van Dijck, G., Hannák, A., Spanakis, G., & Kollnig, K. (2026). "If I had to buy just ONE: Galaxy S26 Ultra": Auditing AI-generated product recommendations (arXiv:2609.18729). arXiv. https://doi.org/10.48550/arXiv.2609.18729 (opens in a new tab)
- Vernick, N. (2026, September 10). Best answer engine optimization agencies for B2B SaaS: Ranked and tested in AI search. Arobis AI. https://arobis.ai/blog/best-aeo-agencies (opens in a new tab)
- Martinez, O. (2026). Optimizing visibility in generative engines: A critical survey of generative engine optimization (2023–2026) (arXiv:2607.14035). arXiv. https://doi.org/10.48550/arXiv.2607.14035 (opens in a new tab)
- Kumar, P. (2026). Generative engine optimization at scale: Measuring brand visibility across AI search engines (arXiv:2606.20065). arXiv. https://arxiv.org/abs/2606.20065 (opens in a new tab)
- HubSpot. (2026, September 8). How much does AEO cost? A breakdown by approach. HubSpot Blog. https://blog.hubspot.com/marketing/how-much-does-aeo-cost (opens in a new tab)
- Bowyer, S., Aitchison, L., & Ivanova, D. R. (2025). Position: Don't use the CLT in LLM evals with fewer than a few hundred datapoints. In Proceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267). PMLR. https://proceedings.mlr.press/v267/bowyer25a.html (opens in a new tab)
- Fletcher, J., & Verckist, D. (2025, October). News integrity in AI assistants: An international PSM study. European Broadcasting Union & BBC. https://www.ebu.ch/research/open/report/news-integrity-in-ai-assistants (opens in a new tab)
- Xu, H., Iqbal, U., & Montgomery, J. M. (2026). Measuring Google AI Overviews: Activation, source quality, claim fidelity, and publisher impact (arXiv:2605.14021). arXiv. https://doi.org/10.48550/arXiv.2605.14021 (opens in a new tab)
- Jaźwińska, K., & Chandrasekar, A. (2025, March 6). AI search has a citation problem. Columbia Journalism Review, Tow Center for Digital Journalism. https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php (opens in a new tab)
- Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating verifiability in generative search engines. In Findings of the Association for Computational Linguistics: EMNLP 2023 (pp. 7001–7025). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-emnlp.467 (opens in a new tab)
- Linehan, L. (2025, February 6). 63% of websites receive AI traffic (new study of 3,000 sites). Ahrefs. https://ahrefs.com/blog/ai-traffic-study/ (opens in a new tab)
- Belson, D., & Rhea, S. (2025, July 1). The crawl before the fall... of referrals: Understanding AI's impact on content providers. Cloudflare Blog. https://blog.cloudflare.com/ai-search-crawl-refer-ratio-on-radar/ (opens in a new tab)
- Google. (n.d.). [GA4] Default channel group. Analytics Help. Retrieved September 27, 2026, from https://support.google.com/analytics/answer/9756891 (opens in a new tab)
- Google. (2025, December 10). AI features and your website. Google Search Central. https://developers.google.com/search/docs/appearance/ai-features (opens in a new tab)
- Birkett, A. (2026, August 28). First-touch attribution captures 15% of our AI-sourced leads [Research]. Omniscient Digital. https://beomniscient.com/blog/first-touch-vs-self-reported-attribution-aeo/ (opens in a new tab)
- OpenAI. (n.d.). Publishers and developers – FAQ. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/12627856-publishers-and-developers-faq (opens in a new tab)
- Madhavan, K., Merchant, M., Canel, F., & Nigam, S. (2026, February 10). Introducing AI Performance in Bing Webmaster Tools (public preview). Bing Webmaster Blog. https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview (opens in a new tab)
- Google. (n.d.). Generative AI performance report (Search). Search Console Help. Retrieved September 27, 2026, from https://support.google.com/webmasters/answer/16984139 (opens in a new tab)
- Google. (2026, June 5). Google Search's guidance on using third-party SEO tools, services, and advice. Google Search Central. https://developers.google.com/search/docs/fundamentals/third-party-seo (opens in a new tab)
- OpenAI. (n.d.). Searching the web with ChatGPT. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/9237897-chatgpt-search (opens in a new tab)
- Microsoft Bing. (n.d.). Bing Webmaster Guidelines. Retrieved September 27, 2026, from https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a (opens in a new tab)
- Google. (2025, December 18). In-depth guide to how Google Search works. Google Search Central. https://developers.google.com/search/docs/fundamentals/how-search-works (opens in a new tab)
- Allsopp, G. (2025, December 4). Do self-promotional "best" lists boost ChatGPT visibility? Study of 26,283 source URLs. Ahrefs Blog. https://ahrefs.com/blog/best-lists-research/ (opens in a new tab)
- Hong, A. (2026, September 5). 12 best AI SEO, AEO & GEO agencies in 2026, scored. Tobe Agency. https://www.tobeagency.co/learn/12-best-ai-seo-agencies-in-2026-scored-on-what-they-can-prove (opens in a new tab)
- Ray, L. (2026, February 3). Is Google finally cracking down on self-promotional listicles? Substack. https://lilyraynyc.substack.com/p/is-google-finally-cracking-down-on (opens in a new tab)
- Allouah, A., Besbes, O., Figueroa, J. D., Kanoria, Y., & Kumar, A. (2025). What is your AI agent buying? Evaluation, biases, model dependence, & emerging implications for agentic e-commerce (arXiv:2508.02630). arXiv. https://arxiv.org/abs/2508.02630 (opens in a new tab)
- Sabbah, J., & Acar, O. A. (2026). Marketing to machines: How AI models respond to promotional cues [Working paper]. SSRN. https://doi.org/10.2139/ssrn.6406639 (opens in a new tab)
- Chu, X., & Hou, Y. (2026). Incumbent advantage: Brand bias and cognitive manipulation dynamics in LLM recommendation systems (arXiv:2606.17443). arXiv. https://arxiv.org/abs/2606.17443 (opens in a new tab)
- Google. (2026, July 10). Optimizing your website for generative AI features on Google Search. Google Search Central. https://developers.google.com/search/docs/fundamentals/ai-optimization-guide (opens in a new tab)
- Google. (2026, September 24). Latest Google Search documentation updates. Google Search Central. https://developers.google.com/search/updates (opens in a new tab)
- TrustRadius. (2026, July 15). TrustRadius 2026 B2B Buying Disconnect report reveals AI has changed how buyers research, but not what they trust [Press release]. PR Newswire. https://www.prnewswire.com/news-releases/trustradius-2026-b2b-buying-disconnect-report-reveals-ai-has-changed-how-buyers-research-but-not-what-they-trust-302825792.html (opens in a new tab)
- Chen, M., Wang, X., Chen, K., & Koudas, N. (2025). Generative engine optimization: How to dominate AI search (arXiv:2509.08919). arXiv. https://doi.org/10.48550/arXiv.2509.08919 (opens in a new tab)
- Sharma, A. P. (2026). The discovery gap: How Product Hunt startups vanish in LLM organic discovery queries (arXiv:2601.00912). arXiv. https://doi.org/10.48550/arXiv.2601.00912 (opens in a new tab)
- Linehan, L. (2025, December 12). Top brand visibility factors in ChatGPT, AI Mode, and AI Overviews (75k brands studied). Ahrefs. https://ahrefs.com/blog/ai-brand-visibility-correlations/ (opens in a new tab)
- Vishwakarma, R., Kumar, S., & Jamidar, R. (2026). What gets cited: Competitive GEO in AI answer engines. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 4950–4954). ACM. https://doi.org/10.1145/3805712.3808445 (opens in a new tab)
- Lee, Y., Ye, X., & Choi, E. (2024). AmbigDocs: Reasoning across documents on different entities under the same name. In Proceedings of the First Conference on Language Modeling (COLM 2024). https://arxiv.org/abs/2404.12447 (opens in a new tab)
- Rosala, M., & Brown, J. (2026, February 27). GenAI for complex questions, search for critical facts. Nielsen Norman Group. https://www.nngroup.com/articles/ai-search-infoseeking/ (opens in a new tab)
- Loktionova, M. (2026, July 8). How AI tools shape the B2B buying process: A survey of 600+ US business professionals. Semrush. https://www.semrush.com/blog/how-ai-shapes-b2b-buying/ (opens in a new tab)
- Kim, S., Jeong, W., Kim, S., Lee, S., & Lee, D. (2026). SAGEO Arena: A realistic environment for evaluating search-augmented generative engine optimization. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (pp. 2342–2353). ACM. https://doi.org/10.1145/3770855.3818146 (opens in a new tab)
- Wan, A., Wallace, E., & Klein, D. (2024). What evidence do language models find convincing? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 7468–7484). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.403 (opens in a new tab)
- Watanabe, K., & Nakayashiki, K. (2026). Disentangling answer engine optimization from platform growth: A log-based natural experiment on ChatGPT referral traffic (arXiv:2606.04362). arXiv. https://doi.org/10.48550/arXiv.2606.04362 (opens in a new tab)
- Zhang, K., He, X., & Yao, J. (2026). From citation selection to citation absorption: A measurement framework for generative engine optimization across AI search platforms (arXiv:2604.25707). arXiv. https://doi.org/10.48550/arXiv.2604.25707 (opens in a new tab)
- Linehan, L. (2026, May 11). We tracked 1,885 pages adding schema. AI citations barely moved. Ahrefs. https://ahrefs.com/blog/schema-ai-citations/ (opens in a new tab)
- Kumar, A., & Palkhouski, L. (2025). AI answer engine citation behavior: An empirical analysis of the GEO-16 framework (arXiv:2509.10762). arXiv. https://doi.org/10.48550/arXiv.2509.10762 (opens in a new tab)
- searchVIU. (2025, December 2). Schema markup and AI in 2025: What ChatGPT, Claude, Perplexity & Gemini really see. https://www.searchviu.com/en/schema-markup-and-ai-in-2025-what-chatgpt-claude-perplexity-gemini-really-see/ (opens in a new tab)
How to cite this page
Maxwell, P. (2026). Answer engine optimization antipatterns. AEO HQ. Last updated September 27, 2026. https://www.aeohq.ai/articles/aeo-antipatterns
More in Answer engine optimization (AEO)
Complete guide
Answer engine optimization (AEO): the complete guide
What answer engine optimization (AEO) is, how AI answer engines choose sources, which tactics the research supports, and how to plan and measure AEO.
Guide
Answer engine optimization examples
Documented AEO examples: page examples from Google and Microsoft, a field study, pages AI answers cite, and results companies reported, with sources.
Guide
Does schema markup help AEO? What the evidence says
Does schema markup help AEO? What Google, Bing, and controlled studies say about structured data, AI citations, and rich results, and what to do instead.
Guide
Featured snippets and People Also Ask in the AI era
How Google picks featured snippets and People Also Ask answers, how AI Overviews changed both, what the data show, and how to write pages Google can quote.
Guide
How to write content for answer engines
How to write AEO content that answer engines retrieve and cite: answer first, use the buyer's words, state checkable facts, and skip rewriting tricks.
Comparison
AEO vs GEO (and AIO, LLMO): are they the same thing?
AEO, GEO, AIO, and LLMO are mostly labels for one practice. What each term means, where it comes from, how it is used, and where the emphasis differs.