Skip to content
AEO HQ

Antipatterns · AI search optimization (SEO+)

SEO antipatterns

Twelve SEO mistakes that also keep pages out of AI answers, from blocked crawlers to facts hidden in JavaScript, with the evidence and a test for each.

By , founder of AEO HQ

Published · Updated

An SEO antipattern is a common search practice that seems harmless or even helpful but keeps pages from being crawled, indexed, trusted, or shown. This page describes 12 that also keep a site out of AI answers. Each entry gives the evidence for why the practice fails and a check you can run to detect it.

Scope and definitions

This page is part of AI search optimization, the approach that treats AI visibility as SEO plus corroboration, facts, and measurement. The mistakes below matter as much for answer engine optimization (AEO) and generative engine optimization (GEO) as for classic SEO, because every major AI assistant finds pages through a crawler or a search index. Retrieval is the step in which an assistant finds candidate pages. Grounding is the step in which it bases its answer on the pages it retrieved.

A crawler identifies itself with a name called a user agent, such as Googlebot or OAI-SearchBot. The entries below are the common ways that pages fall out of these crawlers and indexes.

Two companion pages cover the other families of mistakes. Answer engine optimization antipatterns covers measurement, claims, and third-party corroboration. Generative engine optimization antipatterns covers content tactics such as hidden prompts, mass rewriting, and templated page farms.

How to read this page

Each entry has six parts: what the antipattern looks like, why people do it, why it fails, what to do instead, how to detect it, and its sources. "Why it fails" reports evidence. "What to do instead" is AEO HQ's recommendation.

Every fact links to its primary source. The label after a fact gives the type of source:

  • (peer-reviewed): a journal article or conference paper that passed peer review.
  • (vendor study): research published by a company that sells a related product. Read it as descriptive.
  • (official documentation): what a platform states about its own systems. It is authoritative about policy, not about the size of effects.
  • Other labels, such as (network measurement), (practitioner test), and (practitioner analysis), mark published sources that were not peer reviewed and often use small or self-selected samples.

Each entry also rates the overall evidence as strong (official documentation, or several independent studies that agree), moderate (consistent evidence limited to a few studies or to correlations), or weak (one small or conflicted source).

The 12 antipatterns at a glance

#AntipatternDo this insteadEvidence
1Expecting AI answers to cite pages that are not indexedGet indexed in Google and Bing firstStrong
2Snippet and preview controls on pages you want citedRemove them from those pagesStrong
3Blocking AI search crawlers in robots.txtAllow search crawlers; decide on training crawlers separatelyStrong
4Blocking Google-Extended to protect contentLeave it allowed if you want to be used by the Gemini appStrong
5Firewall rules that block crawlers and agentsAllow verified crawlers and signed agentsStrong
6Key facts only in JavaScript, widgets, PDFs, or JSON-LDPut every decision fact in the server-rendered HTMLModerate to strong
7Request-time sitemap lastmod valuesUse each page's real modified dateStrong
8Ignoring Bing and IndexNowVerify the site in Bing Webmaster Tools and send IndexNow pingsStrong
9Relying on llms.txtFix crawling and indexing firstModerate to strong
10Parasite placements and expired domainsPublish on your own domain and earn editorial coverageStrong
11Tracking only rankings for the prompt as typedCover and track the sub-questionsStrong
12Over-investing in Core Web VitalsFix severe failures, then stopWeak

1. Expecting AI answers to cite pages that are not indexed

What it looks like. A new site or section goes live, and the team starts checking AI answers and producing content for AI without confirming that Google and Bing have indexed the pages. No sitemap has been submitted, and no one has looked at Search Console.

Why people do it. AI assistants seem to read the whole web, so it is easy to assume that they will find a new page without a search index.

Why it fails.

Evidence: strong (official documentation).

What to do instead. Submit the sitemap in Google Search Console and in Bing Webmaster Tools. Request indexing once for each key URL. Confirm that the Search generative AI setting is on "include." Judge AI visibility only after the pages are indexed.

How to detect it. Run URL Inspection in Search Console and in Bing Webmaster Tools on the ten most important URLs. It passes if each one is indexed and allowed to show a snippet. It fails if any key page shows as unknown, "Discovered – currently not indexed," or "Crawled – currently not indexed."

Sources: 1, 9, 10, 11, 12.

2. Snippet and preview controls on pages you want cited

What it looks like. A nosnippet or max-snippet:0 robots directive, data-nosnippet attributes around key content, or noarchive and nocache directives on pages that the business wants AI answers to cite. They often arrive through a CMS default, a plugin, or a legal review.

Why people do it. To limit how content is reused, or because settings were copied from another site without a check of their effect.

Why it fails.

Evidence: strong (official documentation).

What to do instead. Use snippet and archive controls only on content that you do not want quoted, and record why each one is there.

How to detect it. For each key page, check both the HTTP headers and the HTML. Run curl -sI https://example.com/page | grep -i x-robots-tag for the headers, and search the page source for nosnippet, max-snippet, data-nosnippet, noarchive, and nocache. It passes if none applies to content you want cited. It fails if any does.

Sources: 1, 11.

3. Blocking AI search crawlers in robots.txt

What it looks like. A robots.txt file that blocks every AI user agent in one list, blocks search crawlers when the intent was to block training crawlers, or adds a rule to the * group and expects bots that have their own group to follow it.

Why people do it. To keep content out of model training, often by copying published lists of "AI bots" that mix training crawlers with search crawlers.

Why it fails. AI companies run separate crawlers for search and for training. Blocking a search crawler keeps the site out of that assistant's search results or reduces its visibility there.

Evidence: strong.

What to do instead. Decide separately for search crawlers and training crawlers. Allow the search crawlers of every assistant you want to appear in: Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, and PerplexityBot. Treat training crawlers, such as GPTBot and ClaudeBot, as a separate business decision, and write it down.

How to detect it. Test robots.txt with a robots.txt parser for each user agent: Googlebot, Bingbot, OAI-SearchBot, GPTBot, ChatGPT-User, Claude-SearchBot, ClaudeBot, Claude-User, PerplexityBot, and Google-Extended. It passes if every search crawler may fetch the pages you want cited. It fails if any is blocked, including through a named group that was not updated.

Sources: 3, 6, 8, 13, 14.

4. Blocking Google-Extended to protect content

What it looks like. User-agent: Google-Extended followed by Disallow: /, added to opt out of AI training.

Why people do it. It looks like a safe way to opt out of AI training, because blocking it does not remove the site from Google Search.

Why it fails for visibility.

Evidence: strong for Gemini grounding (official documentation). Weak for any effect on AI Overviews.

What to do instead. Leave Google-Extended allowed if you want the Gemini app to use your pages as a source. If training use is a real concern, weigh it against losing Gemini grounding, and record the decision.

How to detect it. Read robots.txt for a Google-Extended group. It passes if there is none, or if it allows the pages you want cited. It fails if it disallows them.

Sources: 2, 15.

5. Firewall rules that block crawlers and agents

What it looks like. A CDN (content delivery network) or hosting firewall that blocks or challenges non-browser traffic through JavaScript challenges, CAPTCHAs, "block AI bots" switches, or strict rate limits. robots.txt says "allow," but crawlers receive a 403 error or a challenge page. Some rules are switched on during an attack and never switched off.

Why people do it. Security teams enable bot protection without a list of the bots the business depends on.

Why it fails.

Evidence: strong (official documentation).

What to do instead. Allow verified crawlers and signed agents explicitly. OpenAI's cloud browser, for example, signs its requests with Web Bot Auth (opens in a new tab), which lets a site verify them. Keep challenges off the pages you want read and off the purchase path, and log bot traffic.

How to detect it. From outside your network, request key pages with each crawler's user agent, for example curl -s -o /dev/null -D - -A "OAI-SearchBot" https://example.com/pricing | head -1. Then check server or CDN logs for 403, 429, and challenge responses to the crawlers' published IP ranges. It passes if every search crawler receives a 200 response with the full page. It fails on any 403, 429, or challenge page. A user-agent test alone does not prove that IP-based rules admit the real crawler, so check the logs too.

Sources: 1, 4, 16, 17, 18.

6. Key facts only in JavaScript, chat widgets, tabs, PDFs, or JSON-LD

What it looks like. Prices that load from an API after the page renders, answers that exist only inside a chat widget, content in tabs or accordions that loads on click, specifications only in a PDF or an image, or facts that appear only in JSON-LD, a format for structured data.

Why people do it. Modern front-end frameworks render content in the browser, and structured data is sometimes treated as a place to put facts for machines.

Why it fails.

Evidence: moderate to strong (official guidance plus small tests; crawler behavior changes over time).

What to do instead. Put every decision fact in the server-rendered HTML: prices, scope, deliverables, turnaround, and contact details. Use JavaScript for interaction, not for content. Mark up only facts that are also visible on the page.

How to detect it. Fetch the page without running JavaScript, and list the prices that appear in the raw HTML: curl -s https://example.com/pricing | grep -o '\$[0-9,]*' | sort -u. Repeat for other key facts. It passes if every key fact appears in the raw HTML. It fails if a fact appears only after rendering, only inside a widget, or only in JSON-LD.

Sources: 1, 11, 19, 20, 21.

7. Request-time sitemap lastmod values

What it looks like. A generated sitemap that stamps every URL's <lastmod> value with the moment the sitemap was requested, so every page appears to change each time a crawler looks.

Why people do it. A generator that fills lastmod with the current time is simple to write, and constant updates look like freshness.

Why it fails.

Evidence: strong for the requirement (official documentation). No study has measured the effect of an inaccurate value on AI visibility.

What to do instead. Set lastmod from each page's real modified date, such as the "updated" timestamp in your content management system, and change it only when the content changes substantively.

How to detect it. Fetch the sitemap twice, a few minutes apart, and compare the values: curl -s https://example.com/sitemap.xml | grep -o '<lastmod>[^<]*' | sort | uniq -c. It passes if the values are identical across fetches and match each page's real last change. It fails if they change between fetches, or if all of them equal the time of the request.

Sources: 11, 22.

8. Ignoring Bing and IndexNow

What it looks like. Search work that covers Google only: no Bing Webmaster Tools account, no sitemap submitted to Bing, and no IndexNow notices when pages change. IndexNow is a protocol for telling search engines that a URL has changed.

Why people do it. Teams concentrate on the search engine that sends them the most visits.

Why it fails. Bing feeds several AI assistants.

Evidence: strong for the dependency (official documentation). Moderate for the effect of IndexNow.

What to do instead. Verify the site in Bing Webmaster Tools, submit the sitemap, send IndexNow notices when pages are published or changed, and review the AI Performance report each month.

How to detect it. It passes if the site is verified in Bing Webmaster Tools, the sitemap shows as processed, the IndexNow key file resolves at the site root, and someone has reviewed the AI Performance report in the last month. It fails if any of these is missing.

Sources: 4, 5, 23, 24, 25, 26.

9. Relying on llms.txt

What it looks like. Publishing an llms.txt file, a Markdown summary of a site written for language models, and counting it as the main AI visibility step, sometimes while crawl or indexing problems remain unfixed.

Why people do it. It is cheap, it is specific to AI, and audit tools check for it. In one log study, SEO audit tools were the largest group of requesters of llms.txt files, at 21.7% of requests (opens in a new tab) (vendor study).

Why it fails.

Evidence: moderate to strong that the file has no current effect on citations in answer engines.

What to do instead. Keep an accurate llms.txt file if you have one; Google calls it "fine if you want to maintain these files for other services" (opens in a new tab). The file may matter more to coding agents than to search: in the same log study, Anthropic's coding agent, Claude-Code, requested these files more often than any retrieval bot (opens in a new tab). Do not count the file as progress on AI visibility, and fix crawling and indexing first.

How to detect it. Check 30 days of server logs for requests to /llms.txt, grouped by user agent. It passes if the file's facts match the site and the plan treats the file as a minor item. It fails if the plan counts llms.txt as an AI visibility result, or if its facts differ from the pages.

Sources: 9, 27, 28, 29.

10. Parasite placements and expired domains

What it looks like. Paying to publish content on a high-authority site that has little editorial involvement in it, or buying an expired domain for its old links and filling it with new content that points to your site.

Why people do it. The host site's or the old domain's ranking strength appears to transfer to the new content.

Why it fails.

Evidence: strong (official policy).

What to do instead. Publish under your own domain, earn coverage from sites that exercise editorial control, and disclose any sponsorship.

How to detect it. List every off-site placement and every domain that you or your vendors control. It passes if each placement had editorial review by the host and each domain has a continuous, relevant history. It fails if a placement was bought on a site with no editorial involvement, or a domain was bought for its old links.

Sources: 1, 30, 31, 32.

11. Tracking only rankings for the prompt as typed

What it looks like. Judging AI readiness by whether a page ranks in Google's top 10 for the full question a buyer might type into an assistant, while ignoring the related questions behind it.

Why people do it. Rank tracking is the familiar SEO report.

Why it fails. AI systems search for rewritten and related queries, not only for the prompt.

Evidence: strong that the overlap is partial. The exact figures vary with method and date.

What to do instead. For each buyer question, list the sub-questions behind it, such as price, scope, comparisons, how it works, and how results are measured, and make sure that a distinct page or section answers each. Track rankings for the sub-questions too. Do not generate a templated page for every variation; see the entry on page farms in Generative engine optimization antipatterns.

How to detect it. For five priority buyer questions, write down the sub-questions a buyer would need answered. It passes if each sub-question has a page or section that ranks for it, and reporting tracks those rankings. It fails if reporting covers only the full question as typed.

Sources: 1, 4, 15, 33, 34, 35.

12. Over-investing in Core Web Vitals

What it looks like. Repeated projects to push already-passing speed scores higher, justified as work for AI search.

Why people do it. Core Web Vitals are measurable and familiar, and a faster page feels like a safe bet.

Why it fails.

Evidence: weak (one correlational analysis).

What to do instead. Fix pages that fail Core Web Vitals badly, then stop. Spend the remaining effort on crawl access, rendering, and content.

How to detect it. Open the Core Web Vitals report in Search Console. It passes if no key page is rated "Poor" and further speed work is not listed as an AI visibility task. It fails if pages that already pass are being optimized further in the name of AI search.

Sources: 36.

Next steps

The Instant AEO Audit ($499) is an automated audit that runs a live crawl of your key pages, checks AI crawler access and schema, compares your site with up to three competitors, and returns a prioritized to-do list.

Sources

  1. Google. (2025, December 10). AI features and your website. Google Search Central. https://developers.google.com/search/docs/appearance/ai-features (opens in a new tab)
  2. Google. (2026, July 14). Google's common crawlers. Google Search Central. https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers (opens in a new tab)
  3. OpenAI. (n.d.). Overview of OpenAI crawlers. OpenAI Developers. Retrieved September 27, 2026, from https://developers.openai.com/api/docs/bots (opens in a new tab)
  4. OpenAI. (n.d.). Searching the web with ChatGPT. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/9237897-chatgpt-search (opens in a new tab)
  5. Microsoft. (2026, August 18). Data, privacy, and security for web search in Microsoft Copilot and Microsoft Copilot Chat. Microsoft Learn. https://learn.microsoft.com/en-us/copilot/microsoft-365/manage-public-web-access (opens in a new tab)
  6. Anthropic. (2026, April 7). Does Anthropic crawl data from the web, and how can site owners block the crawler? Claude Help Center. https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler (opens in a new tab)
  7. Anthropic. (n.d.). Subprocessors. Anthropic Trust Center. Retrieved September 27, 2026, from https://trust.anthropic.com/subprocessors (opens in a new tab)
  8. Perplexity. (n.d.). Perplexity crawlers. Perplexity Docs. Retrieved September 27, 2026, from https://docs.perplexity.ai/guides/bots (opens in a new tab)
  9. Google. (2026, July 10). Optimizing your website for generative AI features on Google Search. Google Search Central. https://developers.google.com/search/docs/fundamentals/ai-optimization-guide (opens in a new tab)
  10. Google. (n.d.). Search generative AI control. Search Console Help. Retrieved September 27, 2026, from https://support.google.com/webmasters/answer/16908024 (opens in a new tab)
  11. Microsoft Bing. (n.d.). Bing Webmaster Guidelines. Retrieved September 27, 2026, from https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a (opens in a new tab)
  12. Google. (2025, December 10). Ask Google to recrawl your URLs. Google Search Central. https://developers.google.com/search/docs/crawling-indexing/ask-google-to-recrawl (opens in a new tab)
  13. Brave. (n.d.). Brave Search crawler. Retrieved September 27, 2026, from https://search.brave.com/help/brave-search-crawler (opens in a new tab)
  14. Koster, M., Illyes, G., Zeller, H., & Sassman, L. (2022). Robots Exclusion Protocol (RFC 9309). RFC Editor. https://doi.org/10.17487/RFC9309 (opens in a new tab)
  15. Grossman, R., Liu, S., Chen, M. K., Smith, M., Borcea, C., & Chen, Y. (2026). How generative AI disrupts search: An empirical study of Google Search, Gemini, and AI Overviews. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 448–459). ACM. https://doi.org/10.1145/3805712.3809667 (opens in a new tab)
  16. Vercel. (2026, September 10). WAF managed rulesets. Vercel Docs. https://vercel.com/docs/vercel-firewall/vercel-waf/managed-rulesets (opens in a new tab)
  17. OpenAI. (n.d.). Using cloud browser in ChatGPT. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/20001280 (opens in a new tab)
  18. OpenAI. (n.d.). ChatGPT Work's cloud browser allowlisting. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/11845367 (opens in a new tab)
  19. Zecchini, G., Moore, A. A., Ubl, M., & Siddle, R. (2024, December 17). The rise of the AI crawler. Vercel. https://vercel.com/blog/the-rise-of-the-ai-crawler (opens in a new tab)
  20. searchVIU. (2025, December 2). Schema markup and AI in 2025: What ChatGPT, Claude, Perplexity & Gemini really see. https://www.searchviu.com/en/schema-markup-and-ai-in-2025-what-chatgpt-claude-perplexity-gemini-really-see/ (opens in a new tab)
  21. Madhavan, K. (2025, October 8). Optimizing your content for inclusion in AI search answers. Microsoft Advertising Blog. https://about.ads.microsoft.com/en/blog/post/october-2025/optimizing-your-content-for-inclusion-in-ai-search-answers (opens in a new tab)
  22. Google. (2026, July 8). Build and submit a sitemap. Google Search Central. https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap (opens in a new tab)
  23. OpenAI. (n.d.). ChatGPT search for Enterprise and Edu. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/10093903-chatgpt-search-for-enterprise-and-edu (opens in a new tab)
  24. Madhavan, K., Merchant, M., Canel, F., & Nigam, S. (2026, February 10). Introducing AI Performance in Bing Webmaster Tools (public preview). Bing Webmaster Blog. https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview (opens in a new tab)
  25. Canel, F. (2023, September 21). Wix now supports IndexNow for faster content indexing. Bing Webmaster Blog. https://blogs.bing.com/webmaster/september-2023/Wix-now-supports-IndexNow-for-Faster-Content-Indexing (opens in a new tab)
  26. IndexNow. (n.d.). FAQ. Retrieved September 27, 2026, from https://www.indexnow.org/faq (opens in a new tab)
  27. Linehan, L. (2026, June 15). We analyzed 137K sites: 97% of llms.txt files never get read. Ahrefs. https://ahrefs.com/blog/llmstxt-study/ (opens in a new tab)
  28. Deda, Y. (2025, November 7). Does LLMs.txt impact your AI visibility and citations? No, according to research. SE Ranking. https://seranking.com/blog/llms-txt/ (opens in a new tab)
  29. Google. (2026, September 24). Latest Google Search documentation updates. Google Search Central. https://developers.google.com/search/updates (opens in a new tab)
  30. Google. (2026, August 28). Spam policies for Google web search. Google Search Central. https://developers.google.com/search/docs/essentials/spam-policies (opens in a new tab)
  31. Tucker, E. (2024, March 5). New ways we're tackling spammy, low-quality content on Search. Google. https://blog.google/products/search/google-search-update-march-2024/ (opens in a new tab)
  32. Google. (n.d.). Google Search Status Dashboard: Ranking incident history. Retrieved September 27, 2026, from https://status.search.google.com/products/rGHU1u87FJnkP6W2GwMi/history (opens in a new tab)
  33. Linehan, L. (2025, August 11). Only 12% of AI cited URLs rank in Google's top 10 for the original prompt. Ahrefs. https://ahrefs.com/blog/ai-search-overlap/ (opens in a new tab)
  34. Linehan, L. (2025, July 21). 76% of AI Overview citations pull from the top 10. Ahrefs. https://ahrefs.com/blog/search-rankings-ai-citations/ (opens in a new tab)
  35. Linehan, L. (2026, March 2). Update: 38% of AI Overview citations pull from the top 10. Ahrefs. https://ahrefs.com/blog/ai-overview-citations-top-10/ (opens in a new tab)
  36. Taylor, D. (2026, January 13). What 107,000 pages reveal about Core Web Vitals and AI search. Search Engine Land. https://searchengineland.com/core-web-vitals-ai-search-visibility-analysis-467456 (opens in a new tab)

How to cite this page

Maxwell, P. (2026). SEO antipatterns. AEO HQ. Last updated September 27, 2026. https://www.aeohq.ai/articles/seo-antipatterns

More in AI search optimization (SEO+)

Next step

Find out who AI recommends.

Book the audit to see where you rank, where AI cites you, and where competitors win. The full audit price credits toward the Blueprint within 30 days.