Antipatterns · AI search optimization (SEO+)
SEO antipatterns
Twelve SEO mistakes that also keep pages out of AI answers, from blocked crawlers to facts hidden in JavaScript, with the evidence and a test for each.
By Paul Maxwell, founder of AEO HQ
Published · Updated
An SEO antipattern is a common search practice that seems harmless or even helpful but keeps pages from being crawled, indexed, trusted, or shown. This page describes 12 that also keep a site out of AI answers. Each entry gives the evidence for why the practice fails and a check you can run to detect it.
Scope and definitions
This page is part of AI search optimization, the approach that treats AI visibility as SEO plus corroboration, facts, and measurement. The mistakes below matter as much for answer engine optimization (AEO) and generative engine optimization (GEO) as for classic SEO, because every major AI assistant finds pages through a crawler or a search index. Retrieval is the step in which an assistant finds candidate pages. Grounding is the step in which it bases its answer on the pages it retrieved.
| AI surface | Where it finds pages |
|---|---|
| Google AI Overviews and AI Mode | Google's search index (opens in a new tab), built by Googlebot |
| Gemini app | Google's crawl. The Google-Extended token controls whether content is used for grounding in the Gemini app (opens in a new tab) |
| ChatGPT search | OpenAI's crawler, OAI-SearchBot, which surfaces websites in ChatGPT's search features (opens in a new tab), plus other search providers, including Microsoft (opens in a new tab) |
| Microsoft Copilot | The Bing search service (opens in a new tab) |
| Claude | Anthropic's Claude-SearchBot and Claude-User, which index and fetch pages for Claude's search results (opens in a new tab), plus Brave Search and TurboPuffer, which Anthropic lists as web search subprocessors (opens in a new tab) |
| Perplexity | PerplexityBot, which surfaces and links websites in Perplexity's search results (opens in a new tab) |
A crawler identifies itself with a name called a user agent, such as Googlebot or OAI-SearchBot. The entries below are the common ways that pages fall out of these crawlers and indexes.
Two companion pages cover the other families of mistakes. Answer engine optimization antipatterns covers measurement, claims, and third-party corroboration. Generative engine optimization antipatterns covers content tactics such as hidden prompts, mass rewriting, and templated page farms.
How to read this page
Each entry has six parts: what the antipattern looks like, why people do it, why it fails, what to do instead, how to detect it, and its sources. "Why it fails" reports evidence. "What to do instead" is AEO HQ's recommendation.
Every fact links to its primary source. The label after a fact gives the type of source:
- (peer-reviewed): a journal article or conference paper that passed peer review.
- (vendor study): research published by a company that sells a related product. Read it as descriptive.
- (official documentation): what a platform states about its own systems. It is authoritative about policy, not about the size of effects.
- Other labels, such as (network measurement), (practitioner test), and (practitioner analysis), mark published sources that were not peer reviewed and often use small or self-selected samples.
Each entry also rates the overall evidence as strong (official documentation, or several independent studies that agree), moderate (consistent evidence limited to a few studies or to correlations), or weak (one small or conflicted source).
The 12 antipatterns at a glance
| # | Antipattern | Do this instead | Evidence |
|---|---|---|---|
| 1 | Expecting AI answers to cite pages that are not indexed | Get indexed in Google and Bing first | Strong |
| 2 | Snippet and preview controls on pages you want cited | Remove them from those pages | Strong |
| 3 | Blocking AI search crawlers in robots.txt | Allow search crawlers; decide on training crawlers separately | Strong |
| 4 | Blocking Google-Extended to protect content | Leave it allowed if you want to be used by the Gemini app | Strong |
| 5 | Firewall rules that block crawlers and agents | Allow verified crawlers and signed agents | Strong |
| 6 | Key facts only in JavaScript, widgets, PDFs, or JSON-LD | Put every decision fact in the server-rendered HTML | Moderate to strong |
| 7 | Request-time sitemap lastmod values | Use each page's real modified date | Strong |
| 8 | Ignoring Bing and IndexNow | Verify the site in Bing Webmaster Tools and send IndexNow pings | Strong |
| 9 | Relying on llms.txt | Fix crawling and indexing first | Moderate to strong |
| 10 | Parasite placements and expired domains | Publish on your own domain and earn editorial coverage | Strong |
| 11 | Tracking only rankings for the prompt as typed | Cover and track the sub-questions | Strong |
| 12 | Over-investing in Core Web Vitals | Fix severe failures, then stop | Weak |
1. Expecting AI answers to cite pages that are not indexed
What it looks like. A new site or section goes live, and the team starts checking AI answers and producing content for AI without confirming that Google and Bing have indexed the pages. No sitemap has been submitted, and no one has looked at Search Console.
Why people do it. AI assistants seem to read the whole web, so it is easy to assume that they will find a new page without a search index.
Why it fails.
- Google's AI Overviews and AI Mode can only show pages that are indexed and eligible to appear with a snippet; no special optimization is required (opens in a new tab) (official documentation).
- Google's guide adds that a site must be included in Search generative AI features in Search Console to be eligible for display in those features (opens in a new tab), and inclusion is the default setting (opens in a new tab) (official documentation).
- Bing states that "Bing and Copilot search experiences rely on the same core crawling, indexing, and ranking foundation as traditional search" (opens in a new tab) (official documentation).
- Indexing takes time. Google says that crawling "can take anywhere from a few days to a few weeks," and that asking it to recrawl the same URL several times "won't get it crawled any faster" (opens in a new tab) (official documentation).
Evidence: strong (official documentation).
What to do instead. Submit the sitemap in Google Search Console and in Bing Webmaster Tools. Request indexing once for each key URL. Confirm that the Search generative AI setting is on "include." Judge AI visibility only after the pages are indexed.
How to detect it. Run URL Inspection in Search Console and in Bing Webmaster Tools on the ten most important URLs. It passes if each one is indexed and allowed to show a snippet. It fails if any key page shows as unknown, "Discovered – currently not indexed," or "Crawled – currently not indexed."
Sources: 1, 9, 10, 11, 12.
2. Snippet and preview controls on pages you want cited
What it looks like. A nosnippet or max-snippet:0 robots directive, data-nosnippet attributes around key content, or noarchive and nocache directives on pages that the business wants AI answers to cite. They often arrive through a CMS default, a plugin, or a legal review.
Why people do it. To limit how content is reused, or because settings were copied from another site without a check of their effect.
Why it fails.
- Google names
nosnippet,data-nosnippet,max-snippet, andnoindexas the controls that limit what its search features show from a page (opens in a new tab), and its AI features show only pages that are eligible for a snippet (official documentation). - Bing says that NOARCHIVE "prevents content from being used in Copilot responses and grounding results," and that NOCACHE limits Copilot to the URL, title, and snippet (opens in a new tab) (official documentation).
Evidence: strong (official documentation).
What to do instead. Use snippet and archive controls only on content that you do not want quoted, and record why each one is there.
How to detect it. For each key page, check both the HTTP headers and the HTML. Run curl -sI https://example.com/page | grep -i x-robots-tag for the headers, and search the page source for nosnippet, max-snippet, data-nosnippet, noarchive, and nocache. It passes if none applies to content you want cited. It fails if any does.
Sources: 1, 11.
3. Blocking AI search crawlers in robots.txt
What it looks like. A robots.txt file that blocks every AI user agent in one list, blocks search crawlers when the intent was to block training crawlers, or adds a rule to the * group and expects bots that have their own group to follow it.
Why people do it. To keep content out of model training, often by copying published lists of "AI bots" that mix training crawlers with search crawlers.
Why it fails. AI companies run separate crawlers for search and for training. Blocking a search crawler keeps the site out of that assistant's search results or reduces its visibility there.
- OpenAI: sites that opt out of OAI-SearchBot "will not be shown in ChatGPT search answers" (opens in a new tab). GPTBot is its separate training crawler (official documentation).
- Anthropic: disabling Claude-SearchBot "may reduce your site's visibility and accuracy in user search results" (opens in a new tab). ClaudeBot is its training crawler (official documentation).
- Perplexity: PerplexityBot "is not used to crawl content for AI foundation models" (opens in a new tab); it exists to surface sites in Perplexity's results (official documentation).
- Brave Search, which Anthropic lists as a web search subprocessor, says that if a page "is not crawlable by Googlebot," Brave's crawler will not crawl it either (opens in a new tab) (official documentation). Blocking Googlebot therefore also blocks Brave's crawler.
- Under the Robots Exclusion Protocol, a crawler follows the group of rules that names it, and the
*group applies only when no group matches (opens in a new tab) (internet standard). A rule added only to*does not reach a bot that has its own group.
Evidence: strong.
What to do instead. Decide separately for search crawlers and training crawlers. Allow the search crawlers of every assistant you want to appear in: Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, and PerplexityBot. Treat training crawlers, such as GPTBot and ClaudeBot, as a separate business decision, and write it down.
How to detect it. Test robots.txt with a robots.txt parser for each user agent: Googlebot, Bingbot, OAI-SearchBot, GPTBot, ChatGPT-User, Claude-SearchBot, ClaudeBot, Claude-User, PerplexityBot, and Google-Extended. It passes if every search crawler may fetch the pages you want cited. It fails if any is blocked, including through a named group that was not updated.
Sources: 3, 6, 8, 13, 14.
4. Blocking Google-Extended to protect content
What it looks like. User-agent: Google-Extended followed by Disallow: /, added to opt out of AI training.
Why people do it. It looks like a safe way to opt out of AI training, because blocking it does not remove the site from Google Search.
Why it fails for visibility.
- The Google-Extended token controls whether content is used for Gemini training and for grounding in Gemini Apps and Vertex AI; it does not affect inclusion in Google Search (opens in a new tab) (official documentation). Blocking it therefore removes the site as a grounding source for the Gemini app, not only from training.
- In a peer-reviewed study presented at SIGIR 2026, 21 publishers that blocked Google-Extended received no Gemini citations and were significantly less likely to be retrieved by AI Overviews (opens in a new tab). The study is observational: the difference may reflect which publishers chose to block rather than the block itself.
Evidence: strong for Gemini grounding (official documentation). Weak for any effect on AI Overviews.
What to do instead. Leave Google-Extended allowed if you want the Gemini app to use your pages as a source. If training use is a real concern, weigh it against losing Gemini grounding, and record the decision.
How to detect it. Read robots.txt for a Google-Extended group. It passes if there is none, or if it allows the pages you want cited. It fails if it disallows them.
Sources: 2, 15.
5. Firewall rules that block crawlers and agents
What it looks like. A CDN (content delivery network) or hosting firewall that blocks or challenges non-browser traffic through JavaScript challenges, CAPTCHAs, "block AI bots" switches, or strict rate limits. robots.txt says "allow," but crawlers receive a 403 error or a challenge page. Some rules are switched on during an attack and never switched off.
Why people do it. Security teams enable bot protection without a list of the bots the business depends on.
Why it fails.
- Google lists making sure that crawling is allowed "in robots.txt, and by any CDN or hosting infrastructure" (opens in a new tab) among the basics for appearing in its AI features (official documentation).
- OpenAI says that to make a site eligible for ChatGPT search, the owner must allow OAI-SearchBot and "confirm that the website host or content delivery network allows traffic from OpenAI's published searchbot IP addresses" (opens in a new tab) (official documentation).
- On Vercel, for example, the "AI Bots" and "Bot Protection" managed rulesets are off by default but can be switched on in one click, and Bot Protection's challenge mode serves a JavaScript challenge to traffic that is unlikely to be a browser (opens in a new tab) (official documentation).
- AI agents that buy on websites, described in how AI agents find and buy services, meet the same walls. OpenAI notes that each website decides whether to allow traffic from its cloud browser (opens in a new tab) (official documentation).
Evidence: strong (official documentation).
What to do instead. Allow verified crawlers and signed agents explicitly. OpenAI's cloud browser, for example, signs its requests with Web Bot Auth (opens in a new tab), which lets a site verify them. Keep challenges off the pages you want read and off the purchase path, and log bot traffic.
How to detect it. From outside your network, request key pages with each crawler's user agent, for example curl -s -o /dev/null -D - -A "OAI-SearchBot" https://example.com/pricing | head -1. Then check server or CDN logs for 403, 429, and challenge responses to the crawlers' published IP ranges. It passes if every search crawler receives a 200 response with the full page. It fails on any 403, 429, or challenge page. A user-agent test alone does not prove that IP-based rules admit the real crawler, so check the logs too.
Sources: 1, 4, 16, 17, 18.
6. Key facts only in JavaScript, chat widgets, tabs, PDFs, or JSON-LD
What it looks like. Prices that load from an API after the page renders, answers that exist only inside a chat widget, content in tabs or accordions that loads on click, specifications only in a PDF or an image, or facts that appear only in JSON-LD, a format for structured data.
Why people do it. Modern front-end frameworks render content in the browser, and structured data is sometimes treated as a place to put facts for machines.
Why it fails.
- In Vercel's measurement of its network in December 2024 (network measurement), none of the major AI crawlers rendered JavaScript (opens in a new tab).
- In an October 2025 test of five assistants on one page (practitioner test), only Gemini executed JavaScript when fetching the live page, and none read facts that existed only in JSON-LD (opens in a new tab).
- Google lists making sure that important content is available in textual form (opens in a new tab) among the basics for its AI features (official documentation).
- Microsoft advises against hiding answers in tabs or expandable menus, PDFs, or images (opens in a new tab) (official guidance), and Bing's guidelines warn against hiding critical content behind client-side rendering (opens in a new tab) (official documentation).
Evidence: moderate to strong (official guidance plus small tests; crawler behavior changes over time).
What to do instead. Put every decision fact in the server-rendered HTML: prices, scope, deliverables, turnaround, and contact details. Use JavaScript for interaction, not for content. Mark up only facts that are also visible on the page.
How to detect it. Fetch the page without running JavaScript, and list the prices that appear in the raw HTML: curl -s https://example.com/pricing | grep -o '\$[0-9,]*' | sort -u. Repeat for other key facts. It passes if every key fact appears in the raw HTML. It fails if a fact appears only after rendering, only inside a widget, or only in JSON-LD.
Sources: 1, 11, 19, 20, 21.
7. Request-time sitemap lastmod values
What it looks like. A generated sitemap that stamps every URL's <lastmod> value with the moment the sitemap was requested, so every page appears to change each time a crawler looks.
Why people do it. A generator that fills lastmod with the current time is simple to write, and constant updates look like freshness.
Why it fails.
- Google says that it uses the
<lastmod>value only if it is "consistently and verifiably" accurate, ignores<priority>and<changefreq>, and treats a sitemap as "merely a hint" (opens in a new tab) (official documentation). A value that changes on every request is not verifiably accurate, so it is likely to be ignored. That last step is our inference. - Bing asks for XML sitemaps with accurate lastmod values, together with IndexNow (opens in a new tab) (official documentation).
Evidence: strong for the requirement (official documentation). No study has measured the effect of an inaccurate value on AI visibility.
What to do instead. Set lastmod from each page's real modified date, such as the "updated" timestamp in your content management system, and change it only when the content changes substantively.
How to detect it. Fetch the sitemap twice, a few minutes apart, and compare the values: curl -s https://example.com/sitemap.xml | grep -o '<lastmod>[^<]*' | sort | uniq -c. It passes if the values are identical across fetches and match each page's real last change. It fails if they change between fetches, or if all of them equal the time of the request.
Sources: 11, 22.
8. Ignoring Bing and IndexNow
What it looks like. Search work that covers Google only: no Bing Webmaster Tools account, no sitemap submitted to Bing, and no IndexNow notices when pages change. IndexNow is a protocol for telling search engines that a URL has changed.
Why people do it. Teams concentrate on the search engine that sends them the most visits.
Why it fails. Bing feeds several AI assistants.
- Microsoft Copilot generates search queries and sends them to the Bing search service (opens in a new tab) (official documentation).
- OpenAI says that ChatGPT search sometimes partners with other search providers, including Microsoft (opens in a new tab), and that ChatGPT Enterprise and Edu may share disassociated search queries with the Bing search engine (opens in a new tab) (official documentation).
- Bing Webmaster Tools includes an AI Performance report with citation counts, cited pages, and grounding queries for Copilot and Bing's AI summaries (opens in a new tab) (official documentation).
- Bing reported in 2023 that 12% of new URLs clicked in its results were first discovered through IndexNow (opens in a new tab) (platform report). An IndexNow submission is shared with all participating search engines, a submission does not guarantee immediate indexing, and Google is not among the participants (opens in a new tab) (protocol documentation).
Evidence: strong for the dependency (official documentation). Moderate for the effect of IndexNow.
What to do instead. Verify the site in Bing Webmaster Tools, submit the sitemap, send IndexNow notices when pages are published or changed, and review the AI Performance report each month.
How to detect it. It passes if the site is verified in Bing Webmaster Tools, the sitemap shows as processed, the IndexNow key file resolves at the site root, and someone has reviewed the AI Performance report in the last month. It fails if any of these is missing.
Sources: 4, 5, 23, 24, 25, 26.
9. Relying on llms.txt
What it looks like. Publishing an llms.txt file, a Markdown summary of a site written for language models, and counting it as the main AI visibility step, sometimes while crawl or indexing problems remain unfixed.
Why people do it. It is cheap, it is specific to AI, and audit tools check for it. In one log study, SEO audit tools were the largest group of requesters of llms.txt files, at 21.7% of requests (opens in a new tab) (vendor study).
Why it fails.
- Across 137,210 domains (vendor study, server logs), 97% of llms.txt files received no requests at all in May 2026, and OAI-SearchBot accounted for 0.74% of AI-bot requests for these files (opens in a new tab).
- Across nearly 300,000 domains (vendor study), removing llms.txt from a model of AI citation frequency improved the model's predictions (opens in a new tab), so the file carried no predictive value.
- Google says that it does not need AI text files for its AI features, "as Google Search itself doesn't use them" (opens in a new tab), and that llms.txt files "won't negatively or positively impact your visibility or rankings" (opens in a new tab) (official documentation).
Evidence: moderate to strong that the file has no current effect on citations in answer engines.
What to do instead. Keep an accurate llms.txt file if you have one; Google calls it "fine if you want to maintain these files for other services" (opens in a new tab). The file may matter more to coding agents than to search: in the same log study, Anthropic's coding agent, Claude-Code, requested these files more often than any retrieval bot (opens in a new tab). Do not count the file as progress on AI visibility, and fix crawling and indexing first.
How to detect it. Check 30 days of server logs for requests to /llms.txt, grouped by user agent. It passes if the file's facts match the site and the plan treats the file as a minor item. It fails if the plan counts llms.txt as an AI visibility result, or if its facts differ from the pages.
Sources: 9, 27, 28, 29.
10. Parasite placements and expired domains
What it looks like. Paying to publish content on a high-authority site that has little editorial involvement in it, or buying an expired domain for its old links and filling it with new content that points to your site.
Why people do it. The host site's or the old domain's ranking strength appears to transfer to the new content.
Why it fails.
- Both are covered by Google's spam policies. The site reputation policy applies where "third-party content is published on a host site mainly because of that host's already-established ranking signals" (opens in a new tab). Expired domain abuse is where an expired domain "is purchased and repurposed primarily to manipulate search rankings by hosting content that provides little to no value to users" (opens in a new tab). Sites that violate the policies "may rank lower in results or not appear in results at all" (opens in a new tab) (official documentation).
- Google introduced these policies in March 2024 with a goal of cutting low-quality, unoriginal results by 40%, and later reported a 45% reduction (opens in a new tab) (company report).
- AI Overviews and AI Mode can only show pages that are indexed and eligible to appear with a snippet (opens in a new tab), so a page removed from results for spam cannot appear there either. That connection is our inference from the two documents.
- Enforcement continues: a Google Search spam update began on September 24, 2026 (opens in a new tab) (official documentation).
Evidence: strong (official policy).
What to do instead. Publish under your own domain, earn coverage from sites that exercise editorial control, and disclose any sponsorship.
How to detect it. List every off-site placement and every domain that you or your vendors control. It passes if each placement had editorial review by the host and each domain has a continuous, relevant history. It fails if a placement was bought on a site with no editorial involvement, or a domain was bought for its old links.
Sources: 1, 30, 31, 32.
11. Tracking only rankings for the prompt as typed
What it looks like. Judging AI readiness by whether a page ranks in Google's top 10 for the full question a buyer might type into an assistant, while ignoring the related questions behind it.
Why people do it. Rank tracking is the familiar SEO report.
Why it fails. AI systems search for rewritten and related queries, not only for the prompt.
- Google says that AI Overviews and AI Mode may issue multiple related searches across subtopics and data sources (opens in a new tab), a technique called query fan-out (official documentation).
- OpenAI says that ChatGPT search typically rewrites your query into one or more targeted queries (opens in a new tab) (official documentation).
- In a study of 15,000 prompts (vendor study), only 12% of the links cited by ChatGPT, Gemini, and Copilot appeared in Google's top 10 results for the same prompt, against 28.6% for Perplexity (opens in a new tab).
- The share of AI Overview citations that also rank in the organic top 10 was 76% in July 2025 (opens in a new tab) and 38% in March 2026 (opens in a new tab), in two vendor studies; the second used an improved parsing method. A peer-reviewed audit found a Jaccard similarity of 0.18 between the sources of AI Overviews and those of organic results (opens in a new tab). Jaccard similarity measures the overlap of two sets, from 0 (none) to 1 (identical).
Evidence: strong that the overlap is partial. The exact figures vary with method and date.
What to do instead. For each buyer question, list the sub-questions behind it, such as price, scope, comparisons, how it works, and how results are measured, and make sure that a distinct page or section answers each. Track rankings for the sub-questions too. Do not generate a templated page for every variation; see the entry on page farms in Generative engine optimization antipatterns.
How to detect it. For five priority buyer questions, write down the sub-questions a buyer would need answered. It passes if each sub-question has a page or section that ranks for it, and reporting tracks those rankings. It fails if reporting covers only the full question as typed.
Sources: 1, 4, 15, 33, 34, 35.
12. Over-investing in Core Web Vitals
What it looks like. Repeated projects to push already-passing speed scores higher, justified as work for AI search.
Why people do it. Core Web Vitals are measurable and familiar, and a faster page feels like a safe bet.
Why it fails.
- Among 107,352 pages that appear in AI Overviews and AI Mode (practitioner analysis), correlations between Largest Contentful Paint (how quickly the main content of a page appears) and AI visibility were small, from −0.12 to −0.18 (opens in a new tab), and the author concluded: "Good performance does not create an advantage. Severe failure creates disadvantage." (opens in a new tab) The sample contained only pages that were already visible.
- By contrast, the rendering problems in entry 6 can hide content from AI crawlers entirely.
Evidence: weak (one correlational analysis).
What to do instead. Fix pages that fail Core Web Vitals badly, then stop. Spend the remaining effort on crawl access, rendering, and content.
How to detect it. Open the Core Web Vitals report in Search Console. It passes if no key page is rated "Poor" and further speed work is not listed as an AI visibility task. It fails if pages that already pass are being optimized further in the name of AI search.
Sources: 36.
Next steps
The Instant AEO Audit ($499) is an automated audit that runs a live crawl of your key pages, checks AI crawler access and schema, compares your site with up to three competitors, and returns a prioritized to-do list.
Sources
- Google. (2025, December 10). AI features and your website. Google Search Central. https://developers.google.com/search/docs/appearance/ai-features (opens in a new tab)
- Google. (2026, July 14). Google's common crawlers. Google Search Central. https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers (opens in a new tab)
- OpenAI. (n.d.). Overview of OpenAI crawlers. OpenAI Developers. Retrieved September 27, 2026, from https://developers.openai.com/api/docs/bots (opens in a new tab)
- OpenAI. (n.d.). Searching the web with ChatGPT. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/9237897-chatgpt-search (opens in a new tab)
- Microsoft. (2026, August 18). Data, privacy, and security for web search in Microsoft Copilot and Microsoft Copilot Chat. Microsoft Learn. https://learn.microsoft.com/en-us/copilot/microsoft-365/manage-public-web-access (opens in a new tab)
- Anthropic. (2026, April 7). Does Anthropic crawl data from the web, and how can site owners block the crawler? Claude Help Center. https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler (opens in a new tab)
- Anthropic. (n.d.). Subprocessors. Anthropic Trust Center. Retrieved September 27, 2026, from https://trust.anthropic.com/subprocessors (opens in a new tab)
- Perplexity. (n.d.). Perplexity crawlers. Perplexity Docs. Retrieved September 27, 2026, from https://docs.perplexity.ai/guides/bots (opens in a new tab)
- Google. (2026, July 10). Optimizing your website for generative AI features on Google Search. Google Search Central. https://developers.google.com/search/docs/fundamentals/ai-optimization-guide (opens in a new tab)
- Google. (n.d.). Search generative AI control. Search Console Help. Retrieved September 27, 2026, from https://support.google.com/webmasters/answer/16908024 (opens in a new tab)
- Microsoft Bing. (n.d.). Bing Webmaster Guidelines. Retrieved September 27, 2026, from https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a (opens in a new tab)
- Google. (2025, December 10). Ask Google to recrawl your URLs. Google Search Central. https://developers.google.com/search/docs/crawling-indexing/ask-google-to-recrawl (opens in a new tab)
- Brave. (n.d.). Brave Search crawler. Retrieved September 27, 2026, from https://search.brave.com/help/brave-search-crawler (opens in a new tab)
- Koster, M., Illyes, G., Zeller, H., & Sassman, L. (2022). Robots Exclusion Protocol (RFC 9309). RFC Editor. https://doi.org/10.17487/RFC9309 (opens in a new tab)
- Grossman, R., Liu, S., Chen, M. K., Smith, M., Borcea, C., & Chen, Y. (2026). How generative AI disrupts search: An empirical study of Google Search, Gemini, and AI Overviews. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 448–459). ACM. https://doi.org/10.1145/3805712.3809667 (opens in a new tab)
- Vercel. (2026, September 10). WAF managed rulesets. Vercel Docs. https://vercel.com/docs/vercel-firewall/vercel-waf/managed-rulesets (opens in a new tab)
- OpenAI. (n.d.). Using cloud browser in ChatGPT. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/20001280 (opens in a new tab)
- OpenAI. (n.d.). ChatGPT Work's cloud browser allowlisting. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/11845367 (opens in a new tab)
- Zecchini, G., Moore, A. A., Ubl, M., & Siddle, R. (2024, December 17). The rise of the AI crawler. Vercel. https://vercel.com/blog/the-rise-of-the-ai-crawler (opens in a new tab)
- searchVIU. (2025, December 2). Schema markup and AI in 2025: What ChatGPT, Claude, Perplexity & Gemini really see. https://www.searchviu.com/en/schema-markup-and-ai-in-2025-what-chatgpt-claude-perplexity-gemini-really-see/ (opens in a new tab)
- Madhavan, K. (2025, October 8). Optimizing your content for inclusion in AI search answers. Microsoft Advertising Blog. https://about.ads.microsoft.com/en/blog/post/october-2025/optimizing-your-content-for-inclusion-in-ai-search-answers (opens in a new tab)
- Google. (2026, July 8). Build and submit a sitemap. Google Search Central. https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap (opens in a new tab)
- OpenAI. (n.d.). ChatGPT search for Enterprise and Edu. OpenAI Help Center. Retrieved September 27, 2026, from https://help.openai.com/en/articles/10093903-chatgpt-search-for-enterprise-and-edu (opens in a new tab)
- Madhavan, K., Merchant, M., Canel, F., & Nigam, S. (2026, February 10). Introducing AI Performance in Bing Webmaster Tools (public preview). Bing Webmaster Blog. https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview (opens in a new tab)
- Canel, F. (2023, September 21). Wix now supports IndexNow for faster content indexing. Bing Webmaster Blog. https://blogs.bing.com/webmaster/september-2023/Wix-now-supports-IndexNow-for-Faster-Content-Indexing (opens in a new tab)
- IndexNow. (n.d.). FAQ. Retrieved September 27, 2026, from https://www.indexnow.org/faq (opens in a new tab)
- Linehan, L. (2026, June 15). We analyzed 137K sites: 97% of llms.txt files never get read. Ahrefs. https://ahrefs.com/blog/llmstxt-study/ (opens in a new tab)
- Deda, Y. (2025, November 7). Does LLMs.txt impact your AI visibility and citations? No, according to research. SE Ranking. https://seranking.com/blog/llms-txt/ (opens in a new tab)
- Google. (2026, September 24). Latest Google Search documentation updates. Google Search Central. https://developers.google.com/search/updates (opens in a new tab)
- Google. (2026, August 28). Spam policies for Google web search. Google Search Central. https://developers.google.com/search/docs/essentials/spam-policies (opens in a new tab)
- Tucker, E. (2024, March 5). New ways we're tackling spammy, low-quality content on Search. Google. https://blog.google/products/search/google-search-update-march-2024/ (opens in a new tab)
- Google. (n.d.). Google Search Status Dashboard: Ranking incident history. Retrieved September 27, 2026, from https://status.search.google.com/products/rGHU1u87FJnkP6W2GwMi/history (opens in a new tab)
- Linehan, L. (2025, August 11). Only 12% of AI cited URLs rank in Google's top 10 for the original prompt. Ahrefs. https://ahrefs.com/blog/ai-search-overlap/ (opens in a new tab)
- Linehan, L. (2025, July 21). 76% of AI Overview citations pull from the top 10. Ahrefs. https://ahrefs.com/blog/search-rankings-ai-citations/ (opens in a new tab)
- Linehan, L. (2026, March 2). Update: 38% of AI Overview citations pull from the top 10. Ahrefs. https://ahrefs.com/blog/ai-overview-citations-top-10/ (opens in a new tab)
- Taylor, D. (2026, January 13). What 107,000 pages reveal about Core Web Vitals and AI search. Search Engine Land. https://searchengineland.com/core-web-vitals-ai-search-visibility-analysis-467456 (opens in a new tab)
How to cite this page
Maxwell, P. (2026). SEO antipatterns. AEO HQ. Last updated September 27, 2026. https://www.aeohq.ai/articles/seo-antipatterns
More in AI search optimization (SEO+)
Complete guide
AI search optimization: SEO plus corroboration, facts, and measurement
What AI search optimization is, how AI assistants choose sources, which tactics hold up in research, and how to measure results. Every claim is sourced.
Guide
Brand mentions and AI recommendations
What brand mentions are, where AI assistants find them, what research shows about mentions and AI recommendations, and how to earn and track them.
Guide
How AI agents find and buy services
How AI agents research and buy services in 2026: what ChatGPT, Google, Copilot, Perplexity, Claude, and Stripe support, and what a service business can do.
Guide
llms.txt: what it is and whether it matters
What llms.txt is, who reads it, and whether it helps AI visibility: the proposal, Google's position, server-log studies, and when a file is worth keeping.
Checklist
Entity and brand consistency checklist
A checklist for giving a company and its people one name and one set of facts across their site, markup, profiles, and the sources AI assistants read.
Checklist
Technical SEO checklist for AI search
A technical SEO checklist for AI search: robots.txt, status codes, canonical URLs, redirects, rendering, page elements, and crawler identity, with sources.