Example engagement ยท B2B SaaS: HR and payroll software
AEO audit for an HR and payroll software company
Hypothetical example, not a client: how AEO HQ's audit would test buyer questions in AI assistants and review crawling, facts, sources, and claims for an HR and payroll software company. No results.
An example engagement for a hypothetical company, showing how the audit runs and what it delivers. It is not a client result.
Industry guide: SEO, AEO, and GEO for B2B SaaS companies
The company
| Item | Hypothetical profile |
|---|---|
| Product | Cloud payroll with federal, state, and local tax filing; benefits administration; time tracking; onboarding; and reporting under the Affordable Care Act (ACA) |
| Market | U.S. employers with 50 to 500 employees, many with staff in several states |
| Company | About 200 employees; sells through a sales team, after a demo |
| Buyers | HR directors, controllers, and CFOs, with IT joining for the security review |
| Pricing | A fee per employee per month plus a base fee. The pricing page shows one "starting at" figure; plan details are in a PDF |
| How buyers find it | Search, software review sites, referrals from accountants and benefits brokers, and peers |
| Current marketing | About 300 blog posts; twelve "[Product] vs [Competitor]" pages built from one template; a review campaign on two review sites; webinars; paid search. The pricing calculator renders in the browser. Google Analytics 4 (GA4) runs with its default channels; the CRM is HubSpot |
Two facts about this market shape the audit. Finance involvement in software decisions rose from 31% to 46% in a year (G2 survey of 1,038 decision-makers, June 2026; G2 runs a review platform), so controllers ask questions too. And many customers of this size are applicable large employers, which the IRS defines by at least 50 full-time employees, including full-time equivalents, on average in the prior year; each must file Forms 1094-C and 1095-C. The panel asks about those forms by number.
The questions buyers ask assistants
The panel follows AEO HQ's measurement design: 40 buyer intents, each written three ways (120 unbranded prompts), plus 20 branded prompts about the company. Wording comes from sales-call notes, support tickets, review-site questions, and Search Console queries. The panel varies the buyer's role and adds constraints such as budget, because 45% of AI-using B2B buyers add constraints such as budget, required features, or compatibility (519 respondents; the publisher sells search-marketing software).
Each prompt runs three times a week on each of the four assistants the audit measures (ChatGPT, Perplexity, Gemini, and Google AI Overviews), in a new chat with a clean session; the ten most important prompts run eight times a week. Results are reported as rates, because the same prompt returned the same list of brands less than once in 100 runs (2,961 runs; industry study; a co-investigator works for a tracking vendor). The table shows 12 of the 140 prompts. The wording is illustrative.
| Intent | Example prompt (illustrative) | What a correct answer depends on |
|---|---|---|
| Category | "Best HR and payroll software for a 150-person company with employees in several states" | Pages stating company-size fit and multi-state payroll |
| Category | "All-in-one HR software with payroll, benefits administration, and ACA reporting" | A feature page for each module |
| Comparison | "[Product] vs [Competitor A] for a 200-person company" | A comparison page with dated, sourced facts |
| Comparison | "Alternatives to [Competitor B] with built-in time tracking" | Third-party lists and review sites |
| Comparison | "Which HR platforms integrate with NetSuite and QuickBooks Online?" | An integrations page in server-rendered HTML |
| Pricing | "How much does [Product] cost per employee per month?" | Prices stated as text |
| Pricing | "Does [Product] charge extra for off-cycle payroll runs or year-end W-2s?" | Plan inclusions and fees |
| Local (state and local taxes) | "Payroll software that handles local income taxes in Ohio and Pennsylvania" | A public page listing tax coverage |
| Local (state and local taxes) | "HR software for a company with remote employees in 15 states" | The same page and the help center |
| Compliance | "Does [Product] file ACA Forms 1094-C and 1095-C?" | The ACA feature page |
| Compliance | "Does [Product] have a SOC 1 Type 2 report?" | A security page naming each report and its period |
| Compliance | "If [Product] makes a payroll tax filing error, who pays the penalty?" | The filing guarantee's terms, with conditions |
How the assistants reach the company
An assistant answers from what its model learned in training or by searching the web and writing from the pages it finds, called retrieval. Prices and tax coverage change after a model is trained, so they reach assistants mainly through search. Anthropic, for example, says Claude searches when a request depends on information that is "current, changing, or outside its training data". Each assistant searches its own index, built by its own crawler; how AI assistants find and cite sources covers each route.
How assistants reach the company's pages
- 01GooglebotGoogle's search crawler; robots.txt applies
- 02Google Search indexIndexed pages eligible for a snippet
- 03AI Overviews and AI ModeLink to supporting pages from the index
OpenAI
- 01OAI-SearchBotOpenAI's search crawler; GPTBot is separate
- 02OpenAI index and partnersOwn crawl; partners include Microsoft
- 03ChatGPT searchRewrites questions into targeted searches
Microsoft
- 01BingbotBing's crawler; sitemaps and IndexNow
- 02Bing indexSame foundation as Bing web search
- 03Microsoft CopilotSends its generated queries to Bing
Anthropic
- 01Claude-SearchBotIndexes pages for Claude's search results
- 02Anthropic index and BraveHow results are split is not documented
- 03Claude web searchSearches for current or company facts
- Assistants in prompt runs
- 4
- Agents checked in logs
- 8
- Google. AI Overviews and AI Mode show only pages that are indexed and eligible to be shown in Google Search with a snippet, from sites included in Search Console's "Search generative AI" setting, the default. Grounding in the Gemini app is controlled by the Google-Extended token. See how to rank in Google AI Overviews and AI Mode.
- ChatGPT. Sites that block OAI-SearchBot "will not be shown in ChatGPT search answers". ChatGPT search also uses other search providers, and OpenAI names Microsoft and Shopify; Enterprise and Edu workspaces may share disassociated search queries with Bing. See how to rank in ChatGPT search.
- Copilot. Copilot sends the queries it generates to the Bing search service, and a NOARCHIVE tag prevents content from being used in Copilot responses.
- Claude. Anthropic lists Brave Search as a web-search subprocessor, and blocking Claude-SearchBot may reduce a site's visibility in Claude's search results. See how to rank in Claude.
- Perplexity. PerplexityBot builds Perplexity's own index; the Perplexity-User fetcher, which opens pages when someone asks a question, "generally ignores robots.txt rules". See how to rank in Perplexity.
What the audit checks
The audit covers four things: whether crawlers can reach and read the pages, whether the facts buyers check are published and consistent, which third-party sources answers draw on, and whether the claims fit the rules. General items are in the AEO checklist; these are specific to this company.
Access, rendering, and indexing
- robots.txt on every host. The site, help center, and developer docs sit on different hosts, and robots.txt rules apply only to the host, protocol, and port number where the file is hosted. The audit reads the groups that apply to OAI-SearchBot, Claude-SearchBot, Claude-User, PerplexityBot, Googlebot, and Bingbot, since a crawler follows the group that names it, and the
*group only when none does. Blocking training crawlers is a separate choice: OpenAI says each of its crawler settings is independent. - The firewall. Google lists allowing crawling "in robots.txt, and by any CDN or hosting infrastructure" among the basics, so the audit reads CDN logs for each crawler's response codes on key URLs.
- Rendering. For pricing, plans, integrations, tax coverage, and security, the HTML the server sends is compared with the rendered page, because none of the major AI crawlers rendered JavaScript in Vercel's December 2024 data.
- Index coverage. About 40 key URLs are checked in Search Console's Page indexing report and URL Inspection tool, and in Bing Webmaster Tools.
- Gated help. Help articles that answer buyer questions are opened in a private window to find those behind a login.
The facts buyers check
The audit writes a fact sheet: each fact a buyer is likely to check, the page that states it, and that page's date. Here it covers plans, prices, fees, the states and localities covered, ACA forms, integrations, security reports, data location, and contract terms. The same sheet is the key for grading answers.
Prices come first: 27% of AI-using B2B buyers say AI answers do not reflect real pricing or contract structures, and in trials across six models, a stated price and a recent date raised a source's odds of being cited first in all six (peer-reviewed; laboratory setting; the authors work for a marketing software vendor). Security follows, since IT security review was the biggest source of delay after a vendor was chosen. The AICPA describes a SOC 1 as an examination of controls at a service organization that are likely to be relevant to user entities' internal control over financial reporting; payroll feeds customers' financial statements, so their controllers may ask for one (our reading).
Third-party sources
The audit lists every URL cited in the baseline runs, groups them (review sites, comparison pages and "best payroll software" lists, communities, news, the company's own site), and checks each against the fact sheet. The map comes from this company's runs, not from general studies, because the studies disagree: comparison pages were the most-cited page type across about 1,000 decision-stage prompts (agency study), while review sites made up 1.1% of 149,912 citations in a 102-brand panel (preprint).
Claims and the rules that limit them
Each claim is listed with its URL and the rule it touches. This is not legal advice; the company's counsel decides.
| Claim | Rule or official source | What the audit checks |
|---|---|---|
| Taking on the customer's tax liability | The IRS: using a payroll service provider "does not relieve the employer of its employment tax obligations or liability for employment taxes" (opens in a new tab); certified professional employer organizations are treated differently (opens in a new tab) | Liability wording against the service's actual role and contract |
| Guarantees and AI features | The FTC: advertisers must have evidence to back up their claims (opens in a new tab), and claims about technology, including artificial intelligence, "need to be backed up" (opens in a new tab) | Each headline next to its conditions and evidence |
| Security reports | AICPA SOC logos are for organizations that received a SOC report issued by a licensed, independent CPA (opens in a new tab) | Report type and period named; no CPA's report called a certification (our reading) |
| Review requests | The FTC rule prohibits incentives conditioned on a particular sentiment (opens in a new tab); its guidance (opens in a new tab) does not address reviews by business buyers | Request emails, incentives, disclosures |
| Posts by employees | The Endorsement Guides: a connection that "might materially affect the weight or credibility of the endorsement" (opens in a new tab), and that the audience would not expect, must be disclosed clearly and conspicuously | Employee reviews and community posts |
| Comparison pages | The FTC judges comparative advertising like all other advertising (opens in a new tab); Google calls separate content for every possible variation (opens in a new tab), made primarily to manipulate rankings or AI responses, scaled content abuse | Sourced, dated competitor facts; shared text across the twelve pages |
The audit also checks GA4 channels, key events, and the demo form's source question against how to track AI referral traffic in GA4.
Sample findings
These are sample findings for the hypothetical company: they show the form a finding takes and the evidence it needs, describe no real company, and were not measured. Impact is our judgment of how directly a finding affects whether assistants can reach, read, and correctly state the company's facts, or how much rule risk it carries; it is not a forecast.
| # | Finding (sample) | Evidence to collect | Impact | Effort | Owner |
|---|---|---|---|---|---|
| 1 | robots.txt puts OAI-SearchBot in the same disallow group as GPTBot | robots.txt history; OAI-SearchBot requests in logs | High: out of ChatGPT search answers | Low | Web engineering |
| 2 | The CDN answers Claude-SearchBot and PerplexityBot with a 403 challenge on /pricing | Firewall events by user agent and published IP range | High for Claude and Perplexity | Low | IT |
| 3 | Prices appear only after the calculator's script runs | Page source with JavaScript off; ChatGPT-User requests | High for pricing prompts | Medium | Web engineering |
| 4 | The help center, which lists local tax coverage, requires a login | Private-window test; articles sales sends to prospects | High for tax-coverage prompts | Medium | Support |
| 5 | The twelve comparison pages share most of their text; competitor prices are undated; several show "Duplicate without user-selected canonical" | Text comparison; Page indexing report; source of each competitor fact | Medium: spam-policy and accuracy risk | High | Product marketing |
| 6 | The home page says the company "takes on your payroll tax liability" | Page copy; service agreement; penalty-guarantee terms | High: conflicts with the IRS statement | Low | Legal, marketing |
| 7 | A "SOC 2 certified" badge; the report on file is a SOC 1 Type 2 | Report cover page and period | Medium: misdescribes the report | Low | Compliance |
| 8 | A review-request email offers a gift card to customers "who love [Product]" | Templates; incentive terms | High: sentiment-linked incentive | Low | Customer marketing |
| 9 | Review profiles show a 2023 price and a 10 to 100 employee fit | Field-by-field comparison with the fact sheet | Medium: conflicting facts | Low | Product marketing |
| 10 | In branded runs, answers call ACA reporting a paid add-on, citing a 2022 post | Run log with cited URLs; the post's status | Medium for accuracy | Low | Content |
| 11 | GA4 uses only the default channel group; demo requests are not a key event; the demo form asks no source question | Traffic acquisition by source / medium; key events; test form | Medium: measurement gap | Low | Marketing operations |
The deliverable
The report is a PDF with a 60-minute readout call. It has ten parts:
- Summary. What was measured, when, and the fixes to make first.
- Method. Panel, assistants, dates, location, account state, run counts, and matching rules, so the baseline can be repeated.
- Baseline rates. Per assistant: mention rate, citation rate, share of voice against eight named competitors, and branded accuracy, each with its count and 95% interval. Answers without a search or citation stay in the denominator. No pooled score and no rank.
- Source map. Cited domains and URLs by type, with the company's facts checked on each.
- Accuracy review. Each branded answer graded against the fact sheet, with the cited page for every error.
- Access and indexing. robots.txt by host, firewall responses, rendering, index status, and crawler requests.
- Claims review. Each claim, its URL, the rule, and a suggested rewording for counsel.
- Findings and fix plan. Owners, effort, and how each fix will be verified.
- Measurement plan. As below.
- Appendices. The run log with every answer and citation, the panel, the fact sheet, and the log extract.
How the example audit runs
Answers
- 01Prompt panel120 unbranded and 20 branded, frozen first
- 02Repeated runsNew chat per run, clean session
- 03CodingMentions, citations, and facts, by fixed rules
- 04Baseline ratesPer assistant, with 95% intervals
- 05Findings and fix planRanked by impact and effort
Site and sources
- 01Crawl and logsrobots.txt, firewall, rendering, fetches
- 02Index coverageAbout 40 key URLs in Google and Bing
- 03Source mapEvery cited URL, grouped by type
- 04Claims reviewFor the company's counsel to decide
- Prompts
- 140
- Assistants run
- 4
- Runs per prompt per week
- 3 (8 for top 10)
- Baseline window
- 2 weeks
Measurement plan
Every measurement uses the same frozen panel, run the same way, following the methodology page and how to measure AI visibility. Two weeks give 820 unbranded and 120 branded runs per assistant, 3,760 in all. The methodology's full design also runs Google AI Mode, Copilot, and Claude; here they are covered by Google's and Microsoft's reports and the logs.
| Metric | Method | Frequency | Tool |
|---|---|---|---|
| Mention rate, per assistant | Runs naming the company on unbranded prompts, divided by all runs; Wilson interval; cluster bootstrap across intents | Baseline: 820 runs per assistant; then rolling four-week windows | Run log (spreadsheet, or a tool that exports raw answers) |
| Citation rate, per assistant | Runs citing any company URL, divided by all runs | As above | As above |
| Share of voice | Company mentions divided by mentions of it and eight named competitors | As above | As above |
| Accuracy | Branded answers graded against the fact sheet, by a person | 120 runs per assistant at baseline, then each window | Run log and fact sheet |
| Index coverage | The 40 key URLs in the Page indexing report, URL Inspection, and Bing | Baseline, then monthly | Search Console; Bing Webmaster Tools |
| AI Overviews and AI Mode impressions | Generative AI performance report (opens in a new tab): impressions, not clicks | Monthly | Search Console |
| Copilot citations | AI Performance report (opens in a new tab): citations and grounding queries, no clicks | Monthly | Bing Webmaster Tools |
| AI referrals and key events | A custom channel group (opens in a new tab) matching assistant domains by source, above Referral, because GA4's default channels name only some assistants and put AI Overviews and AI Mode in Organic Search (opens in a new tab); demo and trial key events | Monthly | GA4 |
| Crawler and fetcher requests | OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Googlebot, and Bingbot on key URLs, with status codes | Monthly | CDN or server logs |
| Self-reported source | "How did you hear about us?" with each assistant as an option | Every demo request | HubSpot |
For scale: at 820 runs, an observed rate of 10% has a 95% Wilson interval of about 8.1% to 12.2% (AEO HQ calculation). The real interval is wider because runs of one prompt are correlated, hence the cluster bootstrap. A change between windows is reported only when the interval for the difference excludes zero, and fixed pages are compared with similar unchanged pages: in the only controlled field study found, ChatGPT referrals to pages that were not changed grew 3.5 times over the same period (preprint; one site).
The fix-and-measure loop
Each fix
- 01Fix one page groupFor example, prices in the server's HTML
- 02Confirm crawl and indexLogs and URL Inspection show the new page
- 03Re-run the frozen panelSame prompts, assistants, and settings
- 04Compare with control pagesSimilar pages left unchanged
- 05Report the differenceOnly if its 95% interval excludes zero
- Measurement window
- 4 weeks, rolling
Engagement timeline
The pricing page gives the audit's current delivery time; this example assumes about two weeks from completed intake.
| Week | Activities |
|---|---|
| Intake | Viewer access to GA4, Search Console, Bing Webmaster Tools, and HubSpot; a log export; eight competitors; sales-call notes; the service agreement and guarantee terms |
| Week 1 | Fact sheet; panel drafted, reviewed by the company, and frozen; access, rendering, and index checks; baseline runs start |
| Week 2 | Runs finish; coding and grading; source map; claims review; findings ranked; report and readout |
| After delivery | The company fixes items in its own order and runs the plan with the handed-over panel, or asks AEO HQ to under a separate engagement |
What this example does not show
- Results. No rates, citations, traffic, or pipeline, because the company does not exist.
- How long changes take. A 2026 review of 45 studies found no technique with "a stable, longitudinal, cross-platform causal effect on organic discoverability" (preprint).
- Legal advice. The claims review flags wording; counsel decides.
- Real prompts. A real panel comes from the company's own calls, tickets, and search data.
- Every buyer's view. Answers vary by account, location, interface, and day; the panel samples fixed conditions on four assistants.
- Ads in assistants, or the product's AI features. Neither is in scope.
Next step
The audit's current scope, price, and delivery time are on the pricing page. The evidence behind these checks is in SEO, AEO, and GEO for B2B SaaS companies.