Vidima AI

measurement

AI visibility audit: how to run one when no two answers match

An AI visibility audit answers two questions: can AI engines read your site, and do they name you? Eight steps, and the sample sizes that make them hold up.

14 min read
A hand holds a magnifying glass up to a laptop screen, enlarging part of a street map while the rest of the desk stays out of focus.

Key takeaways

  • An AI visibility audit answers two separate questions: can AI engines read your site, and do they name you when a buyer asks?
  • Blocking OAI-SearchBot keeps a site out of ChatGPT's search answers; GPTBot is OpenAI's training crawler, a separate line in robots.txt.
  • Across 2,961 runs, SparkToro found under a 1-in-100 chance that two ChatGPT or Google AI answers name the same brands.
  • In 168 answers we recorded, Perplexity cited 22.6 sources per answer, ChatGPT 7.4 and Gemini none.

Ask ChatGPT the same buying question twice and you will almost never get the same list of businesses back. In a January 2026 study, SparkToro had 600 volunteers run 12 prompts through ChatGPT, Claude and Google's AI 2,961 times. For ChatGPT and Google's AI, there was under a 1-in-100 chance that any two answers named the same brands. Getting the same brands in the same order took about 1,000 runs.

That finding decides how an AI visibility audit has to work. Typing your category into ChatGPT and screenshotting the answer is not an audit. A single answer that names you is consistent with a true mention rate anywhere from 21% to 100%.

It is also the thing people most want to know. We mined 7,836 questions from 2,910 Reddit threads about AI search visibility, posted between September 2025 and August 2026. 1,390 of them, 17.7%, were about measurement or tracking. The most repeated question, in 27 near-identical versions, began "how do I measure whether my brand is cited in…".

This guide is the method in eight steps. Part 1 checks whether AI engines can read your site, and you do it once. Part 2 checks whether they name you. That part is a sample, and you repeat it.

What an AI visibility audit checks

Part 1: accessPart 2: answers
The questionCan AI engines read your site?Do AI engines name you when a buyer asks?
What you checkrobots.txt, bot challenges, JavaScript rendering, indexingMentions, recommendations, citations, who is named instead, which pages are cited
The resultPass or failA rate, per engine
How oftenOnce, then after any site, hosting or security changeOn a fixed schedule, with the same questions

Part 1 comes first for a practical reason. If an engine cannot fetch your pages, nothing you learn in Part 2 can be fixed with content.

Part 1: can AI engines read your site?

Step 1: Check robots.txt for search crawlers, not just training crawlers

OpenAI, Anthropic and Google each run separate crawlers for training models and for answering questions. Perplexity says its search crawler is not used for training at all. Blocking a training crawler is a decision about how your content is used. Blocking a search crawler takes you out of that engine's answers. The two are easy to confuse, and one blanket rule blocks both.

Open yoursite.com/robots.txt and look for these names. Each description comes from the company's own documentation, checked on 29 September 2026: OpenAI, Anthropic, Perplexity and Google.

CompanyCrawlerJobWhat the documentation says
OpenAIOAI-SearchBotSearchSites that block it "will not be shown in ChatGPT search answers", though they can still appear as navigational links
OpenAIGPTBotTrainingBlocking it means your content "should not be used in training"
OpenAIChatGPT-UserUser-initiated visitsrobots.txt rules "may not apply"
AnthropicClaude-SearchBotSearchCrawls to improve the relevance and accuracy of Claude's search responses
AnthropicClaudeBotTrainingCollects content that could contribute to model training
AnthropicClaude-UserUser-initiated visitsVisits a site when a Claude user's question calls for it
PerplexityPerplexityBotSearchSurfaces and links sites in Perplexity; "not used to crawl content for AI foundation models"
PerplexityPerplexity-UserUser-initiated visits"Generally ignores robots.txt rules"
GoogleGooglebotSearch, including AI Overviews and AI ModeIts robots.txt rules govern how your site is crawled for Search; AI Overviews and AI Mode use Search's ordinary eligibility
GoogleGoogle-ExtendedGemini training and grounding"Does not impact a site's inclusion in Google Search"

Check the four search crawlers first: OAI-SearchBot, Claude-SearchBot, PerplexityBot and Googlebot. If any of them is blocked from /, you are asking that engine's search to leave your site out. Fix that before you do anything else in this guide.

Watch the fallback rule. A crawler obeys the most specific group that names it, and falls back to the User-agent: * group only when nothing names it. So a Disallow: / under User-agent: * blocks every crawler you have not named in its own group, including all four search crawlers.

Google-Extended catches people out in both directions. Blocking it does not remove you from AI Overviews, which follow Googlebot. It does stop Gemini Apps from grounding answers in your content.

robots.txt can say yes while your firewall says no. If your host or CDN shows visitors a "checking your browser" challenge, a crawler that cannot pass it gets the challenge page instead of your content.

Step 2: Check that your content is in the HTML, not built by JavaScript

Vercel's analysis of AI crawler traffic, published in December 2024, found that none of the major AI crawlers rendered JavaScript. That included OAI-SearchBot, ChatGPT-User, GPTBot, ClaudeBot and PerplexityBot. They fetched JavaScript files but did not execute them. The exceptions were Gemini, which uses Googlebot's infrastructure, and AppleBot. The analysis is almost two years old and crawlers change, so test your own site rather than assume.

The test takes a minute. Open your page, choose View Source (not Inspect, which shows the page after JavaScript has run) and search for a sentence from your main content: what you sell, your prices, your location. If it is not in the source, a crawler that does not run JavaScript does not see it either. From a terminal:

curl -s https://yoursite.com | grep -c "a sentence from your page"

A result of 0 means the sentence is not in the HTML a crawler receives.

Step 3: Check you are indexed, and skip the AI-specific files

Google's requirement for its AI features is its ordinary one. In Google's words: "To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet." Check your key pages with the URL Inspection tool in Google Search Console. Also look for nosnippet, data-nosnippet, max-snippet or noindex, which Google lists as the controls that limit how your content appears.

The same page settles a question that eats a lot of audit time: "You don't need to create new machine readable files, AI text files, or markup to appear in these features," and "there's also no special schema.org structured data that you need to add."

The case for llms.txt is weaker than its popularity suggests. In August 2026, Mark Williams-Cook invented cats.txt, a joke file declaring a company's office cats. It passed the same four tests offered as evidence for llms.txt: bots crawled it, Google indexed it, LLMs cited it and ChatGPT endorsed it. Those, he wrote, are "the four things that happen to literally any text file you put on the open web." Spend your time on Steps 1 and 2 instead.

Part 2: do AI engines name you?

Step 4: Write the questions your buyers ask, without your name in them

A buyer asking for a recommendation does not know your name yet. So the questions that measure visibility leave you out:

  • "Who is the best emergency plumber in Manchester?"
  • "Which bakery in Munich makes the best wedding cakes?"
  • "What is the cheapest accounting software for a two-person company?"
  • "What is a good alternative to [the competitor you lose to most]?"

Asking "What is [your business]?" tests something different: whether AI describes you correctly. That is worth checking, but keep it in a separate list and never mix it into your rates. A branded question inflates a mention rate without telling you anything about discovery.

Keep each question shape (best, cheapest, for a specific need, alternative to a competitor) as its own group, so you can see which kind of question you lose. Write the questions in the language your buyers use, and treat each language as a separate audit rather than merging results.

Five questions is enough to start. Our own citation tracking uses five buying-intent prompts. What matters more is that the list is fixed. Change the questions between audits and you can no longer compare them.

Step 5: Choose your engines, and audit each one separately

AI engines are not interchangeable surfaces. Between 5 and 15 August 2026 we ran five buying-intent prompts against four engines, 42 answers each, and recorded every source they cited.

EngineAnswersSources citedPer answer
Perplexity4294922.6
Google AI Overviews4253412.7
ChatGPT423107.4
Gemini4200

Two consequences for an audit:

Gemini gives you no citations to measure. All 42 of its answers cited zero sources. On Gemini you can audit whether you are mentioned, not whether you are cited.

Each engine reads different sources. TechRadar was cited 33 times, every time by ChatGPT. SE Ranking's Visible site was cited 30 times, every time by Perplexity. A source list from one engine tells you little about another, so report each engine separately and never average them into one score.

Pick the engines your buyers use and write the list down; it is part of the audit's definition. Ask from a clean session each time, logged out or in a temporary chat with memory off, so your own history does not shape the answer. For local questions, record the location you asked from, because a "near me" answer depends on where the engine thinks you are.

Step 6: Ask every question many times

SparkToro's Rand Fishkin put the conclusion of his study plainly: "if you really want to know an AI's set of recommendations, you need to ask over and over again; usually at least 60-100X." In one of his tests, 994 answers to 142 prompts about headphones, Bose, Sony, Sennheiser and Apple each appeared in 55–77% of answers. That is the kind of number an audit can use: a rate, not a list.

The arithmetic shows why. Here is the range of true mention rates each result is consistent with, at 95% confidence:

RunsResultPlausible true rate
1Named in 1 of 121–100%
10Named in 5 of 1024–76%
60Named in 30 of 6038–62%
100Named in 50 of 10040–60%

A manual audit will not reach 100 runs per question per engine. A workable compromise: take your three most important questions and run each 10 times on each engine. Read the results as bands, not scores. Named in 10 of 10 means reliably named. Named in 0 of 10 means rarely named, not never: the true rate could still be as high as 28%. Anything in between means "sometimes", and it needs more runs before you believe a change.

An answer can include you in three different ways, and they are not the same result:

  • Mentioned: your name appears anywhere in the answer.
  • Recommended: you are named as an answer to the question, on the shortlist rather than in a caveat.
  • Cited: a page of yours is linked as a source.

They come apart in practice. Lily Ray studied 100 B2B "best software" queries in Google AI Overviews between April and June 2026. When a company's own self-promotional listicle was cited, that company was left out of the recommendation 69% of the time. Its page was used, and its competitors were recommended.

Record every run as one row:

DateEngineQuestionRunMentionedRecommendedPositionYour page citedNamed insteadSources cited

Your three rates per engine are then simple counts: runs where you were mentioned, recommended or cited, divided by total runs. Skip the citation rate for Gemini.

If you want to see a single run before building the sheet, our free AI visibility checker writes three buyer questions from your site, puts the one you pick to ChatGPT and Claude, lists who they named and which pages they cited, and reads your robots.txt. It is one run, so read it the way Step 6 says: a starting point, not a rate.

Step 8: Turn the answers into a fix list

The last two columns of the sheet are the audit's output.

"Named instead" is your competitor list as AI sees it. Compare it with the competitors you would have named yourself. The ones you did not expect are the ones to study: what pages mention them, and what those pages say.

"Sources cited" tells you where to act. Across the 1,793 citations in our August sample, editorial pages (blogs, guides and roundups) took 53.5%, and a single Zapier roundup was cited 39 times. Being named on a page the engines already read gets you into answers without having to out-rank that page. For a local business, this column shows which directories, review sites and local publications the engines rely on in your market.

Work through the fixes in this order:

  1. Any Part 1 failure. A blocked search crawler or content missing from the HTML comes before everything else.
  2. Cited pages that do not name you. Listings you can claim, reviews you can earn, editors you can pitch.
  3. Your own pages. Write them so a single passage answers the buyer's question directly, with the specifics an engine can lift: what you do, where, for whom, at what price.

One thing not to do: publish a "best of" list that ranks you first. Ray's 69% is what happens next.

How often to re-run an AI visibility audit

Part 1: once, then after any change to your site, hosting, CDN or security settings. A robots.txt edit or a new firewall rule can undo it overnight.

Part 2: on a fixed schedule, such as monthly, with the same questions, the same engines, the same location and the same number of runs. Change any of those and you are comparing two different tests. Engines also change how they find sources without announcing it, which is exactly why the schedule matters: a drop you catch within a month is a drop you can still trace to a cause.

How we sourced this

Our engine data. Five buying-intent prompts about AI visibility tools, in English, run against ChatGPT, Perplexity, Gemini and Google AI Overviews from 5 to 15 August 2026: 42 answers per engine, 168 in total, 1,793 cited sources. Page types were bucketed by URL path. It covers one category and one language, and commercial comparison questions pull more roundups than informational ones, so treat the source shares as a reading from that sample, not a constant.

Our question data. 7,836 questions extracted from 2,910 Reddit threads about AI search visibility, posted between September 2025 and August 2026, and classified by topic automatically.

Everything else. Crawler descriptions come from each company's documentation, checked on 29 September 2026. The JavaScript finding is Vercel's, from December 2024. The ranges in Step 6 are 95% Wilson score intervals, which we calculated. Vidima sells AI visibility tracking; every step in this guide works without it.

Common questions

What is an AI visibility audit?
An AI visibility audit checks two things: whether AI engines such as ChatGPT, Perplexity, Gemini and Google AI Overviews can read your website, and whether they name your business when a buyer asks for a recommendation. The first part is a pass-or-fail technical check. The second is a sample of repeated answers, reported as a rate for each engine.
Can I run an AI visibility audit for free?
Yes. Reading your robots.txt, checking your page source and inspecting URLs in Google Search Console cost nothing, and the consumer versions of the AI engines answer questions for free. The cost is time: three questions, ten runs each, on four engines is 120 answers to read and record.
How many times should I run each prompt in an AI visibility audit?
Far more than once. SparkToro's Rand Fishkin recommends asking the same question 60 to 100 times, because his study found under a 1-in-100 chance that two AI answers name the same brands. At ten runs, a result of five out of ten is still consistent with a true rate anywhere from 24% to 76%.
Does blocking GPTBot remove my site from ChatGPT?
OpenAI documents GPTBot as the crawler that collects content which may be used to train its models. Appearing in ChatGPT's search answers is governed by a separate crawler, OAI-SearchBot, and OpenAI says sites that block OAI-SearchBot will not be shown in ChatGPT search answers. Check both lines in your robots.txt.
Do I need llms.txt or special schema markup to appear in AI answers?
Google says no for its own AI features: you don't need new machine-readable files, AI text files or special schema.org markup to appear in AI Overviews or AI Mode. A page needs to be indexed and eligible to show in Google Search with a snippet.
How often should I run an AI visibility audit?
Run the access checks once, then again after any change to your site, hosting or security settings. Run the answer checks on a fixed schedule, such as monthly, with the same questions, engines, location and number of runs each time. Change any of those and you are comparing two different tests.

Keep reading