AI Ranked
Guide

How to measure AI visibility

Measuring AI visibility means running a fixed set of buyer prompts through each AI engine many times and tracking how often you are mentioned, cited and recommended against competitors.

Short answerAI visibility is how often and how favorably AI engines such as ChatGPT, Perplexity, Gemini and Google AI Overviews mention your brand and cite your pages. To measure it, build a fixed set of buyer prompts, run each one several times per engine, and track mention rate, citation rate, share of voice, position, sentiment and cited sources over time.
Key takeaways
  • In a 2,961-run study by SparkToro and Gumshoe, ChatGPT and Google's AI returned the same brand list for a prompt less than 1 time in 100.
  • The same study found the top 3 brands per prompt appeared in 64% to 73% of runs, so rates are stable even when lists are not.
  • GA4's Default Channel Group has an AI Assistant channel; Google AI Overviews and AI Mode clicks still count as Organic Search.
  • Search Console's generative AI performance reports reached all sites on August 31, 2026 and report impressions.
  • Pew found only 1% of visits with a Google AI summary included a click on a source inside it, so mentions matter as much as clicks.

What is AI visibility?

AI visibility is the share of relevant AI-generated answers that mention your brand, recommend your product or cite your pages. It's the scoreboard for generative engine optimization, the way rankings and clicks are the scoreboard for SEO.

Clicks alone undercount it. Pew Research found that Google users clicked a link inside an AI summary in just 1% of visits where one appeared1 (Pew Research Center). Most of the value of being named never becomes a session in analytics, so you have to measure the answers themselves.

How do you build a prompt set?

Write down the questions real buyers ask at each stage, phrased the way people type into a chat box, and tag each one by funnel stage. This list is the foundation of prompt tracking, and it's where most measurement goes wrong, because teams load the keywords they already rank for and skip the questions buyers ask.

Pull raw language from sales call notes, support tickets, site search, Reddit threads and long conversational queries in Search Console. Then add the constraints buyers actually state, such as company size, industry, budget, region and integrations, because "best CRM for a 10-person agency in the UK" gets a different answer from "best CRM".

Phrasing matters more than most teams expect. In the SparkToro study, 142 people were asked to write their own prompt for the same need, and almost no two were alike2. For your most important questions, track two or three phrasings so one lucky wording doesn't carry the result.

Then freeze it. Keep the core set unchanged for at least a quarter and add new prompts as a separately tagged group, or your trend line compares two different tests.

Example: a 20-prompt set for a payroll SaaS

Here's a starter set for Northwind Payroll, a fictional company we made up for this example: US payroll software for companies with 10 to 200 employees. The competitors named are real. Swap in your own category, constraints and rivals.

#StagePrompt
1ProblemHow do I run payroll for employees in multiple states?
2ProblemWhat happens if a small business files payroll taxes late in the US?
3ProblemHow do small businesses handle payroll for hourly and salaried staff together?
4ProblemShould a 30-person company outsource payroll or run it in-house?
5ProblemHow do I switch payroll providers in the middle of the year?
6CategoryBest payroll software for a 50-person company
7CategoryBest payroll software for restaurants with tipped employees
8CategoryPayroll software that integrates with QuickBooks Online
9CategoryCheapest full-service payroll for a startup with 15 employees
10CategoryPayroll software with multi-state tax filing for a remote team
11CategoryWhich payroll tools do accountants recommend for small business clients?
12ComparisonGusto vs Rippling for a 40-person startup
13ComparisonADP vs Paychex for a small business
14ComparisonAlternatives to Gusto for a company with employees in 10 states
15ComparisonIs Gusto worth it compared with cheaper payroll options?
16ComparisonWhich payroll provider has the best customer support?
17BrandNorthwind Payroll vs Gusto
18BrandIs Northwind Payroll legit?
19BrandNorthwind Payroll pricing
20BrandWhat do customers complain about with Northwind Payroll?

Two design choices are deliberate. The brand prompts (17 to 20) name Northwind, so they'll mention it every time; keep them out of mention rate and share of voice and use them only for sentiment and accuracy. And the problem prompts don't ask for a vendor at all, which is the point: they show whether engines bring you up before the buyer knows to look for software.

A results sheet you can copy

Log one row per prompt, per engine, per run. Paste this header into a spreadsheet or a CSV file:

run_date,engine,mode,model_noted,prompt_id,prompt_text,stage,run_no,brand_mentioned,mention_position,cited_our_domain,cited_urls,competitors_mentioned,sentiment,claims_about_us,notes

Two example rows for the fictional Northwind set:

2026-10-05,chatgpt,search on,not shown,P06,"Best payroll software for a 50-person company",category,1,yes,4,no,"g2.com/...; nerdwallet.com/...","Gusto; Rippling; ADP; Paychex",neutral,"good for multi-state",
2026-10-05,chatgpt,search on,not shown,P06,"Best payroll software for a 50-person company",category,2,no,,no,"forbes.com/...; reddit.com/...","Gusto; Rippling; Paychex",,,list of 5

Use yes/no in the mention and citation columns so a COUNTIF gives you the rate directly. Separate multiple values with semicolons so the commas don't break the CSV. Fill model_noted whenever the interface shows a model name; it's how you'll explain a sudden jump later.

Which metrics should you track?

Track six metrics: mention rate, citation rate, share of voice, position, sentiment and source diversity. Mention rate and share of voice say whether you're in the answer at all; the other four explain why, and point at the page or source to fix.

MetricDefinitionHow to calculate
Mention rateHow often answers name your brandRuns that mention you ÷ total runs
Citation rateHow often answers link to your domain as a sourceRuns citing your domain ÷ total runs
Share of voiceYour share of all brand mentions in the competitor setYour mentions ÷ mentions of all tracked brands
PositionWhere you appear when listed with othersAverage rank in lists where you appear, reported next to mention rate
SentimentHow favorably answers describe youShare of mentions that are positive, neutral or negative, plus recurring claims
Source diversityWhich domains engines rely on for your topicCount and mix of cited domains: yours, competitors, reviews, forums, media

Position is the one people misread. Because lists change on almost every run, an average position only means something next to the mention rate it came from. A brand at position 2 in 10% of runs is weaker than one at position 4 in 80%.

Sentiment is where the qualitative work sits. Log the specific claims engines repeat about you, such as "expensive", "best for enterprise" or a feature limit you fixed a year ago. Those claims usually trace back to a cited page you can update, answer or outweigh. Our analysis of how AI engines choose sources explains why each engine leans on different domains.

Worked example: computing the three headline numbers

Say Northwind runs its 16 non-brand prompts 3 times each in ChatGPT: 48 runs. These results are invented to show the arithmetic.

  • Northwind is named in 16 of the 48 runs. Mention rate = 16 ÷ 48 = 33%.
  • An answer links to a northwindpayroll page in 5 runs. Citation rate = 5 ÷ 48 = 10%.
  • Across the same 48 runs, counting each brand at most once per run: Gusto 39, Rippling 31, ADP 27, Northwind 16, Paychex 14. That's 127 brand mentions in total. Northwind's share of voice = 16 ÷ 127 = 12.6%. Gusto's is 39 ÷ 127 = 30.7%.

Now split mention rate by stage, because that's where the action items come from. Problem prompts (15 runs): 2 mentions, 13%. Category prompts (18 runs): 7 mentions, 39%. Comparison prompts (15 runs): 7 mentions, 47%. Northwind shows up once buyers are comparing vendors but is nearly absent while they're still defining the problem, so the next content work goes into problem-stage pages and the third-party sources cited for them.

The same counting rules have to hold every week: same prompts, same engines, same number of runs, each brand counted once per run. Change any of them and the new number isn't comparable with the old one.

How many runs are enough when answers change every time?

More than most teams run. At 48 runs per engine the margin of error on a mention rate is about ±13 points, so spotting 5 to 10 point moves takes a few hundred runs. Treat every prompt as a sample and report rates, never single outcomes, because the same prompt can return different brands, a different order and a different number of items each time.

The clearest public evidence comes from SparkToro and Gumshoe. Some 600 volunteers ran 12 prompts through ChatGPT, Claude and Google's AI a combined 2,961 times in November and December 2025. There was less than a 1-in-100 chance that ChatGPT or Google's AI would return the same list of brands across runs. Yet the top three brands per prompt still appeared often: on average 73% of the time in Claude, 68% in Google's AI and 64% in ChatGPT2 (Search Engine Land). Each prompt was run 60 to 100 times per platform, and the authors left open how many runs are needed for reliable visibility data14 (Search Engine Journal). The study wasn't peer reviewed.

No one has published a standard sample size, so here's the reasoning we use. A mention rate is a proportion, and the usual 95% margin of error for a proportion is about 1.96 × √(p × (1 − p) ÷ n), where p is the rate and n the number of runs.

Runs per engine per period (n)Margin at a 33% rateMargin at a 50% rate
48 (16 prompts × 3 runs)about ±13 pointsabout ±14 points
100about ±9 pointsabout ±10 points
300about ±5 pointsabout ±6 points
Our takeThose margins are a best case. Repeated runs of the same prompt aren't fully independent, so the real uncertainty is wider. With a 48-run baseline like Northwind's, we'd only treat a change of 15 points or more as real. Seeing 5 to 10 point moves takes a few hundred runs per engine per period, and that's roughly where a paid tracker starts to beat a spreadsheet. If you have to choose, spend the budget on more prompts. 100 prompts run 3 times tells you more than 20 prompts run 15 times, because it also averages out phrasing luck.

A few conditions keep runs comparable. Use the same engine mode, logged-out or clean sessions, the same location and the same day of the week, since memory and personalization change answers. Record the model and mode, because a model update can move every number at once and you don't want to credit your own work for it. Prefer the consumer interface people actually use over the API, since responses can differ; Profound and Peec AI both say they collect answers from the consumer interface for this reason34.

Five measurement mistakes we see often

  • Counting brand prompts in share of voice. "Is Northwind Payroll legit?" mentions Northwind every time and inflates the rate. Report brand prompts separately.
  • Screenshotting one answer for the board deck. One run is one sample; the SparkToro data says the next run will almost certainly differ.
  • Editing the prompt set mid-quarter. A reworded prompt is a new prompt. Add it to a separate group and leave the core set alone.
  • Mixing modes. ChatGPT with search on and off draws on different information, so log the mode and never average across them.
  • Reading GA4's AI Assistant channel as total AI impact. It counts visits, and Pew's 1% figure says most exposure never becomes a visit.

Which tools measure AI visibility?

Four kinds of tool measure AI visibility: dedicated trackers, AI modules inside SEO suites, free one-time checks, and your own analytics. The table lists what each vendor says its tool does. For plan-by-plan prices, prompt caps and engine coverage, see our comparison of paid AI visibility tracking tools.

ToolCategoryWhat it doesEngines named by vendor
ProfoundDedicated trackerAnswer Engine Insights for visibility, sentiment, citations and competitor benchmarks; Prompt Volumes for what people ask; Agent Analytics for AI crawler activity and AI traffic3ChatGPT, Perplexity, Claude, Gemini, AI Overviews, Copilot, Grok, DeepSeek
Peec AIDedicated trackerVisibility, position, sentiment and cited sources across a prompt set; CSV export, Looker Studio connector and API4ChatGPT, Perplexity, Gemini, AI Mode
Otterly.AIDedicated trackerBrand mentions and cited sources, prompt research, competitor benchmarking and content audits5ChatGPT, Perplexity, AI Overviews, AI Mode, Gemini, Copilot, Claude
ScrunchDedicated trackerVisibility monitoring, a real-time feed of AI bot crawling, and an Agent Experience Platform that serves a machine-readable version of a site to AI agents6ChatGPT, Perplexity, Claude, Gemini, Copilot and others
Semrush AI Visibility ToolkitSEO suite moduleAI Visibility Score, competitor research, prompt research, share of voice and sentiment, daily prompt tracking and an AI search site audit7ChatGPT, Google AI Mode and other AI platforms
Ahrefs Brand RadarSEO suite moduleCustom prompts checked daily, weekly or monthly, with the fan-out queries behind them; an AI Visibility Index across a large prompt database8AI Overviews, AI Mode, ChatGPT, Perplexity, Gemini, Copilot, Claude
HubSpot AEO GraderFree one-time graderScores a brand out of 100 on sentiment, presence quality, brand recognition, share of voice and market competition9OpenAI, Perplexity, Gemini
Arobis AI AI visibility checkerFree one-time checkChecks whether AI engines mention a brand for buyer promptsChatGPT and other engines
GA4 and Search ConsoleYour own analyticsAI referral sessions, and impressions in Google's AI featuresSee the next two sections

If you want a snapshot before committing to any of this, start with our roundup of free one-off AI visibility checkers. The CSV above then handles a first baseline. Past a few hundred rows a week, logging by hand eats the time you should spend acting on the results, and that's usually when a paid tracker pays for itself.

How do you track AI referral traffic in GA4?

Start with the built-in AI Assistant channel, then add a custom channel group to catch what it misses. Google's traffic-source documentation defines AI Assistant as traffic from sources like ChatGPT, Gemini, Deepseek, Copilot or Grok, with the medium set to ai-assistant. Google AI Overviews and AI Mode are excluded and fall under Organic Search10.

Three gaps remain. ChatGPT adds utm_source=chatgpt.com to many citation links but no utm_medium, so some of those visits can land in Unassigned11 (SearchPilot). Apps and browsers that strip the referrer send visits to Direct. And Google hasn't published the full list of recognized assistants. A custom channel group placed above Referral, matching session source against this pattern, covers most of it:

chatgpt\.com|chat\.openai\.com|perplexity\.ai|gemini\.google\.com|claude\.ai|copilot\.microsoft\.com|deepseek\.com|grok\.com

Then build an exploration with landing page, session source and key events, so you can see which pages AI-referred visitors land on and whether they convert.

What can Google Search Console show about AI Overviews and AI Mode?

Impressions in Google's AI features, yes; clicks from them in isolation, no. Google's AI features documentation says sites appearing in AI features are included in the Performance report under the Web search type, together with regular results12.

Google announced dedicated generative AI performance reports on June 3, 2026, with separate views for Search and Discover, and rolled them out to all websites on August 31, 2026. The announcement lists impressions as the metric, with breakdowns by page, country, device (Search only) and date13 (Google Search Central Blog). Use it to see which pages appear in AI features and how often, and keep treating AI Overviews and AI Mode clicks as part of organic search in both Search Console and GA4.

A one-page weekly report template

Lead with rates against competitors, then what moved and what you'll do about it. Copy this outline:

1. Headline: mention rate, citation rate, share of voice (all engines)
   this week | last week | 4-week average
2. By engine: same three numbers for ChatGPT, Perplexity, Gemini,
   Claude, Copilot, AI Overviews, AI Mode
3. By stage: problem | category | comparison (brand prompts reported
   separately, sentiment and accuracy only)
4. Competitors: share of voice for you + top 5, leader marked
5. Top 10 cited domains for your prompts: yours | competitor | third party
6. Sentiment: new or repeated claims about you + the source each traces to
7. Traffic: GA4 AI Assistant sessions and key events, top landing pages,
   Search Console generative AI impressions
8. Context: pages shipped, mentions earned, known model or engine updates
9. Next actions: 3 to 5 tasks, each tied to a prompt group or cited domain

For industry numbers to set your results against, see our AI search statistics. For how this sits next to SEO reporting, see GEO vs AEO vs SEO.

Frequently asked questions

How many prompts do I need to track AI visibility?

Start with 30 to 100 prompts a real buyer would ask, spread across problem, category, comparison and brand stages. Below about 30, results swing on a handful of answers; above a few hundred, manual review gets impractical. Keep the core set fixed so week-over-week numbers stay comparable, and put new prompts in a separate group.

How many times should I run each prompt?

At least three times per engine for a baseline, then report the share of runs where you appear. No standard exists: the SparkToro and Gumshoe study ran each prompt 60 to 100 times and left the question open. As a rule of thumb, 48 total runs gives roughly a ±13 point margin, so only trust large changes until you have a few hundred runs.

Can Google Search Console show AI Overviews traffic?

Partly. Clicks and impressions from AI Overviews and AI Mode sit inside the standard Performance report under the Web search type, mixed with regular results. The generative AI performance reports, rolled out to all sites on August 31, 2026, show impressions in AI features by page, country, device and date. Google's announcement lists impressions as the metric.

Does GA4 track ChatGPT traffic automatically?

Mostly. GA4's Default Channel Group includes an AI Assistant channel for visits from sources like ChatGPT, Gemini, Deepseek, Copilot and Grok, but only when a referrer or tag identifies the source. Visits from apps that strip the referrer can land in Direct or Unassigned, and Google AI Overviews and AI Mode clicks count as Organic Search.

What is a good AI share of voice?

There's no universal benchmark, since share of voice depends on how many competitors you track and which prompts you use. Compare against the same competitor set over time and against the category leader. The gap that matters most is on comparison and alternatives prompts, where recommendations most directly shape buying decisions.

What is the difference between a mention and a citation?

A mention is your brand named in the answer text; a citation is a link to one of your pages as a source. You can be mentioned without being cited, for example when a review site is the cited source, and cited without being mentioned, when your page supplies a fact. Track both separately.

Sources

Cite this page AI Ranked Editorial. "How to measure AI visibility." AI Ranked, October 11, 2026. https://airanked.ai/guides/measure-ai-visibility