Guide · Measurement
Measuring your AI visibility: share-of-voice tracking across engines
There is no "position #1" inside a ChatGPT answer. AI engines don't return a ranked list of ten blue links, they return one synthesized answer that may cite you, a competitor, several competitors, or no one at all. Traditional rank tracking has no unit to count. Share-of-voice (SOV) is the metric that replaces it: the share of a defined set of prompts, run repeatedly over time, in which your brand is mentioned or cited versus your competitors. This — not a single blended number — is what auditing your AI visibility actually involves. This guide covers how to build that measurement correctly, and how to read it once you have it.
Abstract · the bottom line
AI share-of-voice = (prompts where you're mentioned or cited) ÷ (total prompts sampled), tracked against the same competitor set over time. Because AI answers are non-deterministic, a single query proves nothing, you need a fixed prompt panel, a repeat cadence, and enough runs per prompt to separate signal from sampling noise. Citation and mention are different events and must be counted separately. Each engine needs its own panel because retrieval sources differ (Perplexity's live web crawl vs. ChatGPT's Bing index vs. Claude's training-time knowledge). Tools built for this, Otterly AI,Profound, Peec AI, automate the panel and the repeat sampling; a manual spreadsheet works too, just slower — a free AI-visibility audit is entirely possible without buying anything (see the FAQ below).
Why "Rank" Doesn't Apply to AI Answers
Google rank tracking counts a position: your page is #3 for a keyword, or it isn't. That model assumes a stable, ordered list you can screenshot and diff day over day. AI answer engines don't produce that list. ChatGPT, Perplexity, Claude and Google AI Overviews each synthesize a single answer per query, assembled from whichever retrieved passages the model judged most relevant at that moment, and the same prompt run twice can retrieve different sources, especially on engines with continuous crawling like Perplexity.
The consequence: a single query, run once, tells you almost nothing. It's one sample from a probabilistic process. The metric that survives this is share-of-voice: not "am I #1," but "across a representative panel of prompts my buyers actually ask, in what share of runs do I appear at all, and how does that compare to my named competitors, tracked over enough repetitions and enough time to be a trend, not a coin flip."
What Share-of-Voice Actually Measures
Share-of-voice is a ratio, but the numerator hides three distinct events that get conflated if you don't separate them:
1. Mention
Your brand name appears somewhere in the answer text, with or without a link. This is the loosest signal, it means the model's training data or retrieved context contains your brand, but says nothing about whether the user could click through to you.
2. Citation
The answer includes a source reference (a footnote, an inline link, a "Sources" panel entry) pointing at your domain. This is the commercially meaningful event: it's the one that can produce a click. Perplexity, Google AI Overviews and Bing-backed ChatGPT answers expose citations structurally; you can parse them. Claude's citation behavior is more conversational and less consistently structured, which is itself a measurement finding worth logging.
3. Sentiment / framing
When you are mentioned or cited, is the framing accurate and favorable, neutral, or wrong? A brand can have high share-of-voice and still lose deals if the engine consistently describes it with an outdated price, a discontinued feature, or a direct "however, X is better for most users" qualifier. Track this qualitatively, it's the field a pure-numeric dashboard misses.
Report all three separately. A brand with a 40% mention rate but a 5% citation rate has a visibility problem in the model's synthesis, not its knowledge, different fix than a brand with low mentions altogether.
Building a Measurement Methodology
The methodology is the whole game, a good panel run inconsistently is worse than a smaller panel run rigorously. Four decisions determine whether your numbers mean anything.
The prompt panel
Build 20-50 prompts that mirror how real buyers actually ask, not your target keywords verbatim. Include category questions ("best tools for X"), comparison questions ("X vs Y"), and problem-first questions ("how do I do X"). A panel built from keyword-research habits will overweight phrasing no one uses conversationally and undercount your real exposure.
The competitor set
Fix a named list of 3-8 direct competitors before you start measuring, and don't change it mid-quarter, swapping competitors resets your baseline and makes trend lines meaningless. Usethe ranked index or your own sales-loss reasons to pick the set that actually contests your deals.
Repeats per prompt
Run each prompt multiple times per measurement window (5-10 runs is a reasonable floor) because engine output varies run to run, especially on Perplexity and ChatGPT. A single run per prompt per week will show you noise, not a trend, it can swing 20 points from a genuinely stable underlying rate purely from sampling variance.
Cadence
Weekly is the practical floor for a fast-moving competitive set; monthly is enough for a mature category with slow content churn. Match cadence to how often the underlying content (yours and competitors') actually changes. Measuring daily on a panel that never moves just spends budget on noise.
Per-Engine Measurement Notes
Each engine retrieves from a different source and exposes citations differently, so the same panel needs engine-specific handling, not one script that assumes a shared response format. See the companion guide on ranking mechanicsfor how each engine chooses sources in the first place, this section covers how to measure the outcome.
ChatGPT
Bing-backed browsing responses include structured citations you can parse programmatically; plain (non-browsing) responses draw on training data and won't reflect anything published after the model's cutoff, which will understate your true share-of-voice if you only sample the non-browsing mode. Sample both modes and report them separately.
Perplexity
Every answer cites sources by default, which makes Perplexity the easiest engine to measure precisely, but its continuous crawl also means the highest run-to-run variance of the four, because a page published an hour ago can outrank one published last year. Budget more repeats per prompt here than on any other engine.
Claude
Without live web access in most consumer contexts, Claude draws on training-time knowledge and is the slowest engine to reflect new content, measured share-of-voice here changes over months, not weeks. Treat a flat Claude trend line as expected, not a tracking failure; look for step changes around model updates instead of week-to-week movement.
Google AI Overviews
Overviews draw on Google's existing index, so share-of-voice here correlates with (but isn't identical to) your organic rank; it's the one engine where a Search Console query-level view can cross-check your AI panel. Discrepancies between the two are themselves a finding: content that ranks in Search but doesn't get pulled into an Overview usually has an answer-first structure problem, not a relevance problem.
Reading the Data Correctly
A share-of-voice dashboard is only useful if you resist three common misreadings.
- One bad week isn't a trend. Given run-to-run variance, treat single-window drops under ~10 points as noise until a second consecutive window confirms it.
- A competitor's spike often has a dated cause. Check for a launch, a press cycle, or fresh content from them before assuming the model's weighting changed generally.
- Citation share and traffic aren't the same curve. A citation can sit below the fold of a long AI answer and get almost no clicks; correlate citation-share trend with actual referral traffic (most AI engines now send a small but growing amount) before treating a citation win as a business outcome.
Tools like Rankscale andNightwatch plot the trend line for you; the reading discipline above is what keeps you from acting on a single noisy data point either way.
Implementation Checklist
Tick items off as you set this up, this page remembers your progress on this device.
Panel design
Sampling rigor
What gets counted
Cross-checks
Setting up a measurement panel?Get the digest, new guides, checklist updates and who's earning the citations, one short email.
Frequently Asked Questions
What is AI share-of-voice?
The share of a fixed panel of prompts, sampled repeatedly across engines over time, in which your brand is mentioned or cited versus a named set of competitors. It replaces "rank" as the core metric because AI answers don't produce an ordered, stable position to track.
How is share-of-voice different from a citation count?
A citation count is a raw number of appearances; share-of-voice is that number normalized against your competitor set and the total panel size, which is what makes it comparable over time and across categories. A rising citation count with a falling share-of-voice usually means the category is growing faster than your visibility inside it.
How many times should I run each prompt?
5-10 repeats per prompt per measurement window is a reasonable floor given how much AI answers vary run to run, Perplexity and ChatGPT show the most variance, Claude the least. A single run per prompt will show you sampling noise, not a trend.
Should I measure mentions or citations?
Both, tracked as separate fields. Mentions show up in the model's general awareness of your brand; citations are the structural source references that can actually produce a click. Reporting only a blended number hides which one is actually moving.
Why does my Claude share-of-voice barely move week to week?
Claude relies primarily on training-time knowledge in most consumer contexts rather than a live crawl, so it reflects new content on the timescale of model updates, not weeks. A flat trend line there is expected, watch for step changes around model releases instead of treating a flat line as a tracking failure.
Do I need a tool, or can I track this manually?
You can run a manual panel in a spreadsheet, it's slower and harder to keep rigorous at scale. Purpose-built tools like Otterly AI,Profound, and Peec AIautomate the repeat sampling and the per-engine parsing, which is the part that's easiest to do inconsistently by hand.
Is there a single "AI visibility score"?
Not one the whole market agrees on. Some tools deliberately publish a single composite —Semrush's AI Visibility Toolkitand Nightwatch each call theirs an "AI Visibility Score," Promptwatch blends mentions, sentiment and share-of-voice into one trackable number of the same name,RankScale publishes its own single visibility score across its dashboards, and Morningscorerolls it into a 0-100 "GEO Score." Others reject that on purpose:Otterly AI, Peec AIand Knowatoa all market a per-engine breakdown "instead of one blended score," on the grounds that blending erases the exact diagnostic split — a high mention rate but a low citation rate, say — that tells you what to actually fix. Treat any single 0-100 number as a convenient dashboard summary, not a scientific unit; ask what's inside it before trusting a change in it.
Can I audit my AI visibility for free, without a tool?
Yes. The manual-panel method described above (a fixed prompt list, run by hand across ChatGPT, Perplexity, Claude and Google, logged in a spreadsheet) is a genuine, free AI-visibility audit — it just costs time instead of a subscription. Budget an afternoon for a first pass: 20-30 prompts in real buyer language, 5 repeats each, across the engines your buyers actually use, scored for mention, citation and sentiment as three separate columns (never one blended score — see above). That gives you a real baseline before deciding whether a paid tracker is worth automating it.
How does this relate to AEO content work?
Share-of-voice measurement is the feedback loop for AEO: it tells you whether the answer-first structure, schema, and credibility work covered inthe ranking guide is actually moving citations, and which engine responds first. Measure before you change content, so you have a baseline to attribute the change to.
Where to go from here
About CitedIndex: We measure how ChatGPT, Perplexity, Claude, and other AI engines see and cite your brand. This guide reflects patterns we observe tracking AI visibility tools in our directory, every listing researched and quality-gated, never scraped. For more, see our methodology.
Last updated: July 2026. This guide reflects current AEO measurement practices as of this date.