# How to Measure Your AI Visibility: Share-of-Voice, Not One Score

> Source: CitedIndex — https://citedindex.com/guides/measuring-ai-share-of-voice
> By: CitedIndex Editorial, Research desk
> Claims last verified: 2026-08-23
> Last updated: 2026-09-05
> Read time: 11 min read
> Covers: ChatGPT, Perplexity, Claude, AI Overviews

What "AI visibility" actually means for a brand, why the tools disagree on a single visibility score, and how to audit and track your share-of-voice across ChatGPT, Perplexity, Claude and AI Overviews, including a free, no-tool method.

AI answers can contain ordered recommendations, but those positions are not a stable search-results rank. An answer may name a product, cite a source, do both, or do neither. Auditing your **AI visibility** means defining those events separately and comparing their rates across a fixed prompt panel, named competitors and documented observation windows. This guide explains that workflow, including a manual spreadsheet method.

### Abstract · the bottom line

Measure **mention rate** and **source-citation rate** separately: eligible completed runs with the event divided by eligible completed runs where that event can be assessed. Compare each rate against the same competitor set; state the denominator whenever you call a result share-of-voice. A saved single answer is evidence of that answer, not a reliable population estimate. Keep a fixed prompt panel, record engine, surface, model, locale, mode and date, and use repeats to describe uncertainty rather than enforce a universal minimum. Claude supports web search; whether any engine searched must come from the observed run, not its brand name. Tools such as [Otterly AI](/otterly-ai), [Profound](/profound) and [Peec AI](/peec-ai) can automate parts of this workflow; a manual spreadsheet works too.

## Why "Rank" Doesn't Apply to AI Answers

Google rank tracking records a position on a particular results page. AI assistants can instead combine generated text with retrieved material, and may answer without web search. The same prompt can produce different wording, recommendation order and source links. Record those as separate observations rather than treating one answer position as a stable search ranking.

**The consequence:** a single saved run can establish that a particular answer mentioned a product or cited a page. It cannot establish how often a wider audience sees that event. A fixed panel asks a narrower, answerable question: for these prompts and observation conditions, how often did each event occur, and how uncertain is the comparison with competitors or an earlier window?

## What Share-of-Voice Actually Measures

Share-of-voice is a ratio, but the numerator hides three distinct events that get conflated if you don't separate them:

### 1. Mention

The answer text names the intended brand or product. Define its accepted names, domains and disambiguating context before scoring. A generic or shared name without enough context is ambiguous, not a confirmed mention. A mention does not establish that the model was trained on your site, retrieved it, or used it as a source. Keep prompted identity checks (where you supplied the name or URL) separate from unprompted category discovery.

### 2. Citation

The response emits a source citation or annotation whose URL points to the domain or page you are measuring. In a consumer UI, save the visible citation link or attributed Sources entry; in an API, save the provider citation fields and their URLs. Keep a broader retrieved-source list separate if it does not establish a citation in the answer. A plain URL in generated prose is not automatically a provider source annotation. For **directory-source citation**, the URL must point to the directory (for example, citedindex.com), not merely to a product the directory lists. An answer that names Peec AI or links to its vendor site has not thereby cited CitedIndex. A source annotation establishes an emitted reference, not independently verified support for every claim or proof of a click.

### 3. Sentiment / framing

When you are mentioned or cited, is the framing accurate and favorable, neutral, or wrong? A brand can have high share-of-voice and still lose deals if the engine consistently describes it with an outdated price, a discontinued feature, or a direct "however, X is better for most users" qualifier. Track this qualitatively, it's the field a pure-numeric dashboard misses.

Report all three separately, with counts and denominators. A high mention rate and low source-citation rate describes two different outcomes; it does not diagnose a training-data or synthesis defect by itself. For a per-run citation rate, count a target at most once per run even if several URLs from it appear. A share of all competitor mentions or citations is a different metric with a different denominator. Name that choice explicitly, and do not assume figures from different panels or categories are comparable.

## Building a Measurement Methodology

The methodology is the whole game: a small, well-documented panel is more useful than a large panel whose conditions keep changing. Make the following choices before interpreting the results.

### The prompt panel

Build a manageable fixed panel in language supported by customer questions or other stated evidence. Include relevant category, comparison and problem-first prompts. Choose panel size for the decision and report its selection method; 20–50 prompts is an illustrative planning range, not a minimum or a representative sample of all buyers.

### The competitor set

Fix a named competitor set before comparing observations. Three to eight names can be a manageable example, not a methodological requirement. Use actual alternatives buyers consider, such as sales-loss reasons, and document additions or removals with a new panel version. Keep an unchanged subset for comparisons when possible; do not silently join different denominators into one trend.

### Observation record

Save the exact prompt and panel version, response text, citation/source URLs and capture time. Record the provider, consumer UI versus API endpoint, model/version if exposed (otherwise unknown), language/locale, location setting, search/mode settings and session context. Log search requested or enabled separately from search observed: a selected tool or a prompt saying "search" does not prove it ran. Save visible search activity or API tool-call evidence where available, and use unknown when the surface does not expose it.

Log failed requests, unavailable modes, truncated captures and unassessable citation fields as missing observations with reasons, not zeros. Zero means a usable response was inspected and the defined event was absent. Report attempted, completed, missing and eligible counts so outages cannot improve a rate. Define whether a no-Overview Google result belongs in the overall search-opportunity denominator or outside an Overview-only rate, and report that state separately.

### Repeats per prompt

Choose repetitions for the question, budget, observed variability and precision you need; there is no universal five-to-ten-run minimum. Repeats show how a prompt varies under the recorded conditions, but near-identical runs from one session or time window are not necessarily independent evidence. Keep sessions and timing consistent or explicitly stratified, report per-prompt results, event counts and denominators, and do not let extra repeats on one prompt silently give it more weight. Where you report an uncertainty interval, state its method and assumptions; a naive independent-trials interval can be too narrow when runs share prompt, session or time effects. A small manual panel remains useful as a dated descriptive audit, not a market-wide estimate.

### Cadence

Choose observation windows around the decision and how quickly the category changes. Weekly or monthly checks can be useful examples, not universal floors or a reason to automate an ongoing schedule. Fix a planned comparison window and repeat allocation before looking at the outcome, and document any later change. More frequent collection is only useful if it answers a question the existing observations cannot.

## Per-Engine Measurement Notes

Use a common core of buyer prompts when comparing engines, with any engine-specific additions reported separately. The consumer app and a provider API are different measurement surfaces, even when their branding or model names overlap. Preserve each response format rather than assuming that every Sources panel, retrieval result or citation array means the same thing. The notes below describe documented capabilities, not new measurements of our own.

### ChatGPT

For an API panel, [OpenAI documents](https://developers.openai.com/api/docs/guides/tools-web-search) web search in the Responses API, a `web_search_call` output item, and `url_citation` annotations. Making the tool available does not by itself show that the model used it; inspect the returned activity and annotations. A ChatGPT app observation is a separate UI record, not a measurement of that API endpoint. Record the app mode and the source links actually shown. Do not label an answer "training only" merely because you did not observe web search: supplied conversation context is another possible input.

### Perplexity

Record the Perplexity product surface, endpoint/model and search settings. The [Sonar API reference](https://docs.perplexity.ai/api-reference/sonar-post) distinguishes `citations` from `search_results` and documents controls that can disable search or let a classifier decide whether it is needed. Inspect the actual response instead of assuming every Perplexity answer searched and cited. The API record is not a substitute for a consumer-app observation, and this guide supplies no measured basis for ranking its variability above other engines.

### Claude

Claude is not a training-only engine. [Anthropic documents web search in the consumer app](https://support.claude.com/en/articles/10684626-enable-and-use-web-search), enabled per chat and subject to workspace settings, with a search indicator and source citations. In the [Claude API](https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool), a developer supplies the web-search tool; Claude decides when to use it, and search-derived output has structured `web_search_result_location` citations. Record enabled/requested search separately from observed tool use or UI activity, and preserve the cited URLs. A flat Claude line is an observation to investigate, not evidence that new content can appear only after a model update.

### Google AI Overviews

Record whether an AI Overview appeared and which source links it displayed; keep AI Overviews, AI Mode and Gemini observations separate. [Google documents a dedicated Search generative AI performance report](https://support.google.com/webmasters/answer/16984139) covering AI Overviews and AI Mode impressions, with pages, countries, dates and devices dimensions. Dates use Pacific Time (PT). Chart totals normally aggregate by property (by URL when URL-filtered); page rows aggregate by page, so their totals can differ. This is impression reporting, not a query-level citation log or a record of every generated answer. Google says a missing report can reflect access rollout or insufficient impressions; it is not proof of zero visibility. [Search generative AI inclusion](https://support.google.com/webmasters/answer/16908024) is a separate control from Google-Extended training policy; check the effective property and inherited settings. This guide describes public documentation, not access to an owned Search Console account.

## Reading the Data Correctly

A share-of-voice dashboard is only useful if you resist three common misreadings.

- **A percentage-point change has no universal noise threshold.** Read it alongside counts, panel composition, missingness, per-prompt variation and uncertainty. A small change can matter in a sufficiently informative design; a large change can be inconclusive in a sparse or changed panel. A second window adds evidence but does not automatically confirm a trend.
- **Before/after is descriptive, not causal by itself.** Log your edits, competitor launches, model or mode changes and other plausible explanations. To argue that an intervention caused a lift, plan a concurrent comparison or holdout and a suitable controlled design; state the remaining assumptions and confounders. Without those controls, say the observed rate changed after the edit, not because of it.
- **Citation rate and traffic are different outcomes.** A cited URL is not a click, a recommendation or a sale. Compare observed source-citation rates with referral and conversion data separately, noting attribution gaps; co-movement alone does not establish a business effect.

Tools like [Rankscale](/rankscale) and [Nightwatch](/nightwatch) plot the trend line for you; the reading discipline above is what keeps you from acting on a single noisy data point either way.

## Implementation Checklist

Tick items off as you set this up, this page remembers your progress on this device.

#### Panel design

- Manageable prompt panel with documented selection method and buyer-language evidence
- Mix of category, comparison, and problem-first prompts
- Named competitor set and counting rules fixed for each comparison
- Panel changes versioned; comparable subsets retained where useful

#### Sampling rigor

- Repeat allocation chosen for the decision and uncertainty, with counts and per-prompt variation reported
- Observation windows chosen for a defined decision; no automatic repeat schedule required
- Consumer UI and API separated; model, locale, mode, date and requested versus observed search recorded
- Missing observations separated from zeros; no universal percentage-point noise threshold

#### What gets counted

- Product mention, directory-source citation, and sentiment/framing recorded separately
- Target identity and citation URL scope fixed; ambiguous names and prompted identity checks labelled
- Misstatements (stale price, dropped feature, wrong comparison) logged qualitatively

#### Cross-checks

- AI Overview captures kept separate from native Search generative AI impressions; reporting dimensions and aggregation recorded
- Citation-share trend correlated against actual AI-referral traffic in analytics
- Before/after movement described without causal attribution unless a suitable controlled design supports it

## Frequently Asked Questions

### What is AI share-of-voice?

A comparison of visibility across a fixed prompt panel and named competitors under stated conditions. Here, mention rate and source-citation rate each count eligible runs with the event divided by eligible assessable runs. Report the two separately. A share of all competitor mentions or citations uses a different denominator and must be labelled explicitly.

### How is share-of-voice different from a citation count?

A citation count is a raw count under a stated counting rule; it might count URLs or runs, so specify which. A per-run source-citation rate counts eligible runs with at least one target source annotation divided by eligible assessable runs. Compare competitors using the same panel and conditions. Neither normalizing a count nor naming it share-of-voice makes different categories or prompt panels automatically comparable.

### How many times should I run each prompt?

There is no universal minimum. Choose repeats for the decision, variability and precision you need, and report the actual counts and uncertainty. Repeated answers from one prompt, session or time window may be dependent, so raw request count is not necessarily independent sample size. One saved run is useful evidence of that answer; a small panel can be a descriptive baseline without supporting a stable rate or a market-wide conclusion.

### Should I measure mentions or citations?

Both, separately. A mention names the intended product or brand; ambiguous names need review. A source citation is an emitted provider citation or visible attributed source link to the target domain or page. A product mention or vendor URL does not establish that its directory was cited. Save the source URLs and keep retrieval-only results separate; neither a mention nor a citation establishes traffic or endorsement.

### Why does my Claude share-of-voice barely move week to week?

A flat line alone does not explain the mechanism. Claude supports web search in the consumer app and API. Check the recorded surface, model, search settings, observed search activity, panel, missing observations and citation URLs before interpreting a flat result. Do not assume all Claude answers are training-only or that visibility can change only at model releases.

### Do I need a tool, or can I track this manually?

You can run a manual panel in a spreadsheet, it's slower and harder to keep rigorous at scale. Purpose-built tools like [Otterly AI](/otterly-ai), [Profound](/profound), and [Peec AI](/peec-ai) automate the repeat sampling and the per-engine parsing, which is the part that's easiest to do inconsistently by hand.

### Is there a single "AI visibility score"?

There is no universal AI visibility score. Vendor composites can blend mentions, citations, sentiment or other inputs with different weights, while per-engine breakdowns answer narrower questions. The linked listings for [Semrush](/semrush-ai-visibility-toolkit), [Nightwatch](/nightwatch), [RankScale](/rankscale), [Otterly AI](/otterly-ai) and [Peec AI](/peec-ai) are starting points for checking current definitions, not newly verified scoring contracts. Ask for inputs, weights, denominator, surface and missing-data policy before interpreting a change.

### Can I audit my AI visibility for free, without a tool?

Yes, within the access and usage limits of the consumer services you can use. Start with a manageable fixed set of buyer prompts and a spreadsheet; save each exact prompt, complete answer, source URLs and timestamp. Add columns for surface, model if exposed, locale, mode, requested/enabled versus observed search, controlled or ambiguous identity, mention, source citation, sentiment and missing-data reason. Choose and report repeats rather than treating five runs as a required minimum. This is a dated baseline of the observations you captured, not a guarantee of precision or a reason to buy a tracker.

### How does this relate to AEO content work?

Measurement can show whether observed citation rates changed after the content work discussed in [the ranking guide](/guides/how-to-rank-in-chatgpt/). Record a baseline and the intervention date, preserve comparable conditions, and document other changes. Before/after movement alone does not show that structure, schema or credibility work caused it. Causal attribution needs a suitable controlled design and explicit assumptions; without that, report the association and uncertainty.

**About CitedIndex:** This is an explanatory measurement workflow, not a new experiment or a report of freshly collected answers. The provider links above support the specific search and source-format corrections. Our [methodology](/methodology/) explains the directory's editorial approach.

Originally published July 2026. Scoped editorial correction: 5 September 2026, covering search modes, metric definitions, sampling uncertainty and causal interpretation. This correction did not collect new AI answers or re-verify every vendor claim in the guide; historical verification dates remain unchanged.
