6 min remaining
0%
SEO

How to Measure AI Search Visibility: Citation Rate, Share of Voice, and the Noise Floor

Learn how to accurately measure AI search visibility with key metrics like Citation Rate and Share of Voice, ensuring your SEO strategies are effective.

6 min read
Progress tracked
6 min read·

TL;DR: There is no single cross-platform AI visibility score, and any dashboard that sells you one is hiding its methodology. Measure three things — Citation Rate, competitive Share of Voice, and Prominence — under a fixed protocol: frozen prompt set, named platforms, market, language, run counts, exclusion rules, all defined before testing starts. Establish a baseline, learn your noise floor, and only then call a change an improvement.

I am Akira. The most expensive sentence in AI SEO right now is "your AI visibility score is 70." Seventy what, measured how, against which prompts, on which platform, in which market, in which language? A score without a protocol is astrology with a dashboard. Here's the GEO measurement framework we actually run — and the rules that keep it honest.

The Three Core Metrics

1. Citation Rate: How Often Do AI Answers Cite Your Domain?

Citation Rate = valid runs citing your domain ÷ total valid runs.

18 cited runs ÷ 60 valid runs = 30% Citation Rate. The percentage only becomes meaningful when the denominator is clear. Sixty prompts tested once and thirty prompts tested twice both produce 60 runs — they are not the same measurement. Platform, market, language, and run count all change what the number means — which is why, in any serious GEO report, Citation Rate never appears without its protocol attached.

2. Competitive Share of Voice: How Often Do You Appear vs Named Competitors?

Competitive SOV = your brand appearances ÷ total appearances by your brand plus a fixed, named competitor set. If you appear 35 times and named competitors appear 65 times combined: 35%.

Two rules make this comparable. First, the competitor set stays frozen between periods — comparing yourself against three competitors in January and ten in March produces two numbers that share a formula and nothing else. Second, one brand appearance per run, whether the answer names you once or five times.

3. Prominence: What Role Did You Play in the Answer?

Citation Rate counts presence; Prominence classifies it. Four levels:

| Level | Classification | Rule | |---|---|---| | P1 | Primary recommendation | The brand is the main recommendation | | P2 | Named alternative | One of several options, not clearly primary | | P3 | Cited source | Domain cited to support a fact; brand not recommended | | P4 | Passing mention | Named without attribution or recommendation |

"Brand A is our top recommendation. Brand B and C are also worth considering" — A is P1, B and C are P2. If the role is genuinely ambiguous, classify downward. Inflating your own prominence is how measurement programs lie to themselves.

The Measurement Rules That Keep It Honest

What counts as a valid run. A complete response under intended conditions — full prompt submitted, right platform, right market, right language, no truncation. Every excluded run gets logged with its reason. And the cardinal rule: never re-run failed prompts until you get an answer you like. That's not measurement; that's shopping.

Citations and mentions are separate fields. A citation attributes information to your domain or links it as a source. A mention names you without attribution. Not cited ≠ not mentioned — models can name you from training data without ever touching your site. Combine the two into one "appearance" metric and you've destroyed the distinction that tells you whether your content is working.

Repeated appearances count once per run. Three links to your domain in one answer = one cited run. The extra prominence is captured in Prominence, not by inflating counts.

The Noise Floor: The Number Nobody Wants to Publish

Run the same frozen prompt set twice under identical conditions. Baseline 1: 30% Citation Rate. Baseline 2: 27%. That 3-point gap is your observed noise floor — the amount AI visibility moves when nothing intentional changed.

Now the discipline: if a later measurement moves from 30% to 32%, that is not an improvement. It's inside the noise. Keep individual baselines in the report — never just the average — because the noise floor itself shifts over time. It's not a formal confidence interval; it's an observed measure of variation that separates signal from weather. Any AI SEO agency reporting month-over-month "gains" smaller than its own noise floor is selling you variance as progress.

The Sampling Protocol Is Part of the Measurement

Before trusting any number — yours, a vendor's, ours — the protocol must name: the exact prompt wording, the AI platforms and surfaces, the market and location, the language, runs per prompt, the competitor set, the testing window and cadence, account/personalization settings, citation and mention definitions, counting rules, and exclusion rules.

Change any one of them between rounds and the two rounds aren't comparable — even with the same metric name. This protocol is the real difference between SEO reporting and GEO reporting: rankings have one source of truth; citations have none. Language deserves special emphasis: an English prompt and a Traditional Chinese prompt don't retrieve the same documents. They're two measurements, not one. Blend them and you've averaged two unrelated systems.

Where First-Party Data Fits

Google Search Console's Generative AI Performance report and Bing Webmaster Tools' AI Performance data are first-party evidence for their own platforms — impressions on one side, citations and grounding queries on the other. They're real, but they're platform-locked. Cross-platform comparison across ChatGPT, Perplexity, and Gemini still requires prompt sampling under your own protocol. Neither replaces the other.

For a small-to-medium measurement program: 30–60 fixed prompts, repeated runs, quarterly cadence, one analyst's rules written down. That's enough to separate movement from noise — and it's what a serious AI SEO measurement program looks like. This is the protocol behind our GEO Audit — the 48-hour baseline every engagement starts from — and the KPI layer is documented in the LLM KPI framework. If you want to see the measurement before you buy it, run the free GEO Scorecard.

Stop collecting scores. Start publishing protocols.

Mercury Technology Solutions: Accelerate Digitality.