Fundamentals · 5 min read

What Is AI Visibility? Definition, Metrics, and How to Measure It

Last updated: Written by Rastislav MolcanMethodologyEditorial policy

Definition

AI visibility is the measurable presence of a brand inside the answers that AI engines generate — ChatGPT, Perplexity, Gemini, Claude, Copilot, and Google AI Overviews. It is to answer engines what rankings are to Google: the observable, trackable footprint your brand leaves in what a buyer actually sees. It answers a simple question: when the market asks AI about your category, are you in the answer, how often, how favorably, and does it move demand you can measure?

The four core metrics

Serious AI visibility programs standardize on four metrics. The first two describe presence, the third describes quality, and the fourth connects presence to a business outcome you can observe in your own analytics. Each is computed against a fixed, representative prompt set — typically 30 to 100 prompts drawn from real buyer language — so numbers are comparable from one measurement period to the next.

  • Share of voice (SoV) - your brand's share of all brand mentions across responses to the tracked prompt set
  • Citation rate - the percentage of responses that cite at least one of your domains as a source
  • Sentiment - how favorably your brand is characterized when it does appear
  • Branded-search lift - the change in branded search demand attributable to your measurement period

The formulas

Fix the prompt set, the engines, and the number of repeat runs before you compute anything, because changing the denominator mid-stream makes trends meaningless. Mentions and citations are different events: an engine can name your brand without linking to you, and cite your domain without naming you prominently. Track both. Prompt coverage — the percentage of prompts where you appear at all — is a useful companion figure to SoV, since a brand can hold decent share of voice while being entirely absent from half the prompt set.

Share of Voice (SoV)
  = (your brand mentions ÷ total brand mentions across all responses) × 100

Citation Rate
  = (responses citing at least one of your domains ÷ total responses) × 100

Net Sentiment
  = ((positive mentions − negative mentions) ÷ your total mentions) × 100

Branded-Search Lift
  = ((branded impressions, measurement period − branded impressions, baseline)
     ÷ branded impressions, baseline) × 100

A worked example (illustrative)

The following example is entirely hypothetical — Acme Scheduling is an invented brand and every number is illustrative, not a benchmark. Acme tracks 20 prompts covering category, comparison, and alternatives intent. Because identical prompts can return different responses across repeat runs — variability Ahrefs has documented in its own citation testing — Acme runs each prompt three times per period, producing 60 responses. Across those 60 responses, engines mention vendors 138 times in total; Acme accounts for 23 of those mentions. Acme's domain is cited as a source in 9 of the 60 responses. Of Acme's 23 mentions, 14 are positive, 7 neutral, and 2 negative. Over the same window, branded impressions in Search Console rise from a 16,800 baseline to 19,320.

Prompt set: 20 prompts × 3 runs each = 60 responses

Share of Voice
  All vendor mentions across 60 responses: 138
  Acme mentions: 23
  SoV = 23 ÷ 138 × 100 = 16.7%

Citation Rate
  Responses citing acme-example.com: 9
  Citation rate = 9 ÷ 60 × 100 = 15.0%

Net Sentiment
  14 positive, 7 neutral, 2 negative (of 23 mentions)
  Net sentiment = (14 − 2) ÷ 23 × 100 = +52%

Branded-Search Lift
  Baseline branded impressions (prior 4 weeks): 16,800
  Measurement period (next 4 weeks): 19,320
  Lift = (19,320 − 16,800) ÷ 16,800 × 100 = +15%

What real citation distributions look like

Your citation rate is shaped by where engines actually pull sources from, and the published research is consistent on the big picture. Muck Rack's May 2026 analysis of more than 25 million links in ChatGPT, Claude, and Gemini across 17 industries found that earned media made up 84% of links while paid or advertorial content contributed just 0.3% — and each engine has a distinct diet: Wikipedia was the top cited domain for ChatGPT, PubMed Central for Claude, and Reddit for Gemini. Profound's 27-million-citation analysis found owned sources supplied only 4.3% of citations on category-level prompts, so a low owned-domain citation rate on generic prompts is normal, not a failure. Search rankings still matter, unevenly by engine: Semrush's study of 5,000 queries and 150,000 citations found Perplexity cited a domain from Google's top ten in 91% of cases, while an Ahrefs test of 3,311 head terms found ChatGPT overlap of only 31.8% at the domain level. Measuring one engine gives you a partial — and skewed — view.

Why repeat runs, turn depth, and caveats matter

Three findings should discipline how you read your own numbers. First, single runs mislead: repeat-run variability means every metric should be an average across multiple samples. Second, position in the conversation matters — Profound's analysis of roughly 730,000 cited US-English ChatGPT conversations from late 2025 found citation incidence fell from 12.6% at the first turn to 4.5% by turn ten, so first-turn visibility is where most citation opportunity lives. Third, treat lift claims skeptically, including your own. Stacker's March 2026 study of 87 distributed stories for 30 brands across eight AI platforms reported a 239% median citation lift, but the figure is vendor-reported and the study itself calls the analysis observational, not causal. Finally, audit citation quality: a 2026 preprint covering 55,393 queries over 40 days found 11% of 98,020 atomic claims were unsupported by their attached citations — a provisional estimate, but reason enough to check what a citation actually says about you.

Measuring branded-search lift with first-party data

Branded-search lift is the metric that ties AI visibility to demand, and it now has a first-party data source. In June 2026, Google added Search Generative AI performance reports to Search Console for a subset of sites, covering AI Overviews and AI Mode by page, country, device, and date. Combine that with a branded-query filter on standard Search Console data: establish a baseline of at least four weeks, then compare like-for-like periods to control for seasonality. Keep expectations honest — lift is correlational unless you can isolate the change, and Google's own AI-search guidance says the same foundational SEO practices apply to its AI features and explicitly warns against inauthentic mentions as a shortcut.

From measurement to action

Metrics without a workstream are noise. Pair each metric with an owner and a lever. Low share of voice points to content gaps on the prompts where competitors dominate. A weak citation rate points to earned-media and authority work — the Muck Rack finding that 84% of AI-read links are earned media tells you where that effort compounds. Negative sentiment points to messaging, reviews, and how third parties describe you. Flat branded-search lift after visibility gains suggests you are appearing on prompts that do not match real buyer intent — revisit the prompt set before revisiting the strategy.

Sources

Frequently Asked Questions

>How many prompts do I need to measure AI visibility?

A set of 30 to 100 prompts covering category, comparison, alternatives, and use-case intent is a practical range. Run each prompt multiple times per period — identical prompts can produce different responses, so single runs produce noisy, misleading metrics.

>What is a good share of voice?

There is no universal benchmark — SoV only means something relative to your own baseline and your competitors on the same prompt set. Track the trend period over period, and pair it with prompt coverage so a few strong prompts don't mask absence everywhere else.

>Do AI citations just follow Google rankings?

It depends heavily on the engine. Semrush's 150,000-citation study found Perplexity cited a domain from Google's top ten in 91% of cases, while Ahrefs' test of 3,311 head terms found ChatGPT domain-level overlap of only 31.8%. Ranking well helps most on Perplexity and Google AI Overviews, least on ChatGPT.

>Can I measure any of this for free?

Partially. Google added Search Generative AI performance reports to Search Console in June 2026 for a subset of sites, covering AI Overviews and AI Mode. That is first-party but Google-only; measuring share of voice and citations across ChatGPT, Perplexity, and other engines still requires running your own prompt set or using a tracking tool.

>Should I expect citation lifts like the ones vendors report?

No — treat them as observed outcomes, not promises. Stacker's 239% median citation lift, for example, is vendor-reported from an observational study of 87 stories, and Stacker itself notes the analysis cannot establish causation. Use such figures as evidence that distribution can move citations, not as a forecast.

Related guides