How to measure AI visibility without inventing one magic score
- AI visibility is five separate measurements: discovery, narrative fidelity, citation, head-to-head choice, and recommendation.
- A single dashboard score hides which one failed, so it hides the fix.
- Record the full evidence per run: prompt, date, model, retrieved sources, asserted citations, and failures.
- One run is an anecdote. Locked prompts, repeated on a cadence, are a measurement.
- Compute a roll-up score last, after the layers, if at all.
Your AI visibility dashboard says 62. Last month it said 58. Nobody in the room can say what moved, why it moved, or what to do next. Meanwhile a prospect asked ChatGPT for a shortlist this morning, and your competitor was the answer.
The gimmick era just lost its cover. In its AI features guidance, Google says plainly that you do not need special AI files, special markup, or a special writing style to appear in its AI search features, and warns against tools that “promise ranking success or claim to use ‘internal’ Google metrics.” What remains is the real work: measuring how AI systems treat your company, and reading the measurement well enough to act.
The takeaway
AI visibility is five separate measurements: discovery, narrative fidelity, citation, head-to-head choice, and recommendation. A single score that averages them hides which one failed, and that hides the fix.
What is AI visibility?
AI visibility is the measured behavior of AI answers about your company: whether ChatGPT, Claude, Perplexity, Gemini, and Grok find you, describe you accurately, use your pages as evidence, prefer you in comparisons, and recommend you when a buyer asks a real buying question. It is a measurement problem before it is a marketing problem. You cannot improve an answer you have not captured, and you cannot trust a capture that threw away the evidence.
Why one score hides the fix
Each layer fails for a different reason, and each failure has a different owner. A company that never appears has a discovery problem, which is usually an evidence and entity problem. A company that appears with a three-year-old description has a narrative problem, which is a content and website problem. A company that gets described but never cited has a source problem. A company that gets cited but loses every comparison has a positioning problem. Average those into one number and the number goes up while the layer that costs you deals goes down.
How much this matters depends on how big your brand is. A June 2026 preprint from Pratyush Kumar of Ranqo, analyzing 100,000+ prompt responses across 100+ brands from March through May 2026, found household-name brands appeared in 73% of relevant AI answers on a first run, mid-market brands in 44%, and niche brands in 11%. Scope those numbers honestly: one vendor's tracked panel, over one spring, in a preprint. The direction still matches what I see in practice: the smaller the brand, the more the work is discovery and evidence, and the less a blended score will tell you.
Break visibility into five questions
This is the model I use. Each layer is a separate question with a separate record. Two terms of art, defined once: narrative fidelity means the answer describes you accurately and currently. Entity confusion means the answer is about a different company with a similar name.
| Layer | The question it answers | What failure looks like | Where the fix usually lives |
|---|---|---|---|
| 1. Discovery | Does the answer include you at all? | You are absent from answers your competitors appear in | Public evidence and entity clarity; owned by marketing and web |
| 2. Narrative fidelity | Is the description accurate, current, and about the right entity? | Stale positioning, wrong products, or a namesake company's facts | Content and website updates; owned by content and web |
| 3. Citation | Are your pages used as evidence? | The model cites third parties, or prints URLs it never retrieved | Citable source pages; owned by content and PR |
| 4. Head-to-head choice | When forced to compare you against a named competitor, who wins? | Cited but never chosen | Positioning and comparative evidence; owned by product marketing |
| 5. Recommendation | Does the model tell the buyer to pick you, and why? | Recommended with hedges, or recommended for the wrong reasons | The whole system above, retested after each change |
Two of these deserve a warning label. A mention count is worthless until you verify the entity behind each mention. And citation has a trap inside it that most reporting misses, which is the next section.
When a citation counts as evidence
The providers themselves distinguish what was retrieved from what was cited. OpenAI's web search documentation exposes inline citations and a separate sourcesfield, “the complete list of URLs the model consulted when forming its response,” and notes the sources list is often longer than the citation list. Anthropic's web search tool returns retrieved results as their own blocks, with per-passage citations carried separately. Gemini's grounding metadata links each cited text segment to a source URL by character position.
The consequence: a URL printed in an answer's prose is an assertion. It becomes evidence only when it matches the source set the provider actually retrieved. Suppose a run reports twelve citations for your brand. You check them against the provider's retrieved source set and eight match, two point at pages the provider never fetched, and two belong to a similarly named company. The honest citation count is eight, and the last two belong in your entity-confusion column. A dashboard that reports twelve is not lying on purpose. It just never checked. In my research work, a finding is accepted only when its cited URL passes that gate; an unmatched citation is preserved as dropped evidence rather than promoted to a result. If your measurement tool cannot tell you which of those two things a citation was, your citation count is a guess.
Record more than the answer
Here is the metric dictionary. For every prompt run, keep:
- The exact prompt, from a locked library, with a version number
- Date and time of the run
- Model and version, and whether the model was allowed to search the web (grounding)
- The full answer text, raw, before any scoring
- The provider-retrieved source set, kept separate from asserted citations
- The entity check: is this answer about you or a namesake?
- The outcome per layer: discovered, described, cited, chosen, recommended
- Non-success states: refusals, errors, timeouts, and empty retrievals
That last line matters more than it looks. Providers treat failure as data. Anthropic's API returns distinct error codes for rate limits and over-limit searches. It also separates a failed search from a search that ran and found nothing. Your measurement should too. A dashboard that silently drops failed runs reports a cleaner market than the one you are actually in.
Google now supplies a first-party slice of this record for its own surfaces: in June 2026 it added a generative AI performance report to Search Console, and its guidance points there for measuring AI-feature visibility. Use it. Then notice what it covers: your content's performance in Google's AI surfaces. That is roughly the discovery and citation layers, on one platform. The remaining layers, and the other four platforms, still need their own record.
And one run of all this is still an anecdote. Answers move with model updates, phrasing, location, and time. Lock the prompts, run the same library on a cadence, and read trends. The measurement is the repeated method, never the screenshot.
What a score is actually for
Executives need a summary, and that is a fair need. Roll the layers up into a score at the end if it helps the board conversation, under three conditions: the components stay inspectable underneath it, the definitions stay stable across runs, and nobody treats a single run as a distribution. A score sold without the layers underneath is a hiding place.
Also keep the score honest about what it cannot claim. Visibility measurement is diagnostic. The Pew Research Center found that U.S. users who saw an AI summary clicked a traditional result in 8% of visits, against 15% without one, with links inside the summaries clicked in just 1% of those visits. That is Google, one country, March 2025. But it points at the operating reality: influence is moving upstream of the click, which is exactly why you measure the answer itself. It is also why no visibility metric, mine included, proves revenue causation on its own. Pair the measurement with pipeline evidence, and say “directional” out loud.
How I run this
My measurement practice comes from a production research engine I helped architect and build at Trinzik, the Austin AI company I co-founded. Trinzik operates it for clients. The engine runs locked prompt libraries across ChatGPT, Claude, Perplexity, Gemini, and Grok, raw and normalized outputs preserved, provider-retrieved sources captured separately from model-asserted citations, official domains verified so the entity is right, and refusals, errors, and empty results recorded as first-class states. The method is the point, and it is the same method this article just handed you: query, decode, engineer, verify.
You can run the first pass yourself. Write twenty real buying questions for your category. Lock the wording. Run them across the five platforms. Record the full evidence per run, score the five layers separately, and look at where you actually lose. Most companies discover the problem is not where the dashboard pointed.
Questions worth asking next
What is an AI visibility score?
A single number a monitoring tool computes from how often AI systems mention, cite, or recommend a brand. The number is a roll-up. It is useful for tracking direction, and useless for choosing an action, unless you can open it up and see the separate measurements underneath: discovery, narrative fidelity, citation, head-to-head choice, and recommendation.
Is AI visibility measurement different from SEO reporting?
Yes. SEO reporting measures your position in a ranked list of links. AI visibility measures the answer itself: whether ChatGPT, Claude, Perplexity, Gemini, and Grok describe you accurately, use your pages as evidence, and recommend you when a buyer asks. The record you keep is different too: answers, retrieved sources, and citations instead of rankings.
How often should a company re-measure AI visibility?
On a fixed cadence with locked prompts, so runs are comparable. Monthly works for most companies. Answers move with model updates, retrieval changes, and your own publishing, and a single run is one sample. Watch the trend across repeated runs instead of a single screenshot.
Sources
- Google Search Central, "AI features and your website" guidance. https://developers.google.com/search/docs/fundamentals/ai-optimization-guide
- OpenAI, Web search tool guide (citations and the sources field). https://developers.openai.com/api/docs/guides/tools-web-search
- Anthropic, Web search tool documentation (retrieved results, citations, error states). https://platform.claude.com/docs/en/docs/agents-and-tools/tool-use/web-search-tool
- Google, Gemini API grounding with Google Search. https://ai.google.dev/gemini-api/docs/google-search
- Perplexity, Search API guide. https://docs.perplexity.ai/guides/search-guide
- xAI, Grok search tools documentation. https://docs.x.ai/docs/guides/tools/search-tools
- Pratyush Kumar (Ranqo), "Generative Engine Optimization at Scale," arXiv preprint 2606.20065, June 2026. https://arxiv.org/abs/2606.20065
- Pew Research Center, "Google users are less likely to click on links when an AI summary appears in the results," July 22, 2025. https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/
- Google Search Central Blog, "Introducing Search Generative AI performance reports in Search Console," June 2026. https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports
About the practice behind this guide
I am Bob Michaels, a Web and AI Systems Architect in Austin, Texas. I have built the web since 1994, and today I run AI visibility measurement, complete web presence transformations, and custom AI system builds for organizations that want one accountable owner across all three. The first engagement is a head-to-head assessment: I run real buying questions across the five major AI platforms and show you who they recommend in your category, and why.
Evaluating me for an AI leadership role instead? The work record is here.