The AI search visibility guide · chapter 6
Measuring AI visibility honestly
How do I measure whether we are being cited in AI answers, without fooling myself?
Karl-Gustav Kallasmaa, Founder & CEOLast updated How do I measure this without fooling myself?
By sampling deliberately and reporting what the sample was. AI visibility is measured the way a poll is measured, not the way a rank is measured - you ask a defined set of questions a defined number of times, count outcomes, and publish the denominator alongside the number.
Everything that goes wrong in this category's reporting goes wrong by dropping the denominator.
Why one screenshot proves nothing
Ask an assistant the same question twice and you can get two different answers citing two different sources. That is not a defect in your process. Two of the reasons are documented by Google itself: AI Mode and AI Overviews "may use different models and techniques, so the set of responses and links they show will vary",[^google-varies] and AI Overviews are shown only when Google's systems judge them additive, so they "often don't trigger."[^google-often-dont-trigger]
So a single observation confounds at least four things: whether an AI answer was generated at all, which surface generated it, model-level variation, and your actual standing. A screenshot in a board deck resolves none of them. Neither does a screenshot of an improvement.
The unit of measurement is a prompt set
Start by writing down the questions. Ten to thirty of them, in the words a buyer would use, covering the shapes that matter:
- Category questions - "best X for Y", where you may not be named at all.
- Comparison questions - "X vs Y", including pairs you are not in.
- Problem questions - the symptom your product fixes, before anyone knows a
category exists.
- Brand questions - "what is X", "is X any good", "X pricing".
- Unflattering questions - "X alternatives", "problems with X".
The set should be stable. Changing the questions changes the number, so a prompt set that drifts quarter to quarter produces a trend line that measures your editing, not your visibility. Add questions in a separate, labelled cohort rather than silently swapping them in.
Record, per run, four things: was the brand mentioned, was our URL cited, were we recommended for the asker's stated situation, and which other URLs were cited. The fourth is the most useful field in the dataset and the one most tools omit.
Sampling depth, and the number nobody wants to hear
Because answers vary, a rate measured over few runs cannot distinguish a real change from ordinary variation. There is no universal minimum - it depends on how big a change you need to see and how variable the platform is - but two rules make the reporting honest regardless of depth:
Always carry the denominator. "Cited in 4 of 14 runs" is a measurement. A bare rate is not. This is the rule Attensira publishes for its own reporting: a rate is a value plus the sample size behind it, and a null value means never measured rather than measured zero.[^attensira-rate-shape]
Do not report movement you cannot distinguish from noise. Attensira's public statement on this is that a delta carries whether the movement is real, and that movement is only reported when it clears the noise floor.[^attensira-noise-floor] Whatever tooling you use, decide the threshold before you look at the results, because afterwards every number is a candidate for a story.
Two corollaries worth adopting:
- Never average across surfaces. The consumer app and the model's API do not
answer alike, so blending them produces a number describing neither. Attensira keeps them separate as a rule.[^attensira-surface-split]
- Never render an untracked platform as zero. If you did not query Perplexity
this month, its visibility is unknown, not nil. Reporting it as 0% invents a failure that did not happen.[^attensira-tracked-readable]
That last one sounds pedantic until you have watched a team spend a quarter "fixing" a platform they never measured.
What Search Console can and cannot tell you
It answers the Google part, partially. Google says sites appearing in AI features are included in overall search traffic in Search Console and reported in the Performance report within the Web search type.[^google-search-console] Included - not broken out. You cannot filter to AI Overviews, so you cannot attribute a change to them from Search Console alone.
What it is still good for:
- Impressions and clicks by query and page, which tell you whether the pages
you are working on are being served at all.
- Coverage and indexing status, which is the eligibility precondition.
- Directional context when combined with analytics. Google says clicks from
result pages with AI Overviews are higher quality, with users more likely to spend more time on the site,[^google-click-quality] so falling clicks with rising engagement is a pattern worth reading rather than panicking about.
What it cannot tell you: anything about ChatGPT, Claude or Perplexity. Those have no console. The only instruments are your own server logs on the request side and sampled answers on the output side, which is precisely why this practice exists as something separate from SEO reporting.
A reporting format that survives scrutiny
Every number you publish internally should carry its own methodology, in one line:
Cited in 6 of 20 runs on ChatGPT (consumer surface), prompt set v3 (24 questions), sampled daily, week of 25-31 August 2026. Previous period: 4 of 20. Change not distinguishable from variation at this sample size.
That last sentence is the one that builds credibility, because it is the one nobody else writes. A team that reports "no detectable change" when there is none is a team whose "detectable change" claims can be believed.
If you need a headline number, derive it from components that stay visible. Ranking a source by how much of an answer it contributed - the GEO study's position-adjusted approach weights contribution by position with exponential decay[^geo-position-metric] - is a more honest composite than a blended percentage, because it is at least defined.
The two instruments, and what each one sees
An honest measurement setup has exactly two data sources, and they answer different questions. Confusing them is the second most common reporting error after dropping the denominator.
Server logs answer "did they read us". Every AI fetch is a request that lands in your access log with a user agent attached. From logs you can tell which agents visit, how often, which pages, and with what status codes - the whole access half of the problem, at full population rather than by sample. What logs cannot tell you is whether any of that reading turned into a sentence in an answer.
Sampled answers tell you "what did they say". Asking the questions and recording the outputs is the only way to observe the output side, and it is inherently a sample. It cannot be complete, because the population of possible questions is unbounded and the answers are generated fresh.
Used together they diagnose. Crawled heavily and never cited means a retrieval or extraction problem, not an access one. Never crawled and never cited means stop writing and go read your CDN rules. Cited without being crawled recently means the model is working from a stale copy or from someone else's description of you - which is the case chapter four is about.
A third source is often proposed and is weaker than it looks: referral traffic from assistant domains. It undercounts badly, because most AI answers are read without a click, and it tells you nothing about the answers where you were mentioned but not linked. Use it as a floor on impact, never as the measurement.
Cadence. Sample often enough to see drift - weekly is usually enough, daily if the category moves - and judge on a longer horizon than you sample. The mismatch is deliberate: frequent measurement gives you the variance, infrequent judgement stops you from acting on it. A team that reviews a weekly citation rate every week will reliably invent explanations for noise, which is how measurement programmes lose the room.
What not to expect to detect
Some things are real and will not show up in your measurement, and pretending otherwise burns the programme's credibility.
- Fast propagation. Google says recrawling can take days to
months.[^google-recrawl-latency] A content change measured after a week is frequently measuring the old page.
- Small edits on a stable question. If you were cited 6 of 20 times before, a
rewrite that makes you slightly better will not clear the noise floor at that sample size. Batch the work and measure the batch.
- Off-domain corrections, immediately. Fixing a directory entry changes what
the model reads after that page is refetched and re-indexed, on the platform's schedule.
- Anything on a question your buyers do not ask. A tracked prompt nobody uses
is a metric with no consequences, and it will still consume your attention when it moves.
Numbers to refuse to report
A short list, because refusing is where credibility comes from.
- A percentage with no denominator. Always the first thing to go.
- A blended score across platforms. It hides which platform moved, which is
the only actionable part.
- An untracked platform rendered as zero. Unknown is not failure.
- Week-on-week movement inside the noise floor. Report the level, not the
wobble.
- A citation rate for a prompt set that changed this period. Changing the
questions changes the number; say so or do not report the comparison.
- Anything attributed to an in-house study that is really one screenshot. If the method
cannot be written in a sentence, the number is not ready.
The last one deserves the most vigilance because it is the most tempting. A single striking answer is far more persuasive in a meeting than a rate with an n, which is exactly why the discipline has to be a rule rather than a preference.
What to take from this chapter
Write the prompt set down. Sample repeatedly. Report rates with denominators, keep surfaces separate, and refuse to report movement you cannot distinguish from noise. The discipline is not expensive and it is what makes the rest of the guide falsifiable - without it, every recommendation in every chapter, including ours, is unverifiable advice.
Questions people ask
- Can Google Search Console show me AI Overview performance?
- Only as part of the whole. Google says sites appearing in AI features are included in overall search traffic in Search Console and reported within the Web search type - not broken out as their own dimension. It tells you nothing at all about ChatGPT, Claude or Perplexity.
- How many times should I run each prompt?
- Enough that a single unusual answer cannot move the number, and enough that you can state the count out loud without embarrassment. One run is an anecdote. The right question is not "what is the magic number" but "what is the smallest change I need to detect", and then whether your sample can detect it.
- Why do I get a different answer every time I ask the same question?
- Because generation is non-deterministic, and because the surfaces differ. Google states that AI Mode and AI Overviews may use different models and techniques so the links they show will vary, and that AI Overviews often do not trigger at all. Variation is the property being measured, not an error to remove.
- What is the difference between an unmeasured platform and a zero score?
- Everything. A platform you never queried has no visibility number; rendering that as 0% invents a failure that did not happen. Attensira publishes this rule for its own reporting - a null value means never measured, and a rate always carries the sample size behind it.
- Should I report a single AI visibility score?
- Only if you also report what went into it. A blended score across platforms, question types and mention-versus-citation hides the one thing you can act on - which question, on which platform, is losing. Report the components; derive the headline from them if someone insists on a headline.
Sources
Every factual statement above, with the page it came from and the date that page was read.
Google states that sites appearing in AI features are included in overall search traffic in Search Console, reported in the Performance report within the Web search type.
developers.google.com · retrieved
“they're reported on in the Performance report , within the "Web" search type”
Google states that AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary.
developers.google.com · retrieved
“AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary.”
Google states that AI Overviews are only shown when its systems determine they are additive to classic Search, and as such often do not trigger.
developers.google.com · retrieved
“AI Overviews are only shown when our systems determine that it is additive to classic Search, and as such, often don't trigger.”
Google states that clicks from search results pages with AI Overviews are higher quality, meaning users are more likely to spend more time on the site.
developers.google.com · retrieved
“when people click from search results pages with AI Overviews, these clicks are higher quality (meaning, users are more likely to spend more time on the site)”
Google states that crawling can take anywhere from several days to several months, depending on how often its systems determine a page needs to be refreshed.
developers.google.com · retrieved
“crawling can take anywhere from several days to several months, depending on how often our systems determine a page needs to be refreshed”
Attensira publishes that a rate it reports is a value plus the sample size behind it, and that a null value means never measured rather than measured zero.
attensira.com · retrieved
“A rate is `{value, n}` - `value: null` means never measured, and `n` is the sample size behind it.”
Attensira publishes that per-platform metrics carry whether the platform was tracked at all, and that reporting an untracked platform as 0% invents a failure.
attensira.com · retrieved
“`tracked: false` means the plan never queried that platform, so it has no visibility number at all. Reporting it as 0% invents a failure.”
Attensira publishes that a reported delta carries whether the movement is real, and that movement is only reported when it clears the noise floor.
attensira.com · retrieved
“A delta is `{value, real}` - movement is only reported when it clears the noise floor.”
Attensira publishes that answers are separated by surface - the consumer app versus the model's API - and that the two are never averaged together because they do not answer alike.
attensira.com · retrieved
“Answers are also separated by surface - the consumer app versus the model's API - and the two are never averaged together, because they do not answer alike.”
The GEO study's position-adjusted word count metric weights a source's contribution to an answer by an exponentially decaying function of the citation's position.
arxiv.org · retrieved
“we propose a position-adjusted count that reduces the weight by an exponentially decaying function of the citation position”