The GEO measurement guide · chapter 5
What a defensible AI visibility report looks like
What has to be in an AI visibility report before I should believe the numbers in it?
Karl-Gustav Kallasmaa, Founder & CEOLast updated The artefact this chapter describes
A defensible AI visibility report is one that survives a hostile reader. Not a sceptical one — a hostile one: someone who does not want the conclusion to be true and who will look for the weakest number on the page.
That reader exists in every company. They are usually the person who has to fund the work. A report built for them is also, conveniently, the report that an AI system will cite and a journalist will quote, because the properties are the same: stated method, visible denominators, and claims that do not exceed the evidence.
The eight fields, and why each is load-bearing
Every number in the report carries these, or the number does not go in.
1. The rate, as a pair with its sample size. Not a bare percentage. Rates should travel as a value and a denominator together, structurally, so that a component cannot render one without the other.[^attensira-value-n] The reason is not pedantry: a rate of one third from three observations and a rate of one third from three hundred look identical on a chart and deserve entirely different responses. A jump from zero to a third at a sample size of three is one answer changing its mind, not a trend.[^attensira-small-n]
2. The window, with explicit boundaries. Which dates, and whether the last one is closed. Partial days are incomplete in a way that is not uniform across cells,[^attensira-partial-day] so a chart whose final point is today has one point that cannot be compared to the others.
3. The prompt set, versioned. The set is the sampling frame. Any comparison that spans a change to it is comparing two different measurements. Name the version, date it, and say whether it changed during the period.
4. The models, countries and surfaces. Each is part of the cell definition, not a filter. Consumer-app answers and API answers should never be averaged together, because they do not answer alike.[^attensira-surface-split]
5. The matching rule. What counts as a mention: exact string, case handling, possessives, plurals, partial matches. Two vendors' mention rates are not comparable until this is stated, and it is the single most commonly omitted field in the category.
6. Uncertainty. At minimum the denominators; better, an interval. The Wilson score interval is the standard recommendation for proportions and behaves sanely at the small sample sizes where this matters most.[^nist-wilson]
7. A significance test on every movement. A change is drawn as a change only after it clears a stated test — a two-proportion z-test at 95% confidence between the two windows is a reasonable default — and otherwise the delta is reported as not real rather than as a number.[^attensira-ztest] Expect "nothing we can prove" to be a frequent finding at shallow sampling depths, and expect that to be uncomfortable rather than wrong.
8. Coverage gaps, named. Which cells were not measured, and why. An unqueried model is marked as untracked and no zero applies to it.[^attensira-tracked-false] The gaps belong in the report body, not in a footnote, because they bound every conclusion above them.
The claims a report must not make
"Our AI visibility is X." Incomplete without the prompt set. The honest form names it: our mention rate across these questions, on these models, in this window, was X on a denominator of N.
"We rank Nth in ChatGPT." There is no rank. Ordering within a single answer's source list is not a ranking against competitors, and a table sorted by it will be read as a league table regardless of the caveat underneath.
"AI search drove X in revenue." No tool in this category can connect a mention to a visit, a signup or revenue.[^attensira-no-attribution] Google's own reporting folds AI feature appearances into overall search traffic rather than separating them,[^gsc-combined] so even the platform closest to the data does not expose the link. Report visibility as a leading indicator held next to funnel data, and label the relationship correlation.
"We hold X share of AI answers." Mention rates do not sum to one across brands, because one answer can name five brands or none, and treating a rate as a slice of a pie leads to wrong conclusions.[^attensira-mention-not-pie] If you report share of voice, publish the brand set in the same view.
"Visibility fell this week." Only if it cleared the test. Otherwise the true statement is that nothing was proven, which is different from nothing having happened and different again from something having happened.
How these numbers are used to mislead
This section is written about the category, including us. Every technique below is available to any vendor in this space, and several are available to us specifically — which is why our own product is designed to make them awkward. No company is named, because no specific vendor was audited in the research behind this page, and asserting that a named competitor does these things would be exactly the kind of unsourced claim this guide argues against.
Zeroes manufactured from nulls. The most consequential one. A plan that does not query a surface produces no data for it; rendered as zero, it becomes an apparent visibility failure. The customer sees a red bar, buys an upgrade, and the bar improves — because the surface is now being queried, not because anything changed. The counter is structural: make the rate type nullable and make "not measured" a distinct rendering, everywhere. The instruction in our own documentation is blunt about this for exactly this reason.[^attensira-null-vs-zero]
Denominator shopping. Filter until the number is good. Every filter divides the sample, so any dataset contains a favourable cell if you slice finely enough, and it will be technically accurate. The counter is showing the denominator at every level of filtering, so that a thin cell announces itself.
The flattering prompt set. Choose the questions where the answer is already you. This does not require dishonesty — it is what happens by default when the person picking the prompts is the person accountable for the number. The counter is sourcing questions externally and requiring the set to contain expected losses.
Comparison-set curation. Share of voice moves when you edit the list of competitors. A narrow set produces a commanding figure. The counter is publishing the brand set beside the number, always.
Noise sold as trend. Present every week-over-week movement as a finding. Without a significance gate, roughly half the arrows will point up, and a quarterly report can be assembled entirely from the favourable half. The counter is a stated test and the willingness to publish an empty chart.
The unfalsifiable composite. A visibility score built from unpublished weights cannot be recomputed, cannot be audited, and can be adjusted. When it moves, only the vendor knows why. The counter is fewer numbers with published definitions.
Re-baselining. Change the window length, the start date, or the model mix, and the trend line changes shape without any underlying movement. The counter is annotating instrument changes on the timeline as events.
Surface blending. Average consumer-app and API answers into one rate. The result describes a system that does not exist, and it is convenient because the two surfaces rarely disagree in the same direction. The counter is keeping them apart end to end.
Attribution by adjacency. Put a visibility chart next to a revenue chart on one slide and say nothing. No false statement is made; the reader supplies the causal claim. The counter is a written label, on the slide, saying these are held together as correlation.
A reporting template
A structure that carries all eight fields without becoming an appendix nobody reads.
Summary. Two or three sentences. The aggregate mention rate and citation rate with denominators, whether either moved by the stated test, and the single largest coverage gap. If nothing cleared the test, say so in the first sentence.
Movement. Only changes that passed. Each with both window values, both denominators, and the test used. If the section is empty this period, leave it empty and label it.
Diagnosis. The mention-versus-citation gap by topic, and the pages the assistants cited instead of yours. This is the section that produces work.
Coverage. Which models, countries and surfaces were measured; which were not and why. Written as cells, not as prose.
Method. Prompt set version and link, sampling depth, window boundaries, matching rule, failed-run treatment, significance test. Short enough that people read it, complete enough that a reader can reproduce the denominator.
Known limits. What this report cannot tell you, stated in your own words rather than boilerplate. Attribution belongs here permanently.
Evaluating a vendor with three questions
The same discipline works as a purchasing test, and it takes about five minutes.
What is the sample size behind this number, and can I see it in the interface? If the denominator is not in the product, it is not in the vendor's thinking either.
Show me a model you do not track for my plan. What does the cell render? A zero here is disqualifying, because the same carelessness will exist in places you cannot inspect.
What test does a change have to pass before you draw it as a change? A vendor without an answer is selling you noise with arrows on it, and the arrows will point wherever the last few runs happened to land.
A vendor who answers all three in writing has thought about measurement as a discipline rather than as a rendering problem. That is a small population, and it is a better filter than any feature matrix.
Back to the guide
The guide's pillar page collects the five rules these chapters build on, and the chapter on what can and cannot be measured is the one to re-read whenever someone asks for a number that does not exist.
Questions people ask
- What is the shortest test of whether an AI visibility report is trustworthy?
- Ask three questions. What is the sample size behind this number? What does an unmeasured cell render as? What test does a change have to pass before it is drawn as a change? A report or a vendor that can answer all three in writing has done the work. One that answers with a composite score has not.
- What should a report never claim?
- That a mention caused a visit, a signup or revenue; that a rate is a ranking; that an unqueried surface is a zero; that a movement is a trend without a significance test; and that a share-of-voice figure means anything without its brand set published. Each of those is a claim the underlying data cannot support.
- Can a vendor mislead without lying?
- Easily, and this is the more common case. Choosing a flattering prompt set, selecting a comparison group, filtering to a cell with a favourable rate, re-baselining a chart, or blending measured and unmeasured cells all produce technically accurate numbers that support a false impression. None of them require a false statement.
- How should uncertainty be shown to a non-technical audience?
- Show the sample size next to every rate, and use a band rather than a line where a band is warranted. If that is too much, the fallback that costs nothing is to simply not draw movements that failed the significance test — an absent arrow communicates uncertainty better than a small-print caveat under a big one.
- What belongs in the methodology appendix?
- The prompt set and its version, the models, countries and surfaces, the sampling depth, the window boundaries, the matching rule for what counts as a mention, the treatment of failed runs, the significance test, and every known coverage gap. If a reader cannot reproduce your denominator from the appendix, the appendix is incomplete.
- Should I put AI visibility numbers in a board deck?
- Yes, with the aggregate rate over a long window, the sample size, and an explicit note that this is a leading indicator with no attribution path to revenue. What does not belong in a board deck is a filtered cell, a composite score, or a week-over-week movement that did not clear a test.
Sources
Every factual statement above, with the page it came from and the date that page was read.
Attensira's documentation instructs consumers never to render an unmeasured cell as zero, because reporting null as zero invents a failure that was never observed.
docs.attensira.com · retrieved
“Never render an unmeasured cell as 0%. Reporting `null` as zero invents a failure that was never observed.”
Attensira returns every rate as a pair of a value and a sample size, and returns a delta as a pair of a value and a flag stating whether the movement is real.
docs.attensira.com · retrieved
Attensira reports a change between windows only when it clears a two-proportion z-test at 95% confidence, and otherwise returns the delta as not real.
docs.attensira.com · retrieved
Attensira's documentation warns that a jump from 0 to 0.33 at a sample size of 3 is one answer changing its mind rather than a trend.
docs.attensira.com · retrieved
“a jump from 0 to 0.33 at `n` of 3 is one answer changing its mind, not a trend”
Attensira's limitations page states the product cannot tell you that a mention produced a visit, that a citation produced a signup, or that any of it produced revenue.
docs.attensira.com · retrieved
“Attensira cannot tell you that a mention produced a visit, that a citation produced a signup, or that any of it produced revenue.”
Attensira marks a model that was never configured for tracking as tracked false, meaning it was never queried and no zero applies to it.
docs.attensira.com · retrieved
Attensira separates answers by surface — the consumer app versus the model's API — and states the two are never averaged together because they do not answer alike.
attensira.com · retrieved
“the two are never averaged together, because they do not answer alike”
Attensira documents that the current day's numbers are incomplete until the day closes, and that the shape of that incompleteness is not uniform.
docs.attensira.com · retrieved
“Today's numbers are therefore incomplete until the day closes, and the shape of that incompleteness is not uniform.”
Attensira documents that mention rate does not sum to 100% across brands and warns that treating it as a slice of a pie will lead to wrong conclusions.
docs.attensira.com · retrieved
“treating it as a slice of a pie will lead you to wrong conclusions”
Google documents that sites appearing in AI features such as AI Overviews and AI Mode are included in overall Search Console search traffic under the Web search type.
developers.google.com · retrieved
The NIST/SEMATECH e-Handbook gives the Wilson score interval for a proportion and notes Agresti and Coull recommended it for virtually all combinations of n and p.
itl.nist.gov · retrieved