For the complete documentation index, see llms.txt. Every page on this site is also served as Markdown: append `.md` to any URL, or send `Accept: text/markdown`.

How to Audit Mentions Across ChatGPT, Claude, and Gemini

Karl-Gustav KallasmaaKarl-Gustav Kallasmaa, Founder & CEOLast updated
How to Audit Mentions Across ChatGPT, Claude, and Gemini

A practical AI search audit playbook: prompt sets, logging columns, competitor comparison, and what missing looks like across ChatGPT, Claude, and Gemini.

Audit mentions across ChatGPT, Claude, and Gemini
Audit mentions across ChatGPT, Claude, and Gemini
📊
An AI search mention audit is not a vanity dashboard. It is a prompt-by-prompt baseline that shows where ChatGPT, Claude, and Gemini name you, cite you, or skip you — so you can fix the underlying gap, not just watch a score move.

When a buyer asks ChatGPT for the best tool in your category, the answer usually names one to three brands. If you are not one of them, there is no SERP position to defend and no Search Console row that explains the miss. You simply were not in the answer.

This playbook is how to audit that gap across ChatGPT, Claude, and Gemini in a way you can re-run. It covers prompt sets, a logging sheet, competitor comparison, what "missing" actually looks like, and when to stop doing it by hand. For ongoing monitoring after the baseline, see AI search tracking.

Why one platform audit is not enough

The three assistants do not share a ranking system, and they do not pull from the same evidence.

Cross-platform disagreement is normal, not a measurement bug. Analyses of multi-million citation sets show models weight owned sites vs directories differently, and brand mentions disagree across assistants a large share of the time — often cited around 62%.[1] Auditing ChatGPT alone tells you almost nothing about Claude or Gemini.

What you are measuring (keep these separate)

Collapse everything into one "AI visibility score" and you lose the diagnosis.

  1. Mention — your brand name appears anywhere in the answer (recalled or listed).
  2. Citation — the assistant links to your URL as a source. Stronger trust signal than a bare name-drop.[2]
  3. Position / framing — first recommendation, mid-list peer, or "alternative if you need X."
  4. Accuracy — features, pricing, ICP, and integrations are true, outdated, or fabricated.
  5. Share of voice — your mention rate on the same prompt set vs named competitors.

A useful directional benchmark used in 2026 audit write-ups: aim for roughly a 20% mention rate on category-defining recommendation prompts before you call the footprint healthy.[3] Treat that as a starting bar, not a law of nature — your category concentration and competitor set matter more than the round number.

Step 1 — Build a buyer-language prompt set

Build 25–40 prompts (enough for a first baseline; expand to 50+ once you automate). Write them the way a stranger would ask, not the way your product marketing deck reads.

Four intent buckets

  • Category / recommendation — "Best [category] for [use case] in 2026"
  • Comparison — "[You] vs [Competitor]", "Alternatives to [Competitor]"
  • Problem / job-to-be-done — the pain with no brand named
  • Branded / accuracy — "What does [Brand] do?", "[Brand] pricing", "Is [Brand] good for [ICP]?"

Prompt set template (copy into your sheet)

Weight the set toward non-branded commercial and problem prompts. Asking "What does [your brand] do?" only proves the model can recall you when prompted by name. Discovery prompts are where revenue is won or lost.

Step 2 — Run the same prompts on ChatGPT, Claude, and Gemini

For each prompt:

  1. Use a fresh / temporary chat (logged out or history cleared). Personalization quietly biases results toward brands you have already discussed.[1]
  2. Run the exact same wording on all three platforms.
  3. Capture the full answer (screenshot or paste), not a yes/no.
  4. Where the product allows search / browsing, note whether it was on. Live retrieval changes citation behavior.
  5. Re-run high-priority prompts 2–3 times. Answers regenerate; treat single runs as directional.

Cadence for a first audit: one focused half-day to full day. For re-checks, ChatGPT (especially with browsing) moves faster week to week; Gemini often tracks closer to Google's update rhythm; Claude's training-data answers usually shift more slowly — monthly is often enough for Claude alone.[1]

Step 3 — Log every run in a sheet (columns that matter)

Use one row per prompt × platform × run. Keep mention and citation as separate columns.

Roll-up metrics (calculate after the run)

  • Mention rate = prompts with mentioned=yes ÷ prompts tested (per platform, then blended).
  • Citation rate = prompts with a non-empty cited_url ÷ prompts tested.
  • SoV = your mentions ÷ (your mentions + competitor mentions) on the same prompt set.
  • Accuracy fail rate = branded prompts marked outdated or fabricated ÷ branded prompts.

Step 4 — What "missing" looks like (taxonomy)

"We're not showing up" is four different problems. Tag each miss.

Off-site presence still dominates many citation patterns. Multiple 2025 PR-industry analyses summarized in audit guides put a large share of AI citations on earned and owned media rather than paid placements — often paraphrased as ~90%.[1] If your only plan is "publish more on our blog," expect Absents to persist.

Step 5 — Competitor comparison that produces a work queue

Pick 3 competitors you actually lose deals to. Run the identical prompt set. For each P0 prompt, record who was named first and which domains were cited.

Then answer three questions in writing:

  1. Where do they win discovery prompts we lose? (list prompt IDs)
  2. Which recurring sources name them and skip us? (Reddit threads, G2 categories, "best of" posts, docs sites)
  3. What is the smallest fix that would change the next re-run? (one page rewrite, one listicle pitch, one schema/entity cleanup, one review-profile completion)

This is the difference between an audit and a report. The output should be a prioritized backlog, not a chart.

If you need a parallel view of how AI Overviews pick sources on Google itself, pair this audit with how to rank in AI Overviews.

Step 6 — Turn findings into fixes (then re-measure)

Map each miss type to an owner and a 30-day action:

  • Absent on recommendation prompts — ship a clear, answer-first category page; pitch 2–3 third-party listicles that already rank in your niche; close review-profile gaps on G2/Capterra where relevant.
  • Named but never cited — add direct-answer openings, FAQ blocks, and clean Organization/Product schema; remove JS-only walls on the pages you want quoted.
  • Wrong story — update pricing/feature pages; request corrections on the stale third-party URL the model is quoting.
  • Gemini-only gap — validate Google entity consistency and schema; check crawlability.
  • Claude-only gap — prioritize durable editorial and community mentions over a one-week blog blitz.

Re-run the same prompt IDs in 30–60 days. Visibility work often takes that long to show up; weekly Claude checks mostly add noise.[2]

When to automate (and what automation still cannot do)

Stay manual when:

  • You are running a first baseline (under 40 prompts).
  • You need qualitative reads on framing and accuracy.
  • You have not yet validated which prompts matter to pipeline.

Automate when:

  • You re-test 50+ prompts across three+ platforms on a weekly cadence.
  • You need trend lines, not anecdotes.
  • Multiple stakeholders need the same prompt library without copy-paste chaos.

Most AI visibility tools tell you whether you were mentioned. They rarely tell you why a competitor was chosen or ship the fix. That diagnosis — source mix, missing corroboration, unquotable pages — stays human unless your stack is built to do the work.

Attensira is built for that second half: find why a brand is missing from AI search, write the fix, and open a PR or CMS draft for approval — not another dashboard to babysit. If you want a quick technical baseline before the prompt audit, run the free AI Readiness Score.

One-week audit checklist

Freeze a 25–40 prompt library across the four intent buckets
Name 3 competitors and lock exact prompt wording
Run every prompt on ChatGPT, Claude, and Gemini in fresh sessions
Log mention, citation, position, sources, accuracy, and missing type
Calculate mention rate, citation rate, and SoV per platform
Extract the top 5 recurring cited domains you are absent from
Write a 30-day fix list with owners (on-site vs off-site)
Schedule the re-run on the same prompt IDs

FAQ

How is this different from SEO rank tracking?

Rank tracking measures position in a list of links. An AI search mention audit measures whether you are chosen inside a generated answer that typically names only a few brands. There is no reliable "#1" — only mention rate, citation rate, and share of voice across a prompt set.[3]

Do I need Perplexity and AI Overviews in the first audit?

If your buyers live there, yes. This playbook focuses on ChatGPT, Claude, and Gemini because they cover a large share of assistant usage for B2B research. Add Perplexity and Google AI Overviews in week two once the sheet and prompt library are stable.

What is a good first score?

Use your competitors as the bar. A 20% mention rate on category prompts is a common directional benchmark in 2026 audit guides,[1] but beating the rival who currently owns your P0 prompts matters more than hitting a round number.

How often should we re-audit?

Full prompt-set re-audit monthly for most B2B teams; weekly spot-checks on the 10–20 prompts closest to pipeline. Match cadence to platform volatility — faster for ChatGPT with browsing, slower for Claude.


Next step: run the checklist on your top 10 commercial prompts this week. If you want the missing-why diagnosis turned into shipped page fixes, start with Attensira.

Related Articles

See where you rank in AI search

Find out how ChatGPT, Claude, and Google AI recommend your brand — and what to fix to rank higher.

Setup in under 2 minutes.