How to measure AI visibility (without fooling yourself)
updated june 16, 2026
Most AI visibility numbers are easy to game and hard to trust. This guide walks through a method you can run yourself: pick stable buyer questions, read the answers consistently, and score what actually predicts a buying decision instead of a flattering mention count.
To measure AI visibility, build a fixed set of buyer-style prompts, run them across the assistants your market uses on a regular cadence, and record four things for each answer: whether your brand appears, which competitors appear with it, what claim is attached to your name, and whether the cited sources support that claim. Track those as answer drift over time. The discipline is keeping the questions stable so movement means something — a rising mention count on generic prompts is the easiest way to fool yourself.
The method, step by step
- 01
Build a stable question set
Write the questions a buyer actually asks before choosing, across category, comparison, evaluation, and objection language. Lock the wording so reads stay comparable.
- 02
Read across assistants
Run the set through the assistants your buyers use — ChatGPT, Claude, Perplexity, Google's AI answers — and capture the full answer, not just whether your name appears.
- 03
Score presence, accuracy, and support
For each answer, record presence, the competitors named, the claim attached to you, and whether the cited sources back it up. Accuracy and support matter as much as presence.
- 04
Track movement, not snapshots
Repeat on a regular cadence and compare against the prior read. The change between reads is the signal; a single day is just a snapshot.
- 05
Separate signal from noise
Citation churn and minor wording shifts are usually noise. A new competitor, a hardening claim, or a question where you newly disappear is signal worth acting on.
The measurement sheet
A useful measurement sheet is boring on purpose. Each row should make the next read comparable to the last one, and every column should explain whether the answer helps a buyer choose.
- 01
Question family
Category, alternative, comparison, implementation, objection, or pricing. The family matters because a brand can win one part of the journey and disappear in another.
- 02
Locked prompt
The exact buyer question, copied without improving it between reads. Changing the wording resets the measurement.
- 03
Surface and date
The assistant or answer surface, the mode if it matters, and the date you ran it.
- 04
Recommendation rung
Absent, mentioned, described accurately, recommended, or recommended with support.
- 05
Claim and source support
The wording attached to your brand, the visible source trail, and whether that evidence backs the claim.
- 06
Work implied
The page, proof, comparison, or source work this row would create if a buyer judged you from this answer.
If a column does not help you decide what changed or what to do, leave it out. Measurement gets weaker when it becomes reporting theater.
Who this guide is for
If you have ever screenshotted one good ChatGPT answer and thought "we're doing great," this is for you. It is written for founders, marketers, and growth leads who want a read on AI visibility they can defend in a meeting — not a flattering moment they happened to catch. You do not need a data team. You need a stable set of questions and the discipline to read them the same way every time.
What you'll need before you start
Before you run a single prompt, get three things in place.
- Your domain and category, in the words a buyer would actually use.
- A short list of the competitors buyers already weigh you against.
- 15 to 30 buyer questions, spread across category education, alternatives, comparisons, pricing, and implementation.
Lock the wording before you start. Change the prompt and you change the answer — and you have lost the ability to compare one read to the next.
The mistakes that inflate the number
Three habits quietly turn a weak position into a flattering chart. A measurement worth trusting survives all three.
- Testing only friendly, branded prompts — the ones you already know you win.
- Counting any mention as a win, even when the claim attached to your name is wrong.
- Re-reading on different prompts each week, so you can never tell a real shift from random noise.
Use prompt families, not random prompts
A good question set is built from families, not a pile of whatever came to mind. The AI visibility playbook is useful here because each family covers a different moment in the buyer's decision:
- Category — "best tools for…"
- Alternatives — "products like X," "alternatives to X"
- Comparison — "X vs Y," "is X better than Y?"
- Implementation — "how hard is X to set up?"
- Objection — the late questions about price, security, or fit
Miss a family and your score quietly overstates visibility in the part of the journey you happened to test.
Score recommendation strength, not mentions
Counting mentions hides the thing that matters: a mention is not a recommendation. Score every answer on a simple ladder instead.
- Absent — your brand is not in the answer at all.
- Mentioned — named in passing, with no real endorsement.
- Described accurately — named, and the claim about you is right.
- Recommended — the answer actually points the buyer toward you.
- Recommended with support — and it cites a reason or source for why.
A brand mentioned ten times as an afterthought is weaker than one recommended three times with a reason and a source. Score the rung, not the mention.
Turn the result into a work list
The output is a brief, not a grade. Read it by where you fell short — each kind of gap points at a different fix:
- Absent on a prompt → a page or a source is missing.
- Misdescribed → a positioning or proof gap.
- A competitor wins with a specific reason → comparison or use-case work.
Each line is a brief, not a verdict — and that is the difference between a number you watch and a number you can actually move.
A baseline row in practice
Imagine a workforce software brand checking whether it belongs in restaurant scheduling recommendations.
Prompt
"Best scheduling software for restaurants with changing shifts" under the comparison family.
Answer
Claude names the brand in a longer list, but recommends a rival first because the rival has clearer restaurant-specific proof.
Score
Described accurately, not recommended. The source trail supports the rival's claim better than yours.
Work
Create or sharpen the restaurant scheduling page, add workflow detail, and re-read the same prompt after the page has had time to be picked up.
The value is not the row itself. The value is that the same row can be read again next week without moving the goalposts.
What a real measurement captures
- A fixed buyer-question set
- Presence and accuracy per answer
- Competitor co-mentions
- Source support behind claims
- Movement between reads
Measuring AI visibility, answered
- How many questions do I need?
- Enough to cover the buying journey without becoming unmanageable — usually 15 to 30. Breadth across question types matters more than raw volume, because a brand can look visible on generic prompts and vanish on specific ones.
- How often should I measure?
- Often enough to catch drift, rarely enough to act between reads. A daily read works when something else does the watching for you; a weekly manual pass is a reasonable floor if you run it yourself. If the wording itself keeps changing, treat that as answer drift, not just a measurement footnote.
- Can I just count mentions?
- A raw count tells you your name appeared, not whether the answer helped. Pair presence with the claim attached to you and the sources behind it, or you will optimize a number that does not move buyers.