Guide
How to measure AI visibility without fooling yourself
Measure AI visibility with a frozen question set: pick a fixed list of the prompts your buyers actually type, run them on a schedule, and track per-provider whether you are mentioned or cited. Do not trust a single one-off score - AI answers move day to day on content that never changed, so you have to know your noise band before you can read real movement.
If you have spent any time in SEO communities lately, you have seen the complaint: "I used three AI visibility tools, gave them the same prompts, and got three different answers." It is a fair criticism. A lot of AI-visibility dashboards look precise and are actually guessing.
But "the tools disagree" does not mean AI visibility is unmeasurable. It means most people are measuring it wrong - a single snapshot, one provider, no baseline. Here is how to do it properly, and how to tell an honest measurement from a flashy number.
Why one-off AI visibility scores mislead you
Ask ChatGPT "who are the best plumbers in Leeds?" twice in the same week and you can get two different lists - same model, same prompt, nothing on your website changed. AI answers vary by:
- Provider - ChatGPT, Perplexity, Gemini, and Google AI Overviews each cite different sources. Being named by one tells you little about the others.
- Prompt phrasing - "best X" and "what is the best X" can return completely different brands.
- Whether the model searched - sometimes it retrieves live sources, sometimes it answers purely from training data. Those are two different games (more on that below).
- Plain randomness - the same query genuinely drifts day to day.
A tool that hands you "Your AI visibility: 63/100" from one run, one provider, one prompt is selling you certainty it does not have. That is exactly the thing sceptics are right to distrust.
The honest method: a frozen question set
The most credible way to measure AI visibility is the one serious practitioners actually use on their own sites. It is not complicated:
Step 1
Freeze a set of buyer-intent questions
Write down 10-20 prompts a real customer would type when they are ready to buy - "best [your service] in [your city]", "who should I use for [problem]", "[your category] recommendations". Fix the list. Do not change it. The whole point is that the questions stay constant so any change in the answers comes from the world, not from you moving the goalposts.
Step 2
Run them on a schedule, per provider
Run the same set across ChatGPT, Perplexity, Gemini, and Google AI Overviews on a regular cadence - daily or weekly. Record three things for each: were you mentioned by name, was your domain cited as a source, and were you recommended (versus just listed).
Step 3
Learn your noise band before you celebrate
Run it enough times to see how much the score naturally bounces when nothing changed. If your "mentioned" rate wobbles by a few points run to run, that is noise. A real improvement is a move clearly bigger than that wobble. Without a baseline you will mistake random drift for progress - or panic over a dip that means nothing.
Step 4
Track each provider separately
There is no single AI-visibility number, because citation is provider-specific. It is normal to be recommended heavily by one engine and invisible in another on the exact same prompts. Report per provider. A blended "overall AI score" hides the only insight that lets you act.
This is unglamorous and that is the point. A frozen question set run often enough to know your noise band is a real measurement. A single flashy dashboard number is not.
Retrieval vs training data: two different things you are measuring
One reason measurement feels slippery is that AI answers come from two different places, and they respond to different work:
- Retrieval - the model searches the live web and cites sources. This is mostly classic SEO territory: crawlable pages, clear content, third-party corroboration. If you rank and get cited elsewhere, you have a shot at being retrieved.
- Training data / model knowledge - what the model already "knows" without searching. If it has never encountered your brand, it will not go looking for you, and no amount of on-page tweaking fixes that. This is an entity-presence and reputation problem, built over time through mentions, reviews, and consensus across sources.
You can check the training-data side directly: ask a model about your brand and your competitors with web browsing turned off. The answer will be stable for months, until the model's knowledge cutoff moves. That stability is a feature - it makes it measurable.
What a score is good for (and what it isn't)
A page-level AI visibility score - like the one Visus produces - is not a prediction that AI will recommend you. Nobody can promise that, and you should be wary of any tool or agency that does. What a score is good for:
- Diagnosis - it shows which readiness signals you are missing (AI crawler blocks, no FAQ, hidden pricing, weak entity clarity, thin corroboration), because those are the fixable things that consistently separate cited pages from ignored ones.
- A before/after baseline - re-run it after you ship fixes and watch the specific gaps close.
- A precise brief - if you do hire help, you hand them a prioritised list instead of paying a retainer to rediscover it.
Pair a deterministic readiness score (does my page have what AI rewards?) with a frozen-question-set monitor (am I actually being mentioned, per provider, over time?) and you have an honest picture. One without the other is half the story.
We put this to the test at scale: in a benchmark of 100 businesses, the gaps that predicted low visibility were mundane and fixable - 79% had no answer-ready FAQ content, 70% had reviews that were not visible as on-page text, and national brands scored lower than local ones. None of that requires a crystal ball to measure.
Get an honest baseline in under 60 seconds
Free audit. A deterministic score out of 100, your top fixes in plain English, no guarantees and no credit card.
Run free AI visibility auditFind GetVisus useful? Make us a preferred source in Google: