Free audit
Measurement · Analysis

The denominator problem in AI visibility tools

By Robert Langford · Founder · 16 years in search

AI visibility measurement produces numbers that look precise and compare to nothing. The reason is almost always the denominator: 17 citations out of 42 runs and 17 out of 100 both get reported as 17, and only one of them is good news.

How AI visibility measurement goes wrong

Language models do not always retrieve. In one measured configuration, 57.8% of repeated ChatGPT runs never activated web search at all — the model answered from what it already held. Those runs contain no citations, which makes them awkward for any tool trying to report a brand's share of citations.

Why are AI visibility scores inflated?

The same 17 citations. Reported against runs where the model actually searched, it reads as 40%. Against every run, 17%. Both get published as a visibility score.Seventeen filled cells out of forty-two, next to seventeen out of one hundred.17/42Runs that searched17/100All runs
The same 17 citations. Reported against runs where the model actually searched, it reads as 40%. Against every run, 17%. Both get published as a visibility score.Our own measurement, 2026

What AI visibility tools do with the awkward runs

The convenient solution is to drop them. A run with no citations looks like noise, it breaks the chart, and excluding it produces a cleaner number. So a large share of AI visibility tooling reports your share of the responses that cited something, rather than your share of all responses.

That is a defensible engineering decision and an indefensible measurement one, because the excluded runs are not noise. They are the majority outcome in some configurations, and they represent a real state of the world: the model was asked, and it answered without going to look.

The arithmetic, made concrete

Suppose a hundred runs of a prompt set. Forty-two of them retrieve and cite; fifty-eight do not. Your brand appears in seventeen of the forty-two that cited.

  • /Reported as a share of citing runs: 17 of 42, or 40.5%. That is the number most dashboards will show you.
  • /Reported as a share of all runs: 17 of 100, or 17%. That is how often a person asking that question actually sees your name.
  • /The difference is not a rounding error. It is a factor of nearly two and a half, and it always points the same way.

Neither figure is fabricated. But only one of them answers the question an operator is actually asking, which is how often a real person putting that question to a real model comes away having heard of you.

A measurement error that never once disadvantages the client is not an error. It is a feature nobody wrote down.

Why this persists

Partly because it is genuinely easier. Partly because there is no agreed standard, so nobody can be accused of deviating from one. And partly for the reason most measurement inflation persists in marketing: the party being measured is also the party paying, and a number that is flattering is a number that renews.

It is worth saying that most vendors are not being cynical about this. Retrieval behaviour varies by model, by prompt phrasing, by account state and by week, and building a denominator you can defend is considerably harder than building one that looks tidy. But the incentive still only points one way.

The second problem: single runs

Denominator choice is the larger error, but it sits next to a smaller one that compounds it. Model output is not deterministic. Ask the same question three times and you can get three different sets of cited sources, which means a single run tells you very little about anything.

A great deal of confident advice in this field is built on single runs. Somebody asks ChatGPT which casino is best in their market, sees a competitor named, and concludes something about visibility. Run it five more times and the answer frequently changes, sometimes including them and sometimes not.

The practical implication is that any measurement claiming precision to a decimal point is either running far more observations than it is admitting to, or presenting variance as signal. We run each prompt three times across 6 AI surfaces and still report ranges rather than points.

A note on what the numbers are worth

None of this makes AI visibility the largest thing on an operator's dashboard. Similarweb counted 1.13 billion referrals from AI platforms across June 2025, a rise of 357% on the year, and still close to nothing in absolute terms.6% of the 191 billion Google sent in the same month.

The argument for measuring it carefully is not volume. It is that citation presence produces roughly a 35% click lift on the same query, and that the visitor arriving from an AI answer has already been told you are worth considering. That is a different visitor, and it is worth counting honestly, not generously.

Report three figures, not one

The fix is not a better single number. It is refusing to produce one. Retrieval rate tells you how often the model went looking. Citation rate tells you how often it cited anything at all. Share tells you how often it named you. Collapse those into one percentage and you lose the ability to explain any movement in it.

This matters practically. A brand whose retrieval rate doubles while its share stays flat has made real progress — the model is now searching on those queries, which is the precondition for being cited. A blended score would show that as nothing happening, and the client would reasonably conclude the work was not working.

Averaging 6 AI surfaces destroys the same information

There is a related error worth naming while we are here. Tools that blend several AI surfaces into one composite score lose the only detail that would tell you what to do next, because the surfaces genuinely disagree with each other.

An analysis of roughly 366,000 citations found that citations concentrate on a small pool of outlets and that different engines pick different ones, with low agreement between platforms. A strong Perplexity figure tells you very little about ChatGPT, and a composite tells you nothing about either.

Scale differs too. AI Mode accounted for about 0.34% of Google searches between January and April 2026 and refers traffic out at a materially lower rate than classic search. Weighting it equally with ChatGPT in a blended score is a defensible engineering choice and a misleading commercial one.

How to test the tool you are already using

  • /Ask directly what the denominator is. Whether runs with zero citations are included. A vendor who cannot answer that quickly has not thought about it, which is itself the answer.
  • /Ask for the raw run data, not the summary. If every row has at least one citation, the non-retrieving runs have been filtered out somewhere upstream.
  • /Run the same prompt yourself, five or six times, in a clean unauthenticated session. Count how often the model searched. If your count is nowhere near the tool's implied retrieval rate, that gap is the story.
  • /Compare two vendors on the same prompt set. Wildly different scores from the same questions almost always trace back to denominator choices and not to genuinely different findings.

What we do, since it is only fair to say

We include every run, including the ones where the model never searched, and we report retrieval, citation and share separately. Our numbers come out lower than several dashboards on the market and clients occasionally ask why, which is a conversation worth having once instead of a chart worth trusting indefinitely.

We also label movements that sit inside the variance band as exactly that and not presenting them as progress. It makes for a duller monthly report and a more useful one.

A closing note on why this matters commercially instead of academically. In iGaming SEO in full the AI surfaces are already deciding which casino, sportsbook or affiliate brand gets named when a player asks for a recommendation. An agency reporting an inflated share to a gambling client is not making a rounding error, it is describing a competitive position that does not exist — and the client budgets against it.

What this piece does not claim

It does not claim every vendor is being cynical. Retrieval behaviour genuinely varies by model, prompt, account state and week, and building a defensible denominator is harder than building a tidy one. It also does not name vendors, because the point is the measurement choice and not any particular product.

Sources for the denominator figures here: ICODA. Every figure links out, which matters here because the denominator is the whole argument.

How we handle the denominator in practice is set out in the published methodology, which gives the prompt bands, the three-run sampling and the rule that refusals are reported rather than counted as zeros.

Sources for the figures here: the third-party citation share is Seer Interactive, 2026; the AI Overview overlap range is Semrush and BrightEdge, early 2026; and the non-retrieval rate is our own measurement across 1,800 observations, with the method in the published methodology.

Related on this site

  • /Analysis — Analysis on iGaming SEO and AI search. Every figure carries a source and a date. No listicles,.
  • /Start with the audit — A free AI visibility audit for iGaming brands. One market, 6 AI surfaces, 3 competitors you.
  • /The four tiers — pricing $900 a month against a market range of $1,500 to $2,500. Four tiers,.
  • /Glossary of terms — an glossary of 51 terms across AI search, technical SEO, links, commercial models.

There is a second denominator underneath this one: whether the model searched at all. On our own runs 57.8% did not, which means the citation rate everyone reports is computed over a pool that is itself split in two.

Questions

Does this mean AI visibility tools are useless?
No, it means the number needs a stated denominator before it means anything. A tool that includes non-retrieving runs and reports three figures separately is genuinely useful. One that reports a single flattering percentage is measuring its own retention instead of your visibility.
Why does the model sometimes not search?
Because it judges that it already knows enough. Retrieval is triggered by the model's own assessment of the question, and it varies by phrasing, topic, model version and week. That variability is exactly why single-run testing produces such confident and such wrong conclusions.
Should I be optimising for retrieval or for citation?
Neither directly, because you cannot control whether a model chooses to search. What you can influence is whether you are found and cited once it does, which is entity strength, fact consistency and presence in the third-party sources models draw on.
How many runs do you need for a reliable reading?
More than most people use. We run a hundred prompts three times across 6 AI surfaces, which is 1,800 observations per market per round, sampled weekly. Single runs are anecdotes, and a great deal of confident advice in this field is built on them.
Free · No obligation

See the same analysis run on your own domain.

The free audit reports retrieval, citation and share separately, with the full denominator included. One market, 6 AI surfaces, 48 hours, raw data attached.

Run my audit

Last reviewed August 2026

Free AI visibility audit · 48 hours Start