AI visibility measurement methodology, published in full
Every agency selling AI visibility reports a number. Almost none publish the AI visibility measurement methodology behind it, which makes the number unfalsifiable and therefore worthless. This is the whole method: how prompts are built, how runs are sampled, what counts as a mention, and where the measurement breaks.
Why publish an AI visibility measurement methodology at all
An AI visibility score is a claim about how often a machine names your brand. Without the prompt set, the sample size and the counting rules, that claim cannot be checked by you, argued with by your team, or compared to anything. Publishing the method is the only thing that turns it into a measurement.
The arithmetic is fixed and published: 100 prompts per market, 3 runs each, 6 surfaces, giving 1,800 observations per cycle. The prompt set is split 30% discovery, 25% comparison, 25% attribute, 20% trust, and those proportions do not move between numbered revisions. Below roughly 60 prompts the trust band drops to 60 observations, which is too few to read.
Why should an AI visibility method be published at all?
Why publish a method a competitor can copy
Because a number nobody can audit is a sales device, not a metric. Three practical consequences follow from keeping a method secret, and all three land on the buyer.
- /You cannot tell a real movement from a change in the prompt set. An agency that quietly adds ten easier prompts can show you improvement without your brand moving at all.
- /You cannot compare two suppliers. Different denominators produce different percentages from identical underlying visibility, which is the denominator problem in its commercial form.
- /You cannot bring the work in-house later. A method you were never shown is a method you cannot maintain, which is a retention strategy rather than a service.
The trade is straightforward. Publishing this costs us a small amount of differentiation and buys the client the ability to check us. In a vertical where nobody can verify anybody's claims, that is the only differentiation we think holds.
How the prompt set is built
One hundred prompts per market, fixed at the start of an engagement and changed only in a numbered revision. The set is built from four bands, in fixed proportions, so that the mix cannot drift toward whatever the brand happens to be winning.
| Band | Share | What it looks like | Why it is included |
|---|---|---|---|
| Category discovery | 30 | best online casino in Ontario, where to bet on the Brasileirão | The queries with the most volume and the least brand intent. Hardest to appear in, and the ones an operator most wants. |
| Comparison | 25 | is X better than Y, which sportsbook pays out fastest | Where a model weighs two named brands. Presence here usually depends on third-party sources, not your own site. |
| Attribute and constraint | 25 | casino that accepts Pix, sportsbook with no withdrawal limit | Long-tail, specific, and the band where a well-documented operator can win against a larger one. |
| Trust and safety | 20 | is X licensed in Germany, has X ever been fined | Where an operator is most exposed. A model with nothing to cite will either omit you or repeat whatever forum thread it found. |
Proportions set August 2026. Any change to them is a new revision number and the report says so.
Prompts are written in the market language, not translated from English. A Brazilian bettor asks a different question from a German one, and a translated prompt measures a query nobody types. Brand names are used only in the comparison and trust bands, because a brand-name prompt measures whether the model knows you exist, not whether it recommends you.
No prompt is written to be winnable. If a prompt only returns your brand because it describes your brand, it is measuring nothing and it is removed.
Sampling, runs and why one run is worthless
Every prompt runs three times per surface, in separate sessions, with no memory carried between them. Three is the smallest number that distinguishes a stable answer from a coin flip, and the variance between runs is reported rather than averaged away.
- /Three runs, one hundred prompts, six surfaces gives 1,800 observations per market per cycle.
- /Runs are spread across the sampling week rather than fired in one batch, because answers move with index freshness and with load.
- /Sessions are clean. No account, no history, no personalisation, because a personalised answer measures the account rather than the brand.
- /Where a surface refuses to answer a gambling prompt, that is recorded as a refusal and reported separately. Refusals are not zeros and folding them in would flatter every operator in the set.
A single run tells you nothing. In our own measurement, a brand can appear in one of three runs on the same prompt in the same hour, which means any report built on one pass per prompt is reporting noise with a decimal point on it.
What counts as a mention, and what does not
The counting rules matter more than the prompt set, because this is where a generous reading can double a score. Ours are deliberately strict.
| Outcome | Counted as | Note |
|---|---|---|
| Brand named in the answer text | Mention | The brand has to be named. A description that fits you is not a mention. |
| Your domain cited in the sources panel | Citation | Counted separately from a mention, because the two come apart constantly. |
| Named and cited in the same run | Both | Recorded once in each figure, never doubled into a combined score. |
| Named only inside a quoted user review | Mention, flagged | Reported, but flagged, because it is somebody else's characterisation of you. |
| Named as an example of what to avoid | Not counted | Negative presence is logged in the report text and excluded from the figures. |
| Domain cited for an unrelated page | Not counted | A citation has to support the answer given. |
Rules fixed August 2026.
Three figures, never averaged
The report carries three numbers and refuses to combine them, because combining them hides the thing you most need to see.
- /Retrieval rate. Of the prompts where the surface searched the live web at all, what share pulled anything from your domain. This is a technical and indexation measure before it is a content one.
- /Citation rate. Of all runs, what share cited your domain in the sources. This is the number most affected by third-party presence, and the one an operator can least influence directly.
- /Share of answers. Of all runs, what share named your brand in the answer text. This is the number that maps closest to commercial outcome, and the slowest to move.
A single blended score would let a strong retrieval rate mask an absent share of answers, which is the most common shape we see: the model can find you and still will not recommend you. Those are different problems with different fixes and they should not share a number.
The six surfaces, and how each is queried
Six surfaces, each queried in its own native interface rather than through an aggregator API, because an aggregator answers a different question from the one a user asks.
- /Google AI Overviews, on a clean session in the market locale, recorded with the citation panel expanded.
- /Google AI Mode, queried separately, because it retrieves differently from Overviews and conflating the two was the most common error we found in competitor reports.
- /ChatGPT with browsing available but not forced, so the refusal to search is itself an observation.
- /Gemini in the market locale.
- /Perplexity, where the source list is the primary signal rather than the prose.
- /Microsoft Copilot, included because it inherits Bing's index and therefore behaves unlike the other five.
In our measurement a large share of runs never searched at all and answered from training data. Those runs are reported as a separate line, because a brand invisible to a model that did not search has a different problem from a brand invisible to one that did.
Known biases in this method
Every measurement has a shape it cannot see. Four of ours, stated so you can weigh the number accordingly.
- /Recency bias in the prompt set. Prompts fixed in one quarter reflect that quarter's language. A market where terminology shifts fast will drift out of the set before the revision catches it.
- /English-language source bias inside the models themselves. A non-English market can show artificially low citation rates because the models lean on English sources for gambling topics.
- /Small-sample noise on the trust band. Twenty prompts across three runs is 60 observations, and a single well-placed forum thread can move that band more than a quarter of content work.
- /Our own selection. We write the prompts. We publish them so the selection can be argued with, but the selection is still ours and that is a limitation, not a disclaimer.
What happens when a model updates
A surface change can move every number in a week without your site changing at all. When that happens the series is broken and we say so on the face of the report, in the month it happens.
- /The affected surface is marked as a break and the prior series is not restated to make the trend look continuous.
- /The comparison for that month is against the same surface only, never against the six-surface aggregate.
- /Where the change is large enough that the prompt set no longer measures the same thing, the set is revised and given a new number, and the old figures stay published under the old number.
Quietly rebasing a series is the single easiest way to make an AI visibility report say whatever the agency needs it to say that quarter. If you take one thing from this page, make it the question you ask your current supplier.
Replicating this yourself
The method set out here is enough to run it in-house, and for some operators that is the right answer. What it costs is roughly a day a month of a competent analyst's time per market, plus the discipline to leave the prompt set alone between revisions.
- /Build the hundred prompts against the four bands. Getting the proportions right matters more than getting individual prompts clever.
- /Run three passes per surface across a week, in clean sessions, and log the raw answers rather than a summary of them.
- /Count against the rules above, including the exclusions. The exclusions are what stop the number inflating over time.
- /Report the three figures separately with the refusal line alongside. Resist the blended score, however much easier it is to present.
If you would rather not, the free audit runs exactly this method on one market, and the report includes the prompt set so you can check the work against what is written here.
Related on this site
- /What we do — Nine iGaming SEO services for casino, sportsbook and affiliate brands. Technical, content,.
- /Tiers and what they buy — pricing $900 a month against a market range of $1,500 to $2,500. Four tiers,.
- /Who we work with — for operators, affiliates, B2B suppliers, crypto casinos and land-based venues. Six.
- /Gambling markets we cover — by market across 40 jurisdictions. Regulator, licence framework, tax, advertising.
- /gambling SEO compliance — Our gambling SEO compliance policy. Licensed operators only, verified at onboarding, prohibited.
- /Glossary of terms — an glossary of 51 terms across AI search, technical SEO, links, commercial models.
Sources for the measurement figures here: the AI Overview overlap range is Semrush and BrightEdge, early 2026; the third-party citation share is Seer Interactive, 2026; and the ChatGPT non-retrieval rate is our own measurement, which is why the prompt set and counting rules above are published rather than summarised.
Common questions
Why publish a method your competitors can copy?
How many prompts do you need for a reliable figure?
Why report three figures instead of one score?
What happens to the trend when a model changes?
Can we run this ourselves instead of buying it?
Is any of this an industry standard?
See the method run on your own domain.
One market, six AI surfaces, three competitors you name. The report includes the prompt set, so you can check it against this page.
Run my audit