LLM Visibility Monitoring
LLM visibility monitoring answers one question every month: how often do the six major AI surfaces name your brand when someone asks what your players ask? 100 prompts, three runs each, sampled weekly.
Most AI visibility scores on the market are calculated on a denominator that quietly excludes the majority of what happened. The number looks like visibility. It is share of citations among the responses that happened to contain citations, which is a smaller and much friendlier universe. Here is our method, including the parts that make our numbers look worse.
What LLM visibility monitoring measures, precisely
The denominator problem
Language models do not always retrieve. In one measured configuration, 57.8% of repeated ChatGPT runs never activated web search at all — the model answered from what it already held, without going out to look.
If a tool throws away those responses because they contain no citations, and then reports how often you appear among the remainder, the resulting percentage is inflated and the direction of the error is always favourable. You are told you appear in 40% of answers. What was measured is 40% of the 42% of answers that cited anything at all.
A dashboard that never shows you a bad month is not measuring. It is reassuring.
What we count
We decompose every run into three questions and report all three separately, so nothing hides inside an average.
- /Retrieval rate. Did the model go and look, or answer from memory? This varies enormously between surfaces and between prompt phrasings, and it sets the ceiling on everything else.
- /Citation rate. Of the runs that retrieved, how many named any brand at all? Some questions produce a generic explanation and no shortlist, which is neither a win nor a loss for you.
- /Your share. Of the runs that named brands, how often were you one of them, and in what position within the answer.
Three numbers, not one. A brand can be flat on share while its retrieval rate doubles, which is real progress that a single headline figure would erase.
Why a single run tells you nothing
Model output is not deterministic. Ask the same question twice and you can get two different shortlists, in a different order, drawn from different sources. A screenshot of one answer is anecdote, not measurement, and screenshots are what most agencies send.
We run every prompt three times per surface. At 100 prompts across 6 AI surfaces that is 1,800 observations per market per sampling round, and 7,200 a month. It is enough to separate a genuine change from ordinary variance, which is the whole point.
6 AI surfaces, and why we never average them
We sample ChatGPT, Google Gemini, Perplexity, Microsoft Copilot, Google AI Overviews and Google AI Mode. Overviews and AI Mode are counted separately because they behave differently despite both being Google — different triggering, different citation patterns, different referral behaviour.
The reason for reporting them apart is empirical. A 2025 analysis of roughly 366,000 citations found that citations concentrate on a small pool of outlets and that different engines pick different ones, with low agreement between platforms. Averaging 6 AI surfaces into one score destroys exactly the information you would act on.
Scale differs too, and we say so on the report. AI Mode accounted for only 0.34% of Google searches between January and April 2026, and refers traffic out at 1.6% to 2.5% against Google's 17% to 19%. It matters, but not equally, and a chart that implies otherwise is misleading you.
Building the prompt set
From your commercial terms
Prompts are derived from the keywords that already earn your deposits, rewritten as the questions a person would actually type into a chatbot. Not a generic industry list reused across clients.
Per market, per language
A Brazilian prompt is written in Portuguese and pinned to a Brazilian locale. Models answer differently by region and language, and a translated prompt measures the wrong market.
Across the funnel
Legality and licensing questions, comparison questions, payment and withdrawal questions, brand questions. Each band behaves differently and each tells you something separate.
Frozen, then versioned
The set stays fixed so month-to-month numbers compare. When we add prompts we version the set and report both series, rather than quietly resetting your baseline.
Keeping the runs clean
Three things contaminate this kind of measurement, and all three are easy to get wrong.
- /Session memory. A model that has already discussed your brand in the same conversation will favour it. Every prompt runs in a fresh, unauthenticated session.
- /Location. Answers change with inferred location, so each market's runs are pinned to that market rather than executed from wherever the tooling happens to sit.
- /Personalisation. Logged-in accounts carry history and preferences. We do not use them, which means our numbers are colder than an in-house test run from someone's own account, and closer to what a new player would see.
What the monthly report contains
- /Retrieval, citation and share figures for each of the 6 AI surfaces, per market, against last month and against your baseline.
- /Three named competitors tracked on the identical prompt set, so the comparison is like for like, not against a different question.
- /Prompt-level detail: which specific questions you win, which you never appear in, and which flipped this cycle.
- /Source tracing. Where a model cited, we log the URL. That turns a score into an instruction, because you can see which page or publication is doing the work.
- /A written read from us on what moved and why, including the months where nothing moved and we think the answer is patience and not activity.
What does a good score look like?
There is no universal good score, and any supplier quoting one is guessing. What matters is your own trend across three separately reported figures. There is no published industry benchmark for this, and anyone quoting one has invented it. What we can give you is context from our own sampling: in most regulated markets a handful of brands hold the majority of mentions, a long tail appears occasionally, and a large group never appears at all.
Operators consistently underestimate how large that third group is. One test across three queries found 55% of live casino providers invisible to AI entirely — not ranked poorly, not mentioned rarely, absent. If your first audit puts you there, it is a common starting point instead of an unusual failure.
So we do not grade you against a made-up standard. We grade you against the 3 competitors you named, on identical prompts, and against your own previous months. Those are the only two comparisons that mean anything right now.
How is this different from a rank tracker?
A rank tracker answers where your URL sits. This answers whether your brand gets named. Those used to be close to the same question and are no longer.
In mid-2025 roughly three quarters of pages cited in AI Overviews also held a top-ten position for that query. By early 2026 the overlap had fallen to between 17% and 38%. Your rank report can be entirely stable while your presence in answers halves, and nothing in the rank report will tell you.
What we do with the number
A score you cannot act on is decoration. Source tracing is what makes this one actionable: if 3 competitors appear in Ontario prompts because a single comparison site keeps getting quoted, the work is getting accurate information onto that site instead of writing another page on your own domain.
Where the decisive source turns out to be a thin page on your own site, the fix is usually content work, never more measurement. The index tells you which; it does not do the writing.
That pattern is common. Around 68% of AI citations point at third-party sources, never brand-owned ones, which means the fix usually lives outside your CMS. Measurement is how you find out where.
What this is not
- /It is not an AI ranking. No such fixed position exists. Anyone selling you a single AI rank number is selling variance dressed as a metric.
- /It is not a traffic forecast. AI platforms sent 1.13 billion referrals in June 2025, set against the roughly 191 billion Google handled over those same thirty days. Visibility here is worth having; it is not a volume replacement.
- /It is not precise to a decimal point. We report ranges and trends, and we will tell you when a movement sits inside normal variance and not dressing it up as a result.
- /It will not fix a Matthew effect on its own. Models favour sources that are already heavily cited, documented in 2025. New brands climb slowly, and measurement shows you the climb instead of shortening it.
How to buy it
Free audit
One market. Six surfaces. Three competitors you choose. Forty-eight hours. No call required first and no obligation afterwards. Most people start here.
Standalone monitoring
Monthly subscription if you want the measurement without the programme, including if you are running the optimisation work in-house or with another agency.
Inside a retainer
Included from the Growth tier upward, where the number drives what we build next and not sitting in a separate report.
Your own team, eventually
We will document the method so you can reproduce it. If you get to the point of running it yourself, that is a reasonable outcome and we will say so.
Prices are published instead of quoted after a discovery call. If the free audit shows your visibility is already strong, we will tell you that too, and you will not hear from us again unless you ask.
Sources for the retrieval and citation figures: the 57.8% non-retrieval measurement and the citation concentration analysis are third-party research published in 2025 and 2026; the 17% to 38% AI Overview overlap comes from Semrush and BrightEdge; the AI Mode share of Google searches is Similarweb, covering January to April 2026. Each is attributed on the page it appears on and re-checked quarterly.
What monitoring will not do
It will not move the number. Measurement is diagnosis and not treatment, and a report nobody acts on is an expensive subscription. It will not produce a single AI ranking, because none exists. And it will not make a small volume large — AI referral traffic remains a fraction of Google's, and we would rather say so than sell the number as bigger than it is.
Sources for the retrieval and citation figures here: ICODA. The links go to the original studies, never to a summary of them.
What monitoring will not tell you
This is a measurement service. Measurement has edges, and here they are.
- /Why a model named someone. We can show the sources it drew on and the pattern across runs. The weighting inside the model is not visible to us or to anyone outside the lab.
- /What will happen next month. A model update can reshuffle every number in a week. The Index tells you where you stand and which sources moved, not where you will stand.
- /A number you can compare to a competitor's report. Different prompt sets produce different denominators. Our figures are internally consistent over time, and not portable to somebody else's methodology.
- /Whether presence became revenue. We report citations. Attribution from an AI answer to a deposit is not something any tool measures honestly today, and we will not pretend otherwise.
Monitoring is worth buying if you intend to act on it. As a number for a board slide it is an expensive way to feel informed.
LLM visibility monitoring rests on a method that is published in full, including the prompt bands, the counting rules and the known biases. Any supplier figure without that alongside it, ours included, should be treated as unverifiable.
Related on this site
- /What we do — Nine iGaming SEO services for casino, sportsbook and affiliate brands. Technical, content,.
- /AI visibility index — The AI Visibility Index: how often ChatGPT, Gemini, Perplexity and Google name your gaming.
- /Who we work with — for operators, affiliates, B2B suppliers, crypto casinos and land-based venues. Six.
- /Gambling markets we cover — by market across 40 jurisdictions. Regulator, licence framework, tax, advertising.
- /gambling SEO compliance — Our gambling SEO compliance policy. Licensed operators only, verified at onboarding, prohibited.
- /Glossary of terms — an glossary of 51 terms across AI search, technical SEO, links, commercial models.
One measurement is worth separating out before you buy any of this: most ChatGPT answers about your brand never searched the web at all. On our prompt sets that is 57.8% of runs, and it needs the opposite work from the runs that did search.
One finding worth knowing before you buy any measurement: in France and Poland, eighteen of twenty-one prompts returned an unlicensed casino. Your competitor set in AI answers is not the licensed operators you benchmark against.
For suppliers the question is narrower: whether you are named when the category is asked. iGaming B2B provider visibility covers what that takes.
Common questions
How often should we be measuring this?
Why do your numbers look lower than the tool we tried?
Can you track brands you are not working with?
Does measuring change anything on its own?
What if we already rank well in Google?
Can we see the raw data?
Get your first reading free.
One market, 6 AI surfaces, 3 competitors, 48 hours. You will see the retrieval, citation and share figures separately, and the named sources behind each one.
Run my audit