How we measure whether AI cites your brand
This page describes the whole method: what gets asked, to which engines, how many times, how each figure is calculated and — above all — what this kind of measurement cannot tell you.
What gets asked
Around 90 real questions from the client's market, spread across business verticals: the ones somebody about to hire would ask, not the ones an algorithm comes up with.
Questions are organised by commercial intent, not by product. That difference matters more than it sounds: "best employment lawyers in Valencia" is a recommendation question, but "I've been fired and I think it was unfair, what do I do?" is a problem question — and that is where a hire is decided. A question set that only covers the product measures half the market.
Every question belongs to a segment, so the result is not a single number but a map: which parts of your business you show up in and which you don't.
To which engines
ChatGPT, Google AI Mode and Perplexity, through each one's real interface and not only through its API, so we see what a person sees.
That distinction is not a technical footnote. The same model answers differently depending on whether it searches the web or replies from memory, and the gap can be enormous: in one internal run, the same 90 questions returned 0% presence without web search and 22% with it. Both are correct numbers for two different things. That is why every measurement records how it was made.
How many times
Three repetitions per question and engine, by default. These engines are noisy: the same question can return different answers on the same day.
With a single pass there is no way to tell a real gain from a coincidence. This behaviour has been measured: when one engine was tested against itself, it contradicted itself in 1 out of every 10 questions, and the sources it cited only overlapped with its own previous run about a third of the time. Any tool handing you a number from a single pass is handing you noise with two decimal places.
How the figure is calculated
Every percentage is published with its Wilson 95% confidence interval and the number of observations behind it.
A percentage without its margin looks firmer than it is. Over 90 questions, 10% presence actually means somewhere between 5% and 18%. Publishing only the 10% invites you to read noise as progress, and to make content decisions on a difference that does not exist.
The effective sample
There is one refinement almost nobody applies, and it changes the result. The three repetitions of the same question are not three independent observations: they resemble each other far more than three different questions do. Counting them as independent would narrow the interval artificially and make a shaky figure look solid.
So the percentage is computed over the full sample — that is the measured value and it is not touched — but the margin is computed over the genuinely independent observations. The result is a wider, more honest interval.
Branded questions are excluded
If the question already names the brand ("what is company X like?"), the answer will mention it almost every time. That is presence by construction: including it raises the percentage without visibility having changed at all.
Those questions are measured and shown separately, but they do not enter the headline figure. It is one of the easiest ways to inflate a visibility report, which is why it is worth asking about when comparing tools.
What gets recorded with every measurement
The method stamp: engine, model, whether web search was used, with what context, and how many repetitions.
Without that record, a time series can be averaging two different things without anyone noticing — and the average of two methods means nothing. When a series mixes measurements taken in different ways, it says so instead of averaging blindly.
Measurements taken before this record existed are marked as "method unknown". We do not invent one for them.
What this measurement does NOT say
Being cited is not being visited. Citation and traffic are two different things and are reported separately.
- Web analytics undercount the AI channel: many of those visits arrive with no referrer and are logged as direct traffic. A zero in analytics is not always a zero.
- Percentages are computed over a question set defined for your sector, not over the world's real search volume. They are comparable with themselves and with your competitors within that same set.
- No measurement tells you why a model picks who it picks. We observe correlations — which sources it cites, which pages it rewards — and act on them, but the internal mechanism is opaque to everyone.
On return, without overselling it The widely repeated claim that ChatGPT traffic "converts nine times better" does not hold up: it comes from a single-site case study. Published evidence points to AI traffic converting roughly the same as organic, in a range from 13% worse — Kaiser & Schulze, Marketing Science, 2026, across 973 sites and peer-reviewed — to 31% better according to other analyses. The case for investing is not the conversion rate: it is the volume and the trajectory of the channel.
Summary
| Decision | What we do | Why |
|---|---|---|
| Questions | ~90 per client, by intent and segment | Problem intent is where hiring happens |
| Engines | ChatGPT, Google AI Mode, Perplexity | That is where customers ask |
| Repetitions | 3 per question and engine | They contradict themselves 1 in 10 times |
| Margin | Wilson 95% over the effective sample | Repetitions are not independent |
| Branded questions | Out of the headline figure | They give presence by construction |
| Traceability | Method stamp on every measurement | An average of two methods means nothing |
Try it on your own site
The free audit applies this methodology to your domain: what share of your sector's questions you show up in today, and who shows up instead of you. No signup, no card.
Run my free audit →Metrics · How to choose · Use cases · Changelog