Same engine, same question, three times in a row: in 12 of 27 cases it did not give the same answer
Measuring once does not give you a measurement. It gives you a number.
Of 269 questions asked three times in a row to the same engine, the brand appeared at least once in 27. In 12 of those 27 it did not appear all three times: 44.4%, with a 95% confidence interval between 27.6% and 62.7%.
- Sample
- 269 questions asked three times in a row to the same engine · 27 of them with the brand present at least once
- Period
- June – September 2026
- Method
- Three repetitions of each question in the same session, same engine and same transport. We check whether the brand appeared all three times.
- Language
- Spanish (Spain)
Why the denominator is 27 and not 269
Because in most questions the brand appears none of the three times, and those agree because nothing happened.
Counting all 269 would put disagreement at 4.5% and sound reassuring. It would be a lie: it would be diluted with the questions where there was nothing to decide. The relevant question is "of the times the engine could have named you, did it always?", and the answer is no, almost half the time.
The interval is wide — 27.6% to 62.7% — because 27 cases are few, and that is stated too. But even the optimistic end leaves disagreement above one question in four.
The engine does not lean on the same sources either
Across those same 269 repeated questions, the sources the engine cites when answering the same question twice overlap by only 70.9%. Almost a third of the bibliography changes between one run and the next, with neither the question nor the moment changing.
What we cannot say about ChatGPT
For ChatGPT we only have 8 repeated questions in which the brand appeared at all. With eight cases nothing can be claimed, so we do not claim it — however striking the raw figure looks. Publishing that percentage would be exactly the error this finding is about.
What we do about it
- Three repetitions per question in the weekly measurement, and the metric is the average of the three. A single pass is a coin toss in the format of a report.
- Wilson interval and n next to every percentage. A bare 23% looks firm and is not.
- A small week-on-week move is not news. If two weeks have overlapping intervals, the honest thing is to say they cannot be told apart, not to draw an arrow.
What this finding does NOT prove
- It does not say the engine is broken. A generative model produces different text on every run by design; variability is not a defect, it is the nature of what is being measured.
- Twenty-seven cases are few. The point estimate (44.4%) is far less reliable than the conclusion that the disagreement exists and is not marginal.
- It does not extrapolate to other engines. The Perplexity figure says nothing about Google AI Mode or ChatGPT.
- It does not measure whether three repetitions are enough. It only proves that one is not.
Frequently asked questions
How many times should a question be repeated to measure properly?
We do not know precisely and we are not going to invent it. What these data prove is that a single pass is not enough: in almost half the questions where the brand could appear, the result depended on which run you looked at. We use three and publish the interval.
Why does an AI search engine not always give the same answer?
Because it generates the text on every query instead of returning a stored list. Two runs of the same question can retrieve different sources and word the answer differently, and with that name or not name a brand.
Does that invalidate AI visibility metrics?
No, as long as they are published with their sample and their interval. What it invalidates is the bare number: a single-pass percentage, with no n and no interval, cannot tell a real change from the engine’s own noise.