← All posts
30 July 2026

What an AI visibility score hides

One number, three kinds of variance. We ran the same measurement on three consecutive days and the headline moved by half its own value, with nothing changed.

Your AI visibility dashboard says 34. Last month it said 28. Somebody is going to ask whether that is progress, and the honest answer is that nobody looking at that dashboard can tell.

This is not a complaint about a particular tool. It is a property of the thing being measured. Answer engines do not return the same answer twice, they disagree with each other by an order of magnitude about how fast their sources change, and a score built on top of all that is an average of quantities that have no business being averaged. Meanwhile the number is going into a report, and somebody is making a budget decision with it.

So here is what is inside the number, what happened when we ran the same measurement three days running, and the smaller set of things you can actually act on.

The short answer

An AI visibility score blends three different kinds of variance into one figure: how volatile the engine is, how volatile the question shape is, and whether you are in an answer’s stable core or its rotating carousel. Those move independently, so the composite cannot be attributed to anything you did. The sample size that would fix this is roughly 81 to 97 repeats per prompt, which nobody sells and almost nobody would buy. The workable alternative is to stop reading the average and read the extremes: run a fixed question set three times on separate days, then act only on the questions where you are cited every time and the questions where you are cited never. The middle bin is real, but it is not evidence, and it is exactly where dashboards point you.

Three different things are inside one number

How volatile the engine is

SISTRIX measured this at scale and published the method: 82,619 prompts, 1,548,213 snapshots, 17 weeks from 17 December 2025 to 8 April 2026, six countries, weekly reference dates. Verified against the live source on 30 July 2026. Weekly source replacement came out as:

EngineSources replaced per week
Google AI Overviews5%
Google AI Mode56%
ChatGPT Search74%

That is a fifteen-fold spread between the two ends. Any sentence of the form “AI search churns X percent a week” is wrong before it finishes, and a score that averages across engines has folded a nearly static surface into a nearly random one and reported the mean.

How volatile the question shape is

Ranked commercial questions behave nothing like specific buyer questions. ONmetrics (Dave De Vries, published 5 July 2026, 740 live model calls, method disclosed in full) sent 20 “best X” prompts to two flagship models ten times each. The result: “Identical full brand set across two runs: 0.0%” on both. Same prompt, same model, minutes apart, never the same list twice.

That is a real finding, and it is also the easiest one to over-read. A question that asks for a ranked list of ten vendors has an enormous space of nearly equally good answers, so it will churn no matter how visible you are. A question like “what insurance do I need when renting a campervan” has a much smaller space of correct answers. The instability everyone quotes is partly a property of the prompt, not of you.

Worth noting on provenance: ONmetrics is a marketing agency rather than a research institute, so we read the page against our usual test before using it. It publishes its own experiment rather than repeating someone else’s, it states the models, prompt counts, run counts and statistical method, and it does not sell an AI visibility tracker. That is enough to cite. A vendor’s blog post quoting a lifespan figure with no disclosed method, for a product that vendor sells, is not, however plausible the number sounds.

This is the distinction that makes the other two make sense, and it is the one nobody puts in a dashboard. SISTRIX found that 86 percent of prompts have a stable core of a few domains, with everything outside that core rotating at 89 percent per week. In 53 percent of AI Overview prompts, not a single source changed across the full 17 weeks, even though the wording of the answer was often rewritten.

The sharpest version of it is on branded queries: the brand’s own domain was present in all 17 weeks for 43 percent of brand queries, while the co-citations alongside it rotated at 70 percent per week. Read that twice, because it has a direct consequence. Your own presence on your own brand query is relatively stable. Who appears next to you is close to random week to week. So “share of answer against named competitors”, measured once, is one of the least stable numbers you can put in a report, while your own presence measured the same way is one of the more stable ones. Most reports treat them as equally solid.

What three consecutive days did to one measurement

We run this measurement on client properties and on ourselves. On 28, 29 and 30 July 2026 we ran an identical probe on one European company’s site: eight buyer-intent questions about its own service, each against Perplexity Sonar, ChatGPT with web search, and Gemini 2.5 Flash with grounding. Twenty-four question-engine pairs, three days, 72 readings. Nothing on the site changed between the runs, and nothing had been published for it in that window.

The count of pairs that cited the site: 10 on day one, 11 on day two, 15 on day three. The headline number moved by half its starting value in 48 hours, with no cause on our side. Any of those three days, taken alone, would have been reported as the site’s AI visibility.

Now the same 72 readings sorted by pair instead of by day:

Across the three daysPairsShare of 24
Cited all three times9roughly three eighths
Cited twice3one eighth
Cited once3one eighth
Never cited9roughly three eighths

Three quarters of the pairs are unambiguous. They either held every single time or never appeared at all. All of the movement in that 10-to-15 headline came from the six pairs in the middle, a quarter of the set.

The engine split is the part that lines up with SISTRIX. Of the eight questions, Gemini cited the site on all three days for five of them, Perplexity for four, and ChatGPT for none. ChatGPT produced no stable pair at all: its two positive readings were one pair cited twice and one cited once. That is a first-party echo, on a completely different method and sample, of SISTRIX finding ChatGPT Search to be the most volatile surface it measured. Two unrelated datasets agreeing that volatility is a property of the engine is a much better reason to stop blending engines than either one alone.

Two things this is not. It is a pre-existing state we measured, not a result we produced: the site was cited before we touched anything, and none of those citations are ours to claim. And three consecutive days measures re-roll noise, not decay. SISTRIX’s weekly replacement figures measure something genuinely different, which is turnover over time. Conflating the two is common and it makes both arguments worse. One sample, one property, one sector, eight questions, so read the shape and not the specific fractions.

The sample you would need, and why nobody sells it

ONmetrics did the arithmetic that the rest of the market skips. Running a power analysis (Wald interval, plus or minus 10 percentage points, 95 percent confidence), a brand appearing in 30 percent of answers needs about 81 repeats of a prompt to pin that 30 percent down to within 10 points. A brand at 50 percent needs 97.

Eighty-one repeats. Per prompt. To get an answer still carrying a 20-point window. Against a tracked set of 30 questions across four engines, that is most of ten thousand calls for one reporting cycle, and it would still not tell you why anything moved.

Nobody sells that, and it would be irrational to buy it. The useful conclusion is not “measurement is impossible”. It is that the middle of the distribution is unaffordable to measure and should therefore not be the thing you act on. Which is fine, because it was never where the decisions were.

Read the extremes, not the average

Three samples on separate days costs almost nothing and cannot give you a valid percentage. What it can do is sort your question set into three bins, and two of those bins are decision-grade with three samples.

Never cited, on any engine, on any day. This is the strongest and cheapest signal in the whole exercise, and it is the one a composite score buries. It does not mean you rank badly for that question. It means you are not in the retrieval set at all: on that question, nothing you own is a candidate. That is a content gap with a specific address, and it is actionable this week. In our 24 pairs, nine were here.

Cited every time, on the same engine. This is your core, in SISTRIX’s sense. The action is not to improve it. The action is to know which pages carry it and not break them, which mostly means not “refreshing” them into something less quotable. Rebuilds and consolidations are how people lose these without noticing.

Cited sometimes. Real, and not evidence. With three samples you cannot distinguish a page entering the core, a page leaving it, and a page sitting in the carousel being re-rolled. The correct action is usually none, and the correct report line is “unresolved, sampling again”. This is the bin every dashboard renders as a number with an arrow next to it.

A zero is also worth more than a yes, which is not obvious. Absence across every engine and every sample has essentially one explanation: no candidate source. Presence has many, including one strong page, one lucky re-roll, or a directory that mentions you. So the negative half of the measurement is the half that reliably tells you what to do, and it is the half that gets deleted when results are compressed into a score.

What Search Console now shows, and what it does not

One genuine improvement, and it is new enough that plenty of published advice is behind it: as of June 2026 Google ships a Generative AI performance report in Search Console, covering AI Overviews and AI Mode as their own view rather than folded silently into web Search. It is rolling out to a subset of properties first, so not seeing it does not mean you have no data.

Turn it on and read it, with the limits stated plainly. It reports impressions only, so there are no clicks, no CTR and no position. It covers Google’s AI surfaces and says nothing about ChatGPT, Perplexity or Claude. It excludes Search Labs experiments. It tells you your links were shown inside a Google AI answer. It does not tell you what the answer said, whether you were recommended or merely listed among others, or who was named beside you.

That last gap is why probing the engines directly is not redundant with Search Console, and Google is unusually direct about the general case: its guidance on third-party tools states that “Third-party tools don’t have access to our internal ranking data” and that “Some third-party services provide data that some users of those tools misinterpret as somehow being from Google”. Two instruments, two different blind spots. Anyone selling you one and calling it complete is selling.

What a report should say instead of a score

Four changes, none of them expensive:

  1. Per engine, never blended. A single figure across engines with a fifteen-fold volatility spread is not a measurement. Three or four columns cost nothing and mean something.
  2. A rate across a window, not a yes on a date. “Cited in four of the last six weekly samples” is interpretable. “Cited: yes (28 July)” is a coin toss with a timestamp.
  3. The bin, not the number. Always, never, or unresolved. Then the count in each bin, which is a real trend line as it shifts.
  4. The questions where a competitor is named instead of you, and the questions where the engine says something false about you. Neither is a score, both are work orders, and the second one is usually the most valuable output of the whole exercise.

When one reading is genuinely enough

There is a case for the cheap version, and pretending otherwise would be selling. If the question is “is any of this happening at all”, one pass over 20 buyer questions answers it in an afternoon. When the answer is that you are absent almost everywhere, that is not noise. Nine of our 24 pairs were absent on all three days, and a single day’s reading would have identified most of them correctly. A first look is a legitimate thing to buy.

What one reading cannot do is tell you whether something worked. That needs the same questions, on separate days, over enough weeks to see a rate. If your report has a single number with a month-over-month arrow on it, the arrow is describing the instrument.

We wrote about the structural reason companies find themselves absent in the first place in why AI assistants cite everyone except you, and the measurement standard behind the bins above is part of how we run AI search visibility as a service. If you would rather run it yourself, everything in this piece is the method, not a teaser for it: a fixed question set, three samples on separate days, per engine, sorted into three bins, acting on two of them.

Tell us what’s broken.
We’ll tell you the truth.

Book a free call →
Reply within one business day ¡ EN / LV