AI Visibility: Distribution, Not a Number
AI SEOThree independent studies, one conclusion: AI visibility should be treated as a probability distribution.
Key Takeaways:
-
AI visibility is a probability distribution: variability is designed into generative engines – it is not a tracker defect.
-
Same-day identical prompts overlap only 32-42% of times on cited sources: variability results from AI models’ own stochasticity, not index or algorithm changes.
-
A single citation-share figure hides a wide error bar: gaps of ≤ 5-7% percentage points are often statistical ties.
-
Only the top 5 ranking positions are defensible: below the top 10, ordering approaches a coin flip between samples.
-
Collection budgets must be set per engine, based on the effective sample size: SearchGPT can require three times Gemini’s data.
Executive Summary: AI visibility dashboards report citation shares as fixed numbers, when they should be treated as a probability distribution due to generative engines being inherently stochastic. Three independent 2026 studies confirm a single reading is mostly noise. Reliable measurement requires repeated sampling, reported ranges, and per-engine collection budgets determined by effective sample size..
Most AI visibility dashboards out there present citation shares and rankings as fixed facts.
They aren’t.
Yet, brands continue to base their AI SEO around them, hoping they’ll win AI visibility.
They won’t.
Three independent studies from 2026 explain why:
AI visibility was never static to begin with.
Why does the same AI query cite different sources every time?
When a user types in a query, answer engines retrieve a pool of candidate sources into their context window. However, they do not use all of those sources to synthesize the answer – they select only a few, varying between them every time.
On the surface, this mechanism may seem random, but it’s actually stochastic: probability-driven and non-deterministic, but statistically analyzable. This variability is not a tracker defect or a consequence of an algorithm update – it’s baked into the system.
This isn’t an assumption either. Three papers from two unaffiliated teams arrived at this conclusion independently:
- “Don’t Measure Once: Measuring Visibility in AI Search (GEO)” – Schulte, Bleeker, and Kaufmann; University of St. Gallen (arXiv 2604.07585)
- “Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement” – Sielinski; IQRush (arXiv 2603.08924)
- “From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement” – Sielinski; IQRush (arXiv 2607.10341)
Although the first two are from a commercial vendor (IQRush), the St. Gallen paper is academic and independent, and it reaches the same result via its own dataset. This is too big of a congruence to wave the finding off as “coincidental” or “circumstantial.”
The cleanest proof: isolating the engine
When St. Gallen researchers issued identical prompts multiple times on the same day, cited sources overlapped in just 32% – 43% of the cases. During this time window, external factors that could otherwise explain citation change (index refreshes, ranking-algorithm updates, edits to the cited pages) did not have time to move
Therefore, every variation left over had to come from inside the model itself, proving the engine’s own stochasticity.
Then, Sielinski’s checksum control sealed the deal: cited pages remained largely unchanged between samples, while the citations reshuffled. Therefore, the churn had to be the engine choosing differently, not the content of the cited pages.
Can you trust a single AI visibility reading?
Not by a long shot. A lone citation-share figure is only a point estimate with a wide error bar. The problem is that standard trackers report a single clean number, not a range. So, unless your tool samples repeatedly and computes a confidence interval (CI), the error bar is typically invisible and the figure looks more certain (and actionable) than it really is.
Sielinski’s running-gear example (via SearchGPT) makes the risk concrete:
- Source #1: ~9.5% citation share, CI roughly 5.5%-12.5%
- Source #2: ~6.0% citation share, CI roughly 4.0%-8.0%
Logically, Source #1 clearly leads over Source #2 by a full 3.5 percentage points. Statistically, it’s a tie, because the entire Source #2’s confidence interval sits inside Source #1’s. However, the real danger here is that the problem isn’t confined to the leaderboard.
The error bar you aren’t shown
A 3.5-point lead that is statistically a tie
Citation share on SearchGPT, with 95% confidence intervals
Source #2’s entire interval sits inside Source #1’s. The gap looks decisive on a dashboard, but the data cannot rule out that Source #2 is actually ahead.
Source: Sielinski, “Quantifying Uncertainty in AI Visibility” (arXiv 2603.08924); ZeroClick Labs.
Instability across the whole ranking
Across all three papers, only the top ~5 positions are defensible, and the top ~10 are usable. Everything below that is practically a coin flip from one sample to the next, with precision falling off rapidly (e.g., by ranks ~26-30, CIs can span 80 positions). Two consequences follow:
- A small gap (approx. ≤ 5-7 percentage points) between a brand and its competitor is real only if it persists across repeated samples, instead of collapsing into overlap on the next run.
- A single pre/post-intervention reading of a brand’s citation share (or Share of Voice) is not enough to validate a content change. A few percentage-point rise (e.g., a 6% → 9% jump) can be pure run-to-run noise, rather than the intervention working.
The practical implication is simple: measure both the before AND the after, multiple times.
How many times do you need to measure AI visibility?
There’s no single “universal” number, because measuring AI visibility requires sampling along two different axes, and since every engine sports a distinct citation behavior, the number of samples differs. However, the St. Gallen and Sielinski’s research provided us with working floors:
Axis 1: Repetition (# of times to re-run the same prompt to average out the noise)
- Brand presence: minimum of 7 runs per prompt;
- Source coverage: minimum of 8 runs per prompt;
Axis 2: Breadth (# of distinct prompts to run to capture a whole topic / decision space)
- Citation-share precision: roughly 40-50 queries for Gemini; 150+ for SearchGPT;
- Convergence overall: 33-94 responses across platform-topic combinations;
Time: St. Gallen research suggests layering both over a rolling window of 2-4 weeks to smooth the day-to-day turnover.
The reason why SearchGPT needs three times the data of Gemini lies in the effective sample size (the volume of independent information each response actually carries). This figure depends on two factors:
- Citation density: number of citations per answer (Gemini ~40, SearchGPT ~6).
- Citation clustering: number of same domains repeated within one answer (repeat citations to same domain are correlated, not independent, adding little new information).
Effective sample size can be far less than the raw citation count, meaning that a per-engine collection budget cannot be set by counting citations. Doing so can result in over-collecting on one engine and under-sampling on the other, and under-sampling equals noise. Therefore, we can expand on the previous practical implication.
How to sample for reliability?
A brand should sample every prompt 7-8 times per prompt, run a large and diversified prompt set (dozens to 150+, depending on the engine), and extend the measurement window to at least two weeks.
Working floors from the research
How much sampling AI visibility actually needs
Two independent axes, plus a time window. Minimums — not targets.
| Axis | What it measures | Minimum |
|---|---|---|
| RepetitionRe-running the same prompt to average out run-to-run noise | Brand presence | 7 runs / prompt |
| Source coverage | 8 runs / prompt | |
| BreadthDistinct prompts needed to cover the full topic and decision space | Citation-share precision — Gemini | ~40–50 queries |
| Citation-share precision — SearchGPT | 150+ queries | |
| TimeRolling window layered over both axes | Smoothing day-to-day turnover | 2–4 weeks |
The Gemini–SearchGPT gap is not about citation volume. It is effective sample size: how much independent information each response actually carries. Budget per engine, never by citation count.
Sources: Schulte, Bleeker & Kaufmann (arXiv 2604.07585); Sielinski (arXiv 2603.08924, 2607.10341); ZeroClick Labs.
Brands adopting this methodology reduce the risk of acting on noise, making decisions based on false precision, and most importantly, unknowingly misallocating budget toward movement that was never real to begin with.
Tina Clarke is the AI SEO Manager at ZeroClick Labs, specializing in AI search optimization and Generative Engine Optimization (GEO). With a strong foundation in content strategy, technical SEO, and operations, she leverages her expertise to help brands shift from traditional rankings to discoverability and excel in AI-driven ecosystems.
Your dashboard isn’t telling you how confident it is
It’s hiding it
And the budget allocated against a figure that was never real is a budget wasted.
ZeroClick Labs measures your AI visibility the way it should be measured – repeatedly, per engine, and with the uncertainty made visible.
So stop wasting your budget on noise.
“Our agency had no idea how to approach AI visibility. ZeroClick only does this one thing so they actually know what works. Worth every penny just to not waste time figuring it out ourselves.” – Jay