LIVE
UTC
DataBackedNews
Technology

AI Search Citations Change 60% Day to Day — Why One Check Lies

A St. Gallen study tracked four AI engines for six weeks and found the sources they cite overlap only 34–42% between consecutive days. AI visibility is a distribution, not a ranking — and needs 7–8 measurements to pin down.

Marcus Okonkwo · Technology correspondent · ·peer-reviewed

If you check whether your brand appears in an AI answer today, then check again tomorrow, you may get two different answers — with no change to your site, the news, or the world. That instability is not a glitch. According to a six-week study from the University of St. Gallen, it is the defining property of AI search.

The measurement

Schulte, Bleeker and Kaufmann (2026) queried four engines — ChatGPT, Gemini, Google AI Mode and Perplexity — daily across four campaigns for 45–46 days, then measured how much the set of cited sources overlapped from one observation to the next, using the Jaccard index and Rank-Biased Overlap (RBO).

Sources churn ~60% per day

Overlap metricSourcesBrands
Day-to-day (Jaccard)0.34–0.420.45–0.59
Same-day repeat runs0.32–0.430.41–0.49
Rank overlap (RBO)0.21–0.300.19–0.30

A Jaccard of 0.35 means that, on average, 65% of cited sources change between two consecutive days. The critical control: even prompts run twice on the same day, minutes apart, overlapped only 32–43%. That rules out news cycles and index refreshes as the main cause — the volatility is intrinsic to the model’s generation process.

Note the second column: brand mentions are more stable than source URLs. Being named as a brand survives the churn better than being cited as a specific link — a useful asymmetry for anyone tracking presence.

A few domains win most citations

Citations are highly concentrated. The mean Gini coefficient was 0.715 across campaigns and engines — Google AI Mode most concentrated (0.782), Perplexity least (0.671). A handful of authoritative domains capture the bulk of visibility, which compounds the big-brand bias documented elsewhere in the GEO literature.

The rule this forces

Because a single query is a noisy sample of a distribution, the authors ran a bootstrap convergence analysis to ask how many runs you actually need:

TargetRuns needed
Per-brand detection (SE < 0.10)≈ 7 runs
Source coverage (SE < 0.10)≈ 8 runs
Stable per-brand rate over time2–4 week rolling window

“Single observations of AI visibility are misleading and risk over- or under-estimating true brand presence.” — Schulte et al. (2026)

So the next time someone reports that a brand “ranks #2 in ChatGPT,” ask how many times they checked. If the answer is once, the number is inside the noise. This is why our methodology treats any single AI-visibility reading as provisional until it’s repeated.

§ Frequently asked questions

Why do AI search citations keep changing? +

Because generative engines are stochastic. Schulte et al. (2026) found cited sources overlap only 34–42% between consecutive days, and 32–43% even when the same prompt is run twice in one day — so most of the churn comes from the model itself, not from news events or index updates.

How many times should I measure AI visibility? +

At least 7 runs per prompt for brand-level visibility and about 8 for source-level coverage, aggregated over a two-to-four-week window. A single observation has a standard error too large to distinguish real visibility from noise.

#AI search#measurement#GEO#visibility#statistics
M
Marcus Okonkwo

Technology correspondent · MSc, Computer Science

Marcus Okonkwo covers artificial intelligence and search for DataBackedNews, focused on what the measurements actually show once the marketing is stripped away. He reads the papers so you don't have to — and cites them so you can.

§ More from the Technology desk