AI Search Citations Change 60% Day to Day — Why One Check Lies
A St. Gallen study tracked four AI engines for six weeks and found the sources they cite overlap only 34–42% between consecutive days. AI visibility is a distribution, not a ranking — and needs 7–8 measurements to pin down.
If you check whether your brand appears in an AI answer today, then check again tomorrow, you may get two different answers — with no change to your site, the news, or the world. That instability is not a glitch. According to a six-week study from the University of St. Gallen, it is the defining property of AI search.
The measurement
Schulte, Bleeker and Kaufmann (2026) queried four engines — ChatGPT, Gemini, Google AI Mode and Perplexity — daily across four campaigns for 45–46 days, then measured how much the set of cited sources overlapped from one observation to the next, using the Jaccard index and Rank-Biased Overlap (RBO).
Sources churn ~60% per day
| Overlap metric | Sources | Brands |
|---|---|---|
| Day-to-day (Jaccard) | 0.34–0.42 | 0.45–0.59 |
| Same-day repeat runs | 0.32–0.43 | 0.41–0.49 |
| Rank overlap (RBO) | 0.21–0.30 | 0.19–0.30 |
A Jaccard of 0.35 means that, on average, 65% of cited sources change between two consecutive days. The critical control: even prompts run twice on the same day, minutes apart, overlapped only 32–43%. That rules out news cycles and index refreshes as the main cause — the volatility is intrinsic to the model’s generation process.
Note the second column: brand mentions are more stable than source URLs. Being named as a brand survives the churn better than being cited as a specific link — a useful asymmetry for anyone tracking presence.
A few domains win most citations
Citations are highly concentrated. The mean Gini coefficient was 0.715 across campaigns and engines — Google AI Mode most concentrated (0.782), Perplexity least (0.671). A handful of authoritative domains capture the bulk of visibility, which compounds the big-brand bias documented elsewhere in the GEO literature.
The rule this forces
Because a single query is a noisy sample of a distribution, the authors ran a bootstrap convergence analysis to ask how many runs you actually need:
| Target | Runs needed |
|---|---|
| Per-brand detection (SE < 0.10) | ≈ 7 runs |
| Source coverage (SE < 0.10) | ≈ 8 runs |
| Stable per-brand rate over time | 2–4 week rolling window |
“Single observations of AI visibility are misleading and risk over- or under-estimating true brand presence.” — Schulte et al. (2026)
So the next time someone reports that a brand “ranks #2 in ChatGPT,” ask how many times they checked. If the answer is once, the number is inside the noise. This is why our methodology treats any single AI-visibility reading as provisional until it’s repeated.
§ Frequently asked questions
Why do AI search citations keep changing? +
Because generative engines are stochastic. Schulte et al. (2026) found cited sources overlap only 34–42% between consecutive days, and 32–43% even when the same prompt is run twice in one day — so most of the churn comes from the model itself, not from news events or index updates.
How many times should I measure AI visibility? +
At least 7 runs per prompt for brand-level visibility and about 8 for source-level coverage, aggregated over a two-to-four-week window. A single observation has a standard error too large to distinguish real visibility from noise.
Technology correspondent · MSc, Computer Science
Marcus Okonkwo covers artificial intelligence and search for DataBackedNews, focused on what the measurements actually show once the marketing is stripped away. He reads the papers so you don't have to — and cites them so you can.