Generative Engine Optimization: What Four 2026 Studies Actually Found
A synthesis of four peer-reviewed studies on how AI search decides what to cite — covering visibility instability, the discovery gap, content structure, and the earned-media bias. Every figure sourced and dated.
Every marketing blog has a theory about how to rank in ChatGPT. Very few cite data. This is a synthesis of four studies published between late 2025 and mid-2026 that actually measured how generative engines decide what to cite — and several of their findings contradict the standard advice.
We read the primary sources and pulled the headline numbers into one place. Each is linked in the source list.
The one-paragraph summary
On-page optimization makes a page citeable but does not make it discoverable. Discovery is driven off-page — by referring domains and community presence — while the biggest on-page lever is structure, not word choice. And whatever you publish competes inside a system that is both unstable (citations churn day to day) and biased toward earned, big-brand media. Optimize the page, yes, but do not expect the page alone to win the query.
Finding 1 — On-page GEO score does not predict discovery
The most uncomfortable result comes from Sharma’s Discovery Gap study (IIT Patna, 2025), which tested 112 products across 2,240 queries on ChatGPT and Perplexity. It separated direct queries (“What is [product]?”) from discovery queries (“What are the best tools in this category?”).
| Query type | ChatGPT | Perplexity |
|---|---|---|
| Direct (by name) | 99.4% | 94.3% |
| Discovery (by category) | 3.32% | 8.29% |
| Visibility gap | 30 : 1 | 11 : 1 |
A composite on-page GEO score — statistics, citations, technical terms, structured data — showed no significant correlation with discovery (r = -0.108 on ChatGPT, r = -0.102 on Perplexity; both non-significant). What did correlate, for the web-search engine Perplexity, was off-page: referring domains (r = +0.319), and cleaned Reddit presence (r = +0.395). The author’s framing is blunt: GEO is a multiplier, not a catalyst — “you can’t multiply zero.” Read the full breakdown of the discovery gap.
Finding 2 — Visibility is an unstable distribution
Even once you are cited, you can’t trust a single measurement. Schulte, Bleeker and Kaufmann (Univ. St. Gallen, 2026) tracked four campaigns across four engines and found the set of cited sources is remarkably volatile.
| Stability metric | Sources | Brands |
|---|---|---|
| Day-to-day overlap (Jaccard) | 0.34–0.42 | 0.45–0.59 |
| Same-day repeated runs | 0.32–0.43 | 0.41–0.49 |
| Citation concentration (Gini) | ~0.715 | — |
In plain terms: 58–66% of cited sources change from one day to the next, and they churn even when the identical prompt is run twice minutes apart. Brand mentions are more stable than source URLs. The practical implication is a measurement rule, not a content tactic: you need 7–8 runs over two to four weeks to estimate visibility with any confidence. Why AI citations are so unstable →
Finding 3 — Structure alone lifts citations 17.3%
Yu, MuFeng, Ding and Sato (2026) did something clean: they held meaning constant and changed only structure, then measured citation rate across six engines.
| Structural level | Contribution to lift |
|---|---|
| Macro (heading hierarchy, flow) | 44.9% |
| Meso (chunking, tables, lists) | 39.7% |
| Micro (emphasis, keyword placement) | 15.4% |
| Total citation lift | +17.3% |
The actionable specifics: keep paragraphs to 150–300 words (beyond 300, model attention degrades ~31%), use lists and tables (they showed 43% higher extraction), and maintain a clean heading depth of three to five levels. This is the cheapest GEO win available, and it is purely mechanical. The structure finding, in full →
Finding 4 — AI over-cites earned, big-brand media
Chen, Wang, Chen and Koudas (Univ. of Toronto, 2025) compared AI-search sourcing to Google across regions, languages and verticals. The consistent pattern: AI answers lean hard on earned media — third-party reviews, publishers, institutions.
| Source type | AI search | |
|---|---|---|
| Earned (third-party) | 72–92% | ~35–50% |
| Brand-owned | low | balanced |
| Social / community | often ~0% | substantial |
They also documented a big-brand bias: major brands accounted for 56–68% of mentions in unbranded “best/most popular” prompts, niche brands just 6–12%. For a new or niche entity, the lever that matters is earning third-party coverage from higher-authority domains — not polishing your own page.
What this means if you publish
- Fix structure first — it’s the one on-page change with a measured effect (+17.3%).
- Don’t expect on-page work to get you discovered — that’s an off-page problem (referring domains, community).
- Earn third-party media — it’s the category AI over-weights, and the only real path past big-brand bias.
- Measure over weeks, not once — a single check is inside the noise band.
- Foundations before flourish — as Aggarwal et al. (2024) showed, content levers help once you are in the retrieval set; getting into that set is a separate, off-page job.
None of the four studies contradicts the others once you separate the two questions they answer: what gets retrieved (off-page, unstable, earned-media-biased) versus what gets cited once retrieved (structure, statistics, sourcing). Optimize for both, and measure honestly.
§ Frequently asked questions
Does on-page GEO optimization make you appear in AI answers? +
Not on its own. A 112-product study (Sharma, 2025) found the on-page GEO score had no statistically significant correlation with whether a product was discovered in ChatGPT or Perplexity (r≈-0.10). On-page optimization raises citation quality once you are already retrieved; it does not get you retrieved. Discovery was predicted instead by referring domains and community presence.
Why do AI citations change every time I check? +
Because AI-search visibility is a distribution, not a fixed ranking. Schulte et al. (2026) found the set of cited sources overlaps only 34–42% from one day to the next, and 32–43% even for the same prompt run twice in one day. Reliable measurement needs 7–8 runs over a two-to-four-week window.
What is the single highest-leverage on-page change? +
Structure. Yu et al. (2026) isolated content structure from wording and measured a +17.3% citation lift, with document-level heading hierarchy contributing about 45% of the gain and tables/chunking another 40%. Clean headings, 150–300 word chunks, and data tables are the cheapest wins.
What kind of sources do AI engines prefer? +
Earned media. Chen et al. (2025) found AI answers draw 72–92% of citations from third-party, authoritative sources (reviews, publishers, institutions), far more than Google's balanced mix. Brand-owned and social sources are systematically down-weighted, and major brands dominate over niche players.
Technology correspondent · MSc, Computer Science
Marcus Okonkwo covers artificial intelligence and search for DataBackedNews, focused on what the measurements actually show once the marketing is stripped away. He reads the papers so you don't have to — and cites them so you can.