LIVE
UTC
DataBackedNews
Technology ◆ Reference hub

Generative Engine Optimization: What Four 2026 Studies Actually Found

A synthesis of four peer-reviewed studies on how AI search decides what to cite — covering visibility instability, the discovery gap, content structure, and the earned-media bias. Every figure sourced and dated.

Marcus Okonkwo · Technology correspondent · ·peer-reviewed

Every marketing blog has a theory about how to rank in ChatGPT. Very few cite data. This is a synthesis of four studies published between late 2025 and mid-2026 that actually measured how generative engines decide what to cite — and several of their findings contradict the standard advice.

We read the primary sources and pulled the headline numbers into one place. Each is linked in the source list.

The one-paragraph summary

On-page optimization makes a page citeable but does not make it discoverable. Discovery is driven off-page — by referring domains and community presence — while the biggest on-page lever is structure, not word choice. And whatever you publish competes inside a system that is both unstable (citations churn day to day) and biased toward earned, big-brand media. Optimize the page, yes, but do not expect the page alone to win the query.

Finding 1 — On-page GEO score does not predict discovery

The most uncomfortable result comes from Sharma’s Discovery Gap study (IIT Patna, 2025), which tested 112 products across 2,240 queries on ChatGPT and Perplexity. It separated direct queries (“What is [product]?”) from discovery queries (“What are the best tools in this category?”).

Query typeChatGPTPerplexity
Direct (by name)99.4%94.3%
Discovery (by category)3.32%8.29%
Visibility gap30 : 111 : 1

A composite on-page GEO score — statistics, citations, technical terms, structured data — showed no significant correlation with discovery (r = -0.108 on ChatGPT, r = -0.102 on Perplexity; both non-significant). What did correlate, for the web-search engine Perplexity, was off-page: referring domains (r = +0.319), and cleaned Reddit presence (r = +0.395). The author’s framing is blunt: GEO is a multiplier, not a catalyst — “you can’t multiply zero.” Read the full breakdown of the discovery gap.

Finding 2 — Visibility is an unstable distribution

Even once you are cited, you can’t trust a single measurement. Schulte, Bleeker and Kaufmann (Univ. St. Gallen, 2026) tracked four campaigns across four engines and found the set of cited sources is remarkably volatile.

Stability metricSourcesBrands
Day-to-day overlap (Jaccard)0.34–0.420.45–0.59
Same-day repeated runs0.32–0.430.41–0.49
Citation concentration (Gini)~0.715

In plain terms: 58–66% of cited sources change from one day to the next, and they churn even when the identical prompt is run twice minutes apart. Brand mentions are more stable than source URLs. The practical implication is a measurement rule, not a content tactic: you need 7–8 runs over two to four weeks to estimate visibility with any confidence. Why AI citations are so unstable →

Finding 3 — Structure alone lifts citations 17.3%

Yu, MuFeng, Ding and Sato (2026) did something clean: they held meaning constant and changed only structure, then measured citation rate across six engines.

Structural levelContribution to lift
Macro (heading hierarchy, flow)44.9%
Meso (chunking, tables, lists)39.7%
Micro (emphasis, keyword placement)15.4%
Total citation lift+17.3%

The actionable specifics: keep paragraphs to 150–300 words (beyond 300, model attention degrades ~31%), use lists and tables (they showed 43% higher extraction), and maintain a clean heading depth of three to five levels. This is the cheapest GEO win available, and it is purely mechanical. The structure finding, in full →

Finding 4 — AI over-cites earned, big-brand media

Chen, Wang, Chen and Koudas (Univ. of Toronto, 2025) compared AI-search sourcing to Google across regions, languages and verticals. The consistent pattern: AI answers lean hard on earned media — third-party reviews, publishers, institutions.

Source typeAI searchGoogle
Earned (third-party)72–92%~35–50%
Brand-ownedlowbalanced
Social / communityoften ~0%substantial

They also documented a big-brand bias: major brands accounted for 56–68% of mentions in unbranded “best/most popular” prompts, niche brands just 6–12%. For a new or niche entity, the lever that matters is earning third-party coverage from higher-authority domains — not polishing your own page.

What this means if you publish

  1. Fix structure first — it’s the one on-page change with a measured effect (+17.3%).
  2. Don’t expect on-page work to get you discovered — that’s an off-page problem (referring domains, community).
  3. Earn third-party media — it’s the category AI over-weights, and the only real path past big-brand bias.
  4. Measure over weeks, not once — a single check is inside the noise band.
  5. Foundations before flourish — as Aggarwal et al. (2024) showed, content levers help once you are in the retrieval set; getting into that set is a separate, off-page job.

None of the four studies contradicts the others once you separate the two questions they answer: what gets retrieved (off-page, unstable, earned-media-biased) versus what gets cited once retrieved (structure, statistics, sourcing). Optimize for both, and measure honestly.

§ Frequently asked questions

Does on-page GEO optimization make you appear in AI answers? +

Not on its own. A 112-product study (Sharma, 2025) found the on-page GEO score had no statistically significant correlation with whether a product was discovered in ChatGPT or Perplexity (r≈-0.10). On-page optimization raises citation quality once you are already retrieved; it does not get you retrieved. Discovery was predicted instead by referring domains and community presence.

Why do AI citations change every time I check? +

Because AI-search visibility is a distribution, not a fixed ranking. Schulte et al. (2026) found the set of cited sources overlaps only 34–42% from one day to the next, and 32–43% even for the same prompt run twice in one day. Reliable measurement needs 7–8 runs over a two-to-four-week window.

What is the single highest-leverage on-page change? +

Structure. Yu et al. (2026) isolated content structure from wording and measured a +17.3% citation lift, with document-level heading hierarchy contributing about 45% of the gain and tables/chunking another 40%. Clean headings, 150–300 word chunks, and data tables are the cheapest wins.

What kind of sources do AI engines prefer? +

Earned media. Chen et al. (2025) found AI answers draw 72–92% of citations from third-party, authoritative sources (reviews, publishers, institutions), far more than Google's balanced mix. Brand-owned and social sources are systematically down-weighted, and major brands dominate over niche players.

#generative engine optimization#AI search#LLM citation#GEO#retrieval
M
Marcus Okonkwo

Technology correspondent · MSc, Computer Science

Marcus Okonkwo covers artificial intelligence and search for DataBackedNews, focused on what the measurements actually show once the marketing is stripped away. He reads the papers so you don't have to — and cites them so you can.

§ More from the Technology desk

Technology 2 SRC

AI Search Citations Change 60% Day to Day — Why One Check Lies

A St. Gallen study tracked four AI engines for six weeks and found the sources they cite overlap only 34–42% between consecutive days. AI visibility is a distribution, not a ranking — and needs 7–8 measurements to pin down.

Marcus Okonkwo