Parse
Work with usPricing
Sign inCheck your brand
Research/The citation cliff: a model upgrade halved the sources behind AI answers

The citation cliff: a model upgrade halved the sources behind AI answers

When ChatGPT's flagship model upgraded, the sources cited per answer fell from 23 to 12. An unchanged control model held steady, isolating the cut to the model, not to anything brands did to their own pages.

By Dimitry Apollonsky · June 25, 2026 · 7 min read

Sources cited per answer
  • Before the upgrade23.4
  • After the upgrade12.1
Mean sources cited per answer across the matched prompt set, meaning the same prompts run under both versions, before and after the upgrade.
▸Contents
  • A model upgrade halved the sources behind each answer
  • On a matched prompt set, 99.7% of answers cited fewer sources
  • An unchanged control model held steady, isolating the cause
  • The control sat where the flagship used to, then the flagship dropped under it
  • Every industry category fell
  • After the upgrade, answers landed in a narrow range, whatever the topic
  • Citation counts went from a wide spread to a near-fixed number
  • The cut was broad and even, not a shift toward big domains
  • A third fewer citations, and a fifth fewer distinct domains
  • What this means for your AI visibility
  • How we measured this
  • Get the data
  • Sources
  • Related research
Contents
  • A model upgrade halved the sources behind each answer
  • On a matched prompt set, 99.7% of answers cited fewer sources
  • An unchanged control model held steady, isolating the cause
  • The control sat where the flagship used to, then the flagship dropped under it
  • Every industry category fell
  • After the upgrade, answers landed in a narrow range, whatever the topic
  • Citation counts went from a wide spread to a near-fixed number
  • The cut was broad and even, not a shift toward big domains
  • A third fewer citations, and a fifth fewer distinct domains
  • What this means for your AI visibility
  • How we measured this
  • Get the data
  • Sources
  • Related research

We compared 423,341 ChatGPT answers from steady windows before and after a flagship model upgrade, with 16,161 prompts run under both versions and an unchanged control model.

23 → 12
Sources cited per answer
Mean, before and after
−48%
Drop on the upgraded flagship
99.7%
Of matched prompts cited fewer sources
+4%
Change on the unchanged control

A model upgrade halved the sources behind each answer

A model upgrade on the vendor's side is a real cause of AI-visibility change, and it can move more than anything a brand does to its own pages. When ChatGPT logoChatGPT's flagship version upgraded in early March 2026, the evidence behind each answer thinned out across the changeover.

Sources cited per answer fell from a mean of 23.4 to 12.1, a 48% cut. The drop was not gradual. It lined up with the model change, not with anything on the sites being cited.

23.4 → 12.1
mean sources cited per answer, before and after
a 48% cut

Takeaway

Before you read a citation drop as a content problem, check whether it lines up with a model change. The cause may be upstream of your site.

On a matched prompt set, 99.7% of answers cited fewer sources

Window averages can hide a shifting mix of questions. To rule that out we held the questions fixed and compared 16,161 prompts that ran under both versions of the flagship.

Of those matched prompts, 16,116 cited fewer sources after the upgrade and only 45 rose. That is 99.7% moving the same direction, with a mean drop of about 11 sources per prompt. This was not noise.

16,161
Prompts run under both versions
99.7%
Cited fewer sources after
−11.3
Mean drop per matched prompt

An unchanged control model held steady, isolating the cause

If the drop were a quirk of measurement, a model that did not change should have fallen alongside the flagship. A smaller, unchanged model ran throughout the same weeks as a control.

The upgraded flagship fell 48% while the control held steady, even ticking up 4% over the same window. The gap between the two is about 12 sources per answer. The cut tracks the model upgrade, not how we measured it.

  • Upgraded flagship−48.1%
  • Unchanged control+4.4%
Change in sources cited per answer over the same window, upgraded flagship versus unchanged control.

Takeaway

When a number changes, look for a control. The control held while the flagship fell, which is how we know the model moved it.

The control sat where the flagship used to, then the flagship dropped under it

Before the upgrade the two models cited almost the same number of sources, 23.4 for the flagship and 23.6 for the control. After, the control still cited 24.6 while the flagship had fallen to 12.1. Same weeks, same measurement, two very different paths.

23.4 → 12.1
Upgraded flagship, before to after
23.6 → 24.6
Unchanged control, before to after
~12
Sources apart after the upgrade

Every industry category fell

The cut was not concentrated in one kind of question. Every industry group in the sample fell, and answers after the upgrade landed at roughly 11 to 13 sources regardless of topic.

The biggest losers were the categories that had been citing the most. Software and artificial intelligence answers, which led before, fell the hardest. The model pulled the top categories down toward the rest.

Mean sources cited per answer by industry group, before and after. Click a column to sort.
Software25.412.5-12.9
Artificial intelligence25.412.7-12.7
Financial services23.512.8-10.7
Commerce and shopping2211.4-10.7
Health care22.211.8-10.4
Community and lifestyle20.411.6-8.8

After the upgrade, answers landed in a narrow range, whatever the topic

Across all 31 industry segments in the sample, the averages after the upgrade landed in a tight 11 to 13 source range. Where a question came from used to shape how much evidence it drew. After the upgrade it barely did.

31 / 31
industry segments that fell, all landing at ~11–13 sources

Citation counts went from a wide spread to a near-fixed number

Before the upgrade, source counts spread widely. Answers ranged from the mid-teens to past 40, with a long tail of deep-citing answers above 35. The spread, measured as standard deviation, was 7.6.

After, the range collapsed. 88% of answers now cite between 10 and 15 sources, and the spread roughly halved to 3.0. Answers stopped occasionally reaching deep and settled into a fixed band.

88%
Of answers now cite 10–15 sources
7.6 → 3.0
Spread in source counts, before to after
47%
Of old answers cited 20–25 sources

The cut was broad and even, not a shift toward big domains

A natural worry is that fewer citations means the model narrowed onto a handful of giant sites. The concentration figures say otherwise.

The share of citations held by the top 20 domains barely moved, 9.6% to 9.9%, and the top 100 share moved just as little. The cut landed roughly proportionally across the whole field rather than favoring the biggest sites.

  • Top 20 share, before9.6%
  • Top 20 share, after9.9%
Share of all citations held by the top 20 domains, before and after the upgrade.

Takeaway

Fewer citations, but the same order. Everyone lost roughly the same share.

A third fewer citations, and a fifth fewer distinct domains

Halving sources per answer shrinks the whole citation footprint. Total citations across the window fell 37% and the number of distinct domains cited fell 23%. The web that AI draws from got both shallower per answer and narrower overall.

−37%
Total citations
−23%
Distinct domains cited
9.6% → 9.9%
Top 20 domain share (held flat)

What this means for your AI visibility

A citation drop can be the model, not you. Before you treat a fall in citations as a content problem, check whether it lines up with a model change. The cause may be upstream of anything on your site.

Fewer citations raises the bar. When each answer cites fewer sources, the contest to be one of them is sharper. Breadth of presence matters more, not less.

Track across model releases. A single visibility number read in isolation will blame the wrong cause when a model shift happens. Watch the trend across releases, not just week to week.

How we measured this

We counted the number of distinct source domains behind each ChatGPT logoChatGPT answer across a steady window before the flagship upgrade and a steady window after, then compared 16,161 prompts that ran under both versions. A smaller, unchanged model ran throughout as a control. The upgrade is anchored on the observed changeover in early March 2026.

This covers one engine and one upgrade event, so it does not predict the direction of future upgrades. Source count measures breadth, not influence per source, so fewer citations does not strictly mean less downstream effect. Figures describe an observed sample of AI answers, not Perplexity, Gemini, or Copilot.

Get the data

Dataset CSVHeadline metrics behind every figure in this report.

Sources

  1. OpenAI ships flagship model updates that change ChatGPT's default behavior, OpenAI · accessed 2026-06-25
  2. AI answer engines cite a variable number of web sources per answer, Search Engine Land · accessed 2026-06-25

Related research

How long an AI citation lasts
Ask AI the same question twice and it rarely cites the same sources. About half of a query's citations vanish by the next run, but a durable core persists for weeks.
AI recommendation movers: how sticky is the top?
Two snapshots of where AI ranks brands, five months apart, on the same set of brands. The top is sticky but not frozen: four in ten of the top-100 brands churned out. A one-time AI-visibility reading is a weak signal.
How AI picks a winner in head-to-head comparisons
When AI compares two brands on more than one thing, it picks a different winner about half the time. There is no single winner, only a winner per axis.

About this research

Dimitry Apollonsky

Founder, Parse

I built Parse to track where AI answers really come from: the sources they cite and the brands they name. DM me on LinkedIn to talk shop.

See the sources behind your brand in AI answers.

Run a free check against live AI answers — no account needed.

Parse

Parse indexes AI recommendations so brands know where they stand.

Products

  • Brands
  • Markets
  • Integrations
  • Work with us
  • Pricing
  • MCP

Resources

  • Research
  • Methodology
  • Blog

© 2026 Parse. All rights reserved.

LegalPrivacy PolicyTerms of Service