The citation cliff: a model upgrade halved the sources behind AI answers
When ChatGPT's flagship model upgraded, the sources cited per answer fell from 23 to 12. An unchanged control model held steady, isolating the cut to the model, not to anything brands did to their own pages.
By Dimitry Apollonsky · June 25, 2026 · 7 min read
Contents
- A model upgrade halved the sources behind each answer
- On a matched prompt set, 99.7% of answers cited fewer sources
- An unchanged control model held steady, isolating the cause
- The control sat where the flagship used to, then the flagship dropped under it
- Every industry category fell
- After the upgrade, answers landed in a narrow range, whatever the topic
- Citation counts went from a wide spread to a near-fixed number
- The cut was broad and even, not a shift toward big domains
- A third fewer citations, and a fifth fewer distinct domains
- What this means for your AI visibility
- How we measured this
- Get the data
- Sources
- Related research
We compared 423,341 ChatGPT answers from steady windows before and after a flagship model upgrade, with 16,161 prompts run under both versions and an unchanged control model.
A model upgrade halved the sources behind each answer
A model upgrade on the vendor's side is a real cause of AI-visibility change, and it can move more than anything a brand does to its own pages. When ChatGPT's flagship version upgraded in early March 2026, the evidence behind each answer thinned out across the changeover.
Sources cited per answer fell from a mean of 23.4 to 12.1, a 48% cut. The drop was not gradual. It lined up with the model change, not with anything on the sites being cited.
Takeaway
On a matched prompt set, 99.7% of answers cited fewer sources
Window averages can hide a shifting mix of questions. To rule that out we held the questions fixed and compared 16,161 prompts that ran under both versions of the flagship.
Of those matched prompts, 16,116 cited fewer sources after the upgrade and only 45 rose. That is 99.7% moving the same direction, with a mean drop of about 11 sources per prompt. This was not noise.
An unchanged control model held steady, isolating the cause
If the drop were a quirk of measurement, a model that did not change should have fallen alongside the flagship. A smaller, unchanged model ran throughout the same weeks as a control.
The upgraded flagship fell 48% while the control held steady, even ticking up 4% over the same window. The gap between the two is about 12 sources per answer. The cut tracks the model upgrade, not how we measured it.
Takeaway
The control sat where the flagship used to, then the flagship dropped under it
Before the upgrade the two models cited almost the same number of sources, 23.4 for the flagship and 23.6 for the control. After, the control still cited 24.6 while the flagship had fallen to 12.1. Same weeks, same measurement, two very different paths.
Every industry category fell
The cut was not concentrated in one kind of question. Every industry group in the sample fell, and answers after the upgrade landed at roughly 11 to 13 sources regardless of topic.
The biggest losers were the categories that had been citing the most. Software and artificial intelligence answers, which led before, fell the hardest. The model pulled the top categories down toward the rest.
| Software | 25.4 | 12.5 | -12.9 |
| Artificial intelligence | 25.4 | 12.7 | -12.7 |
| Financial services | 23.5 | 12.8 | -10.7 |
| Commerce and shopping | 22 | 11.4 | -10.7 |
| Health care | 22.2 | 11.8 | -10.4 |
| Community and lifestyle | 20.4 | 11.6 | -8.8 |
After the upgrade, answers landed in a narrow range, whatever the topic
Across all 31 industry segments in the sample, the averages after the upgrade landed in a tight 11 to 13 source range. Where a question came from used to shape how much evidence it drew. After the upgrade it barely did.
Citation counts went from a wide spread to a near-fixed number
Before the upgrade, source counts spread widely. Answers ranged from the mid-teens to past 40, with a long tail of deep-citing answers above 35. The spread, measured as standard deviation, was 7.6.
After, the range collapsed. 88% of answers now cite between 10 and 15 sources, and the spread roughly halved to 3.0. Answers stopped occasionally reaching deep and settled into a fixed band.
The cut was broad and even, not a shift toward big domains
A natural worry is that fewer citations means the model narrowed onto a handful of giant sites. The concentration figures say otherwise.
The share of citations held by the top 20 domains barely moved, 9.6% to 9.9%, and the top 100 share moved just as little. The cut landed roughly proportionally across the whole field rather than favoring the biggest sites.
Takeaway
A third fewer citations, and a fifth fewer distinct domains
Halving sources per answer shrinks the whole citation footprint. Total citations across the window fell 37% and the number of distinct domains cited fell 23%. The web that AI draws from got both shallower per answer and narrower overall.
What this means for your AI visibility
A citation drop can be the model, not you. Before you treat a fall in citations as a content problem, check whether it lines up with a model change. The cause may be upstream of anything on your site.
Fewer citations raises the bar. When each answer cites fewer sources, the contest to be one of them is sharper. Breadth of presence matters more, not less.
Track across model releases. A single visibility number read in isolation will blame the wrong cause when a model shift happens. Watch the trend across releases, not just week to week.
How we measured this
We counted the number of distinct source domains behind each ChatGPT answer across a steady window before the flagship upgrade and a steady window after, then compared 16,161 prompts that ran under both versions. A smaller, unchanged model ran throughout as a control. The upgrade is anchored on the observed changeover in early March 2026.
This covers one engine and one upgrade event, so it does not predict the direction of future upgrades. Source count measures breadth, not influence per source, so fewer citations does not strictly mean less downstream effect. Figures describe an observed sample of AI answers, not Perplexity, Gemini, or Copilot.
Get the data
Sources
- OpenAI ships flagship model updates that change ChatGPT's default behavior, OpenAI · accessed 2026-06-25
- AI answer engines cite a variable number of web sources per answer, Search Engine Land · accessed 2026-06-25