ChatGPT's move from gpt-5-2 to gpt-5-3 cut average cited source domains per answer from 23.4 to 12.1 across 16,161 matched prompts, a 48% decline. An unchanged control model held steady, which separates the model change from a collection problem. Citation monitoring should record model versions so teams do not mistake a platform-wide shift for a site-specific loss.
What changed when gpt-5-3 replaced gpt-5-2
In the three weeks before the changeover, ChatGPT's flagship cited a mean of 23.4 source domains per answer (191,224 answers, median 22). In the three weeks after gpt-5-3 stabilized, that mean was 12.1 (232,117 answers, median 12). The same answer that used to reach for two dozen sources now reaches for a dozen.
This is a behavior change, not a quality judgment. gpt-5-3 did not get worse at citing; it got terser. But for a brand, terser is the whole story: every source a model drops is a slot your domain can no longer win. When the slot count falls by half, the average brand's odds of being one of the cited sources fall with it, before anyone touches a page.
- ChatGPT's flagship upgrade (gpt-5-2 to gpt-5-3, early March 2026) cut sources cited per answer from 23.4 to 12.1, a 48% reduction.
- An unchanged gpt-5-mini control held at roughly 24 over the same weeks and pipeline, isolating the cut to the model rather than to measurement.
- On 16,161 prompts run under both models, 99.7% cited fewer sources after the upgrade. Only 45 rose.
- gpt-5-3 cites a near-fixed budget: 88% of answers now cite 10 to 15 sources, regardless of topic.
- The cut was broad, not concentrating. Big domains did not gain share; the long tail was not wiped out.
How we know it was the model, not our measurement
A 48% drop is exactly the kind of number a broken pipeline produces. So we leaned on a control. ChatGPT runs more than one model: alongside the flagship, a lighter gpt-5-mini tier handled a share of the same prompt set throughout. That tier did not change across the window.
If the drop were an artifact of how Parse extracts citations, the control would have dropped too. It did not.
| Model (role) | Sources per answer, Feb 1 to 23 | Sources per answer, Apr 1 to 25 | Change |
|---|---|---|---|
| Flagship (gpt-5-2 to gpt-5-3) | 23.4 | 12.1 | -48% |
| gpt-5-mini (unchanged control) | 23.6 | 24.6 | +4% |
The flagship lost 11.3 sources per answer while the control gained one over the identical period and pipeline. The difference-in-differences is 12.3 sources per answer, and it sits entirely on the model that was upgraded.
Every prompt lost citations, in every category
The window numbers above mix different prompts before and after, so we also ran the strict version: 16,161 prompts executed under both gpt-5-2 and gpt-5-3, compared to themselves. The paired drop was 11.3 sources per answer (from 23.4 to 12.1). Of those 16,161 prompts, 16,116, or 99.7%, cited fewer sources after the upgrade. Forty-five rose. The effect is statistically unambiguous (paired t of about 265; Cohen's d of about 2.1), but the share matters more than the test: this was not a tendency, it was nearly every prompt.
It was also uniform across industries. Every one of the 31 category groups we segment fell, and all of them landed near 11 to 13 sources regardless of topic.
| Category | Before (gpt-5-2) | After (gpt-5-3) | Change |
|---|---|---|---|
| Software | 25.4 | 12.5 | -12.9 |
| Artificial Intelligence | 25.4 | 12.7 | -12.7 |
| Financial Services | 23.5 | 12.8 | -10.7 |
| Health Care | 22.2 | 11.8 | -10.4 |
| Commerce and Shopping | 22.0 | 11.4 | -10.7 |
| Community and Lifestyle | 20.4 | 11.6 | -8.8 |
The categories that cited the most before lost the most. Software and AI prompts started near 25 sources and shed almost 13; lifestyle prompts started near 20 and shed 9. Everything converged on the same narrow band. That convergence is the second finding.
If you want to see how AI engines describe your own brand, run a free brand check — it takes a minute.
gpt-5-3 cites a near-fixed budget of about 12 sources
gpt-5-2 cited a variable number of sources: its per-answer counts spread from 15 to past 40, with 47% landing in the 20 to 25 range and a long tail above 35. gpt-5-3 collapsed that spread. After the upgrade, 88% of all answers cited between 10 and 15 sources (median 12, standard deviation 3.0, down from 7.6).
The model did not just cite less; it cited a consistent amount. gpt-5-3 behaves as if it has a roughly fixed citation budget per answer, applied whether the question is about CRM software or running shoes. For monitoring, that is useful: the metric got less noisy, so a real shift is now easier to see against a tighter baseline. For brands, it means citation slots are now a scarce, predictable quantity, closer to 12 seats per answer than to "as many as the topic warrants."
The cut was broad, not concentrated
A natural worry is that fewer slots favor the giants, that gpt-5-3 kept Wikipedia and the big publishers and dropped everyone else. The data does not support that. The top 20 domains held 9.6% of all citation slots before the upgrade and 9.9% after; the top 100 went from 15.7% to 16.4%. Source diversity barely moved: distinct cited domains fell 23% (360,147 to 278,086) while total citation slots fell 37%.
The decline was roughly proportional across frequently and infrequently cited domains. Large domains did not gain share, and smaller domains did not disappear. Every domain was cited less often because the model returned fewer citation slots per answer.
What a model upgrade means for your AI visibility
The practical lesson is not "gpt-5-3 cites less, so do X." The next upgrade could just as easily cite more, or reshuffle which sources it trusts. The lesson is that a vendor-side model swap is a first-class cause of AI visibility change, as large here as anything you could do to your own content in a quarter, and it leaves no announcement in your analytics.
So three things follow. First, segment AI visibility by model version, not just by platform; a ChatGPT drop that lands on one model and not another is almost certainly an upgrade, not your content. Second, keep a control in view: when one model moves and a comparable one does not, the cause is the model. Third, do not react to a release-week dip as if it were a ranking loss. If you cut citations on a drop that is real or noise, the model getting terser looks identical to your page losing authority until you check the version. Even between identical runs on a stable model, citations churn more than most teams expect (how long an AI citation lasts). Parse tracks AI visibility across ChatGPT and Google AI Overviews over a public index of 4.76 million AI responses, 603,000 brands, and 57 million citations, which is what makes a per-model, controlled comparison like this one possible. You can watch how many sources cite your category, and which ones, in Citations and Movers.
A note on what we did and did not measure
We measured citation breadth: the count of distinct source domains in each answer's citation array, which has above-95% extraction coverage in both windows. We deliberately did not report brand-mention counts for this upgrade.
Brand-name extraction coverage shifted across this same window for every model we track, including the unchanged control, a measurement artifact rather than model behavior. Citation-domain extraction stayed above 95% in both windows, so we report only citation breadth. The model labels are the versions ChatGPT reported inside our pipeline; we anchor the event on the observed changeover (gpt-5-2 last seen March 1, gpt-5-3 from March 6), not a public release date. This is one platform and one upgrade: the finding is that the upgrade moved citations by half, and the takeaway is to measure upgrades, not assume their direction.
Frequently asked questions
Did a ChatGPT update change my citations?
It can, and the change can be large. When ChatGPT's flagship moved from gpt-5-2 to gpt-5-3, sources cited per answer fell 48% across 16,161 matched prompts, and 99.7% of them dropped. If your AI citations fell around a model release, compare the period against a model that did not change. A drop that lands on the upgraded model alone points to the model, not your content.
How much did gpt-5-3 reduce citations compared to gpt-5-2?
From a mean of 23.4 source domains per answer to 12.1, a 48% reduction, measured on 16,161 prompts run under both models. A gpt-5-mini control held at roughly 24 over the same weeks and pipeline, which isolates the drop to the flagship upgrade rather than to measurement.
Does a model upgrade favor big brands over small ones?
Not in this case. The top 20 domains held about 9.6% of citation slots before the upgrade and 9.9% after, essentially flat. The cut fell roughly proportionally across large and small sources, so brands lost citations because there were fewer slots overall, not because the model started preferring incumbents.
How should I track AI visibility across model releases?
Segment by model version, not just by platform, so a change on one model is visible against the others. Keep at least one comparable model as a control, and treat release-week dips as version events to verify before acting. See why AI sometimes stops recommending a brand and the weekly AI visibility review for a routine that catches these.