Ask ChatGPT the same buying question twice and you get two different source lists. Across more than 16,000 questions we re-ran about 22 times each, two answers to the same question shared only about a fifth of their cited sources. The churn is worst in software, AI, and commerce, and it shows up within a single week, so it is mostly model noise, not the web changing. A one-time AI visibility check is a snapshot of a moving target, not a reading you can bank.
Two answers to the same question share only a fifth of their sources
We took every prompt Parse tracks on ChatGPT, kept the ones run at least 12 times in a 30-day window, and measured how much each pair of repeat answers overlapped in the sources they cited. The average overlap was 21%. Put plainly: re-run the same question and roughly four out of every five sources in the answer change. Google AI Overviews was steadier but not stable, at 31% overlap, so even Google's more deterministic surface swaps out about two-thirds of its sources between repeat answers. This is not a tail effect from a few weird prompts. It held across 16,143 ChatGPT prompts and 15,805 Google AI Overviews prompts spanning every major industry. The cited-source list behind an AI answer is a sample, not a fixed fact, and most monitoring treats it as fixed.
- Two ChatGPT answers to the same question shared only ~21% of their cited sources on average; Google AI Overviews shared ~31% (Parse first-party data, 16,000+ prompts re-run ~22 times each, March-April 2026).
- A typical question had just 1-2 "anchor" sources that appeared in at least 80% of its repeat answers, out of roughly 80 distinct domains cited across the runs.
- Volatility was worst in commerce, software, AI, and IT (about 81% of sources changed between repeat answers) and lowest in travel, media, and financial services (about 75-77%).
- ChatGPT was roughly 1.5× more volatile than Google AI Overviews in every industry we measured.
- The churn appears within a single week (only 27% source overlap on ChatGPT), so it is mostly run-to-run model noise, not the source landscape changing.
How we measured it
We read the citation arrays Parse stores for every tracked AI answer. For each prompt run at least 12 times between March 26 and April 25, 2026 (about 22 runs per prompt on average), we compared the set of domains cited in each answer against the others for the same question, and measured the share of sources any two repeat answers held in common. That overlap, averaged across every pair, is our stability number; one minus it is volatility. We ran the same calculation on ChatGPT (355,371 answers) and Google AI Overviews (338,138 answers), then split prompts by their primary industry. Two caveats shape the read. First, this blends two things: the model's run-to-run randomness and genuine change in the underlying web over the month. The week-level check below shows randomness dominates. Second, we measure sources at the domain level, and a domain being cited is not the model endorsing it. The headline, that the source list moves a lot, survives both.
A typical question has one or two anchor sources and a rotating cast of eighty
The volatility has a clear structure. Across about 22 repeat answers, a single ChatGPT question pulled from a pool of roughly 80 distinct domains, but only 1 to 2 of them showed up in at least 80% of the answers. Those one or two "anchor" sources accounted for just 22% of the citation volume; the other ~78% came from a large rotating cast that drifted in and out run to run. Google AI Overviews leaned harder on its anchors, which carried 33% of its citations, the mechanical reason it looks steadier. For a brand, this is the part that matters: being cited once tells you almost nothing. The sources that move your visibility are the anchors, the handful a model returns to nearly every time. Everything else is a coin flip you will win sometimes and lose sometimes, and reading a single answer cannot tell the two apart. The flip side of which sources anchor is how long they stay anchored: how long an AI citation lasts finds the same split over time, a durable core that persists for weeks against a long tail that appears once and never returns.
If you want to know when AI changes its answer about your brand, start with a free brand check — it takes a minute.
Which industries have the most volatile AI citations?
Volatility is not evenly spread. We ranked every industry with at least 150 tracked prompts by how much two repeat ChatGPT answers overlapped. Tech-heavy categories churn the most; categories with stable, authoritative reference sources churn the least.
| Industry | Sources two repeat answers share | Anchor sources per question |
|---|---|---|
| Commerce and Shopping | 18.9% | 1.0 |
| Software | 19.3% | 1.1 |
| Artificial Intelligence | 19.7% | 1.2 |
| Information Technology | 19.9% | 1.3 |
| Internet Services | 20.2% | 1.4 |
| Data and Analytics | 20.6% | 1.3 |
| Consumer Goods | 20.9% | 1.5 |
| Health Care | 21.7% | 1.6 |
| Consumer Electronics | 22.8% | 1.7 |
| Financial Services | 22.9% | 1.7 |
| Media and Entertainment | 23.2% | 1.8 |
| Travel and Tourism | 24.6% | 1.9 |
The gap between the most and least volatile category is real but bounded: every industry sat between 75% and 81% source churn, so no category is stable in absolute terms. Software, AI, and commerce are the hardest to read because their source pools are large, fast-moving, and thin on durable anchors. Travel, media, and financial services hold steadier because answers lean on a small set of recurring authorities (booking platforms, established publishers, regulated finance sites) that the model returns to. If you sell software or run ecommerce, assume your AI visibility signal is noisier than average and size your sample accordingly.
ChatGPT is noisier than Google AI Overviews
The surface you measure on changes the noise floor. ChatGPT swapped about 79% of its sources between repeat answers; Google AI Overviews swapped about 69%. The gap held in every industry: ChatGPT was the more volatile surface for software, finance, health care, travel, and all the rest, by a consistent margin. The reason is the anchor structure above: Google AI Overviews concentrates a third of its citations on a few recurring sources, while ChatGPT spreads attention across a wider, looser pool. For measurement, this has a direct consequence. A given number of checks buys you a more reliable reading on Google AI Overviews than on ChatGPT, and a week-to-week wobble on ChatGPT is more likely to be noise. If you track both surfaces against one combined score, the ChatGPT noise is quietly inflating your error bars, which is one of several reasons a single AI visibility score is misleading.
The churn is noise, not news
A fair objection: maybe the source list changes because the web changes, with new articles, new competitors, and fresh coverage. Over a month, some of that is real. But when we narrowed the window to a single week, the overlap barely moved up: two ChatGPT answers to the same question still shared only 27% of their sources, and Google AI Overviews 37%. Seven days is not enough time for the source landscape to turn over by two-thirds, so most of what we are seeing is the model sampling differently from the same underlying pool, not the pool itself changing. A genuine shift in the corpus is a different event, the kind Parse measured when a flagship model upgrade halved the sources cited per answer. That distinction decides how you read your dashboard. A citation that appears Monday and is gone Wednesday usually has not been "lost"; it was a low-frequency source that was never a reliable part of the answer. The real signal is the rate at which a source shows up across many runs, and you cannot see a rate in a single observation. This is the same trap behind most false alarms when teams ask whether an AI visibility drop is real or noise.
What this means for measuring your AI visibility
Stop reading one-shot AI answers as ground truth, and start measuring presence as a rate. Three rules follow from the data. First, sample repeatedly before you conclude anything: a single check on a volatile category like software can miss or invent a citation that a ten-run average would smooth out. Second, watch the anchors, not the long tail: the one or two sources a model cites for your question nearly every time are where durable visibility lives, and earning a place among them is worth more than a dozen sources that flicker. Third, size your alarms to the surface and the category: a five-point weekly move on ChatGPT in commerce is well inside the noise band, while the same move on Google AI Overviews in financial services is more likely to mean something. The right unit is share of model across many runs, not appearance in one answer, the idea behind share of model as a KPI.
How Parse measures through the noise
Parse tracks AI visibility across ChatGPT and Google AI Overviews, covering a public index of more than 4.7 million AI responses, 603,000 brands, and 57 million citations. The reason the index is built on repeated runs rather than spot checks is exactly the volatility in this study: a single answer is a noisy sample, so Parse re-runs each tracked prompt and reports how often your brand and the sources around it appear, not whether they showed up once. Parse's Citations view ranks the sources cited for your prompt set by how frequently they recur, which separates the durable anchors from the rotating cast and shows where competitors hold anchor positions you do not. That frequency-based reading is what turns a volatile raw signal into something you can plan a budget against. Track which sources cite your brand over time.
How volatile are AI citations?
Very. In Parse's first-party data, two ChatGPT answers to the same question shared only about 21% of their cited sources on average, meaning roughly four in five sources changed between repeat answers. Google AI Overviews was steadier at about 31% overlap. The volatility held across more than 16,000 prompts and every major industry, so a single AI answer is best treated as one noisy sample, not a fixed result.
Which industries have the most volatile AI citations?
Commerce, software, artificial intelligence, and IT were the most volatile on ChatGPT, with about 81% of sources changing between repeat answers. Travel, media, and financial services were the steadiest, at roughly 75-77% churn, because their answers lean on a small set of recurring authorities. Every industry was volatile in absolute terms; the differences are in degree, not kind.
Why do AI answers cite different sources each time?
Mostly because the model samples differently from a large pool of candidate sources each run, not because the web changed. A typical question drew on about 80 distinct domains across repeat answers but only 1-2 appeared in most of them. When Parse narrowed the window to a single week, source overlap barely rose, which shows the churn is run-to-run model noise rather than genuine change in the source landscape.
Is ChatGPT or Google AI Overviews more consistent in its citations?
Google AI Overviews. It shared about 31% of sources between repeat answers versus ChatGPT's 21%, and it was the steadier surface in every industry Parse measured. The reason is that Google AI Overviews concentrates about a third of its citations on a few recurring anchor sources, while ChatGPT spreads citations across a wider, looser pool. A fixed number of checks buys a more reliable reading on Google AI Overviews.
How many times should I check an AI answer before trusting it?
More than once. Because a single answer reflects only about a fifth of the sources that could appear, one check can easily miss a source that shows up most of the time or capture one that rarely does. Track presence as a rate across many runs and weight your attention toward the anchor sources a model returns to nearly every time. A five-point weekly move on a volatile category like software is usually noise, not a real change.