Do ChatGPT model updates change brand recommendations?
We compared the same prompts answered before and after two ChatGPT model updates. Each update changed the top brand in about 10 more answers out of 100 than asking the same model again. Google AI Mode, which did not change, held steady on the same day.
Published Updated
A ChatGPT model update changed the top brand in 10 more answers out of 100
Answers that kept the same top brand
- Same model, Feb to Mar48.2%
- Across gpt-5-2 to gpt-5-338.6%
- Same model, Jul to Aug (average)50.0%
- Across gpt-5-5 to gpt-5-642.6%
Reviewed by
Dimitry ApollonskyFounder, Parse
In search and growth marketing since 2015. Reviews every Parse research report.
Reviewed
A ChatGPT model update changed the top brand in 10 more answers out of 100
We watched the same prompts get answered before and after two ChatGPT model updates. The top brand is the brand an answer names first. At each update, the share of answers that kept the same top brand fell by about 10 percentage points more than it does when the same model is asked again.
At the gpt-5-2 to gpt-5-3 update the drop was 9.6 points against an unchanged control model. At the gpt-5-5 to gpt-5-6 update it was 10.0 points against the same model re-asked, after taking out the small move Google AI Mode made on the same date. Two updates, five months apart, gave nearly the same number.
Takeaway
Both updates kept the top brand less often than the same model asked again
Each row below is a set of answer pairs: the same prompt on the same engine, asked twice. A pair crosses an update when the first answer came from the old model and the second from the new one.
Across gpt-5-2 to gpt-5-3, 38.6% of pairs kept the same top brand. An unchanged smaller model, gpt-5-mini, answered the same prompts on the same days and kept it in 48.2% of pairs, after matching the time between asks. Across gpt-5-5 to gpt-5-6, 42.6% kept it. The same models re-asked a few days apart kept it in 52.8% (gpt-5-5) and 47.1% (gpt-5-6) of pairs. Each rate is accurate to within about 1 point at 95% confidence.
Both updates kept the top brand less often than the same model asked again
| 1 | ChatGPT, gpt-5-2 then gpt-5-3 | 14,044 | 38.6 | 0.8 |
| 2 | ChatGPT, gpt-5-mini both times (same dates) | 14,643 | 47.8 | 0.8 |
| 3 | ChatGPT, gpt-5-5 both times | 46,400 | 52.8 | 0.5 |
| 4 | ChatGPT, gpt-5-5 then gpt-5-6 | 15,494 | 42.6 | 0.8 |
| 5 | ChatGPT, gpt-5-6 both times | 51,359 | 47.1 | 0.4 |
| 6 | Google AI Mode, same dates as row 4 | 15,643 | 50.5 | 0.8 |
Without an update, the top brand held about half the time
AI answers change even when the model stays the same. We measured that normal variation before looking at the effect of an update.
Asked again about four days later by the same model, ChatGPT Search kept the top brand in 52.8% of pairs on gpt-5-5 and 47.1% on gpt-5-6. Google AI Mode kept it in 49.2% of pairs over the same weeks. This fits outside work: SparkToro found that repeated AI prompts almost never return the same brand list twice. An update adds change on top of this floor. It does not replace it.
- ChatGPT Search on gpt-5-5, re-asked
- 52.8%ChatGPT Search on gpt-5-5, re-asked
- ChatGPT Search on gpt-5-6, re-asked
- 47.1%ChatGPT Search on gpt-5-6, re-asked
- Google AI Mode, same weeks
- 49.2%Google AI Mode, same weeks
Takeaway
Google AI Mode did not drop on the day ChatGPT updated
A change in how answers are collected would hit both engines on the same day. A model update hits only the engine that changed.
On August 8, 2026, the first day of gpt-5-6, the share of ChatGPT Search answers keeping the previous answer's top brand fell from 51.6% to 42.3%. Google AI Mode held at 51.0%. Over all pairs that crossed that date, Google AI Mode kept the top brand 2.7 points more often than its own average, while ChatGPT Search kept it 7.3 points less often.
Google AI Mode did not drop on the day ChatGPT updated
Answers keeping the previous answer's top brand, by day
- ChatGPT Search
- Google AI Mode
Explore the data
| 2026-07-29 | 54.4 | 55 |
| 2026-07-30 | 54.6 | 53.5 |
| 2026-07-31 | 52.6 | 53.7 |
| 2026-08-01 | 53.3 | 41 |
| 2026-08-02 | 52.9 | 44.4 |
| 2026-08-03 | 53.4 | 43.1 |
| 2026-08-04 | 51.2 | 42.1 |
| 2026-08-05 | 52.1 | 53.5 |
| 2026-08-06 | 53.2 | 51.8 |
| 2026-08-07 | 51.6 | 51.6 |
| 2026-08-08 | 42.3 | 51 |
| 2026-08-09 | 44.3 | 52.4 |
| 2026-08-10 | 46 | 51.1 |
| 2026-08-11 | 40.3 | 45.4 |
| 2026-08-12 | 47.8 | 48.7 |
| 2026-08-13 | 45.2 | 49.5 |
| 2026-08-14 | 46.7 | 48.3 |
| 2026-08-15 | 45.8 | 48.4 |
| 2026-08-16 | 46.4 | 48.3 |
| 2026-08-17 | 46.1 | 50.2 |
| 2026-08-18 | 49.3 | 51.7 |
| 2026-08-19 | 47 | 50.2 |
| 2026-08-20 | 47.8 | 45.9 |
The brands named in both answers fell by a sixth to a fifth
Shortlist overlap is the share of brands that appear in both answers of a pair, out of all brands named in either. It measures the whole list, not just the first brand.
Across gpt-5-2 to gpt-5-3, shortlist overlap was 30.9%, against 38.9% for the unchanged control. That is 20.5% lower. Across gpt-5-5 to gpt-5-6 it was 34.1%, against an average of 39.8% for the same model re-asked. That is 14.5% lower. Google AI Mode, over the same dates, rose to 41.4% from its 38.4% average.
The brands named in both answers fell by a sixth to a fifth
- Unchanged control, Feb to Mar38.9%
- Across gpt-5-2 to gpt-5-330.9%
- Same model re-asked, Jul to Aug39.8%
- Across gpt-5-5 to gpt-5-634.1%
- Google AI Mode, same dates41.4%
Both updates named fewer brands and wrote shorter answers
On the same prompts, gpt-5-3 named 7.71 brands per answer where gpt-5-2 had named 8.67, an 11% cut. The unchanged control stayed at 8.68. gpt-5-6 named 4.57 brands where gpt-5-5 had named 5.47, a 16% cut. Google AI Mode moved 1% over the same dates.
Answers also got shorter. Text length fell 19% at the first update, while the control moved less than 1%. It fell 5% at the second update, while Google AI Mode moved 1%. Fewer brands per answer means fewer places for any one brand to appear.
- Brands per answer, gpt-5-2 to gpt-5-3
- −11%Brands per answer, gpt-5-2 to gpt-5-38.67 to 7.71; control flat at 8.68
- Brands per answer, gpt-5-5 to gpt-5-6
- −16%Brands per answer, gpt-5-5 to gpt-5-65.47 to 4.57; Google AI Mode −1%
- Answer length, gpt-5-2 to gpt-5-3
- −19%Answer length, gpt-5-2 to gpt-5-3
- Answer length, gpt-5-5 to gpt-5-6
- −5%Answer length, gpt-5-5 to gpt-5-6
gpt-5-6 changed the brands but not the sources or the searches
The gpt-5-3 update also cut the sources cited per answer, which our citation cliff study covers. The gpt-5-6 update did not. On the same prompts, ChatGPT Search cited 4.52 sources per answer before and 4.54 after. It ran 1.01 web searches per answer both times.
So the second update changed which brands the model chose from roughly the same evidence. A brand can lose its place in the answer without losing a single citation.
- Sources cited per answer
- 4.52 → 4.54Sources cited per answer
- Web searches per answer
- 1.01 → 1.01Web searches per answer
- Brands named per answer
- 5.47 → 4.57Brands named per answer
Takeaway
gpt-5-3 named large platforms more and smaller software tools less
The table shows the brands whose appearances moved most across the gpt-5-2 to gpt-5-3 update, after subtracting what the same brand did on the unchanged control over the same dates. Change is counted per 1,000 answer pairs.
Amazon gained the most, adding 12.3 appearances per 1,000 pairs across 401 markets. OpenAI, Slack, GitHub, Microsoft and Stripe also gained. Zoho lost the most, followed by Zapier, Recurly, Facebook, Trello and Stripe Billing. The 100 most-named brands took 15.1% of all brand mentions from gpt-5-3, up from 13.7% from gpt-5-2. On the control model their share fell from 13.6% to 13.2%.
gpt-5-3 named large platforms more and smaller software tools less
| Amazon | 723 | 919 | 12.3 | 401 |
| OpenAI | 108 | 225 | 6.9 | 91 |
| Slack | 195 | 304 | 6.5 | 116 |
| GitHub | 405 | 480 | 6.4 | 173 |
| Microsoft | 996 | 1,046 | 5.5 | 363 |
| Stripe | 224 | 313 | 4.8 | 82 |
| Snowflake | 41 | 111 | 4.7 | 52 |
| Notion | 175 | 242 | 4.5 | 88 |
| Target | 262 | 363 | 4.4 | 241 |
| Walmart | 668 | 714 | 4 | 376 |
| Apple | 347 | 323 | -2.5 | 171 |
| PayWhirl | 55 | 5 | -2.7 | 1 |
| Wave | 61 | 30 | -2.8 | 10 |
| Stripe Billing | 89 | 40 | -2.9 | 5 |
| App Store | 96 | 44 | -2.9 | 51 |
| Trello | 121 | 71 | -3 | 30 |
| Recurly | 150 | 87 | -3.9 | 7 |
| 255 | 204 | -3.9 | 94 | |
| Zapier | 197 | 133 | -5.5 | 79 |
| Zoho | 436 | 290 | -11.4 | 105 |
gpt-5-6 named large enterprise suites less often
Across the gpt-5-5 to gpt-5-6 update, the biggest losses went to large enterprise software. SAP, Oracle, Microsoft Power BI, Microsoft Dynamics 365 and Microsoft 365 Apps each lost between 45% and 56% of their appearances. SAP fell from 136 appearances to 60 in 17,416 pairs.
Gains were smaller, because gpt-5-6 named fewer brands overall. QuickBooks, Shopify, Apple and Salesforce Sales Cloud gained the most. Change is net of what the same brand did when the same model was asked again.
gpt-5-6 named large enterprise suites less often
| SAP | 136 | 60 | -4.2 | 50 |
| ChatGPT Work | 207 | 137 | -3.8 | 57 |
| Microsoft Power BI | 114 | 50 | -3.7 | 41 |
| Oracle | 110 | 52 | -3.7 | 53 |
| Microsoft 365 Apps | 106 | 58 | -3 | 49 |
| Microsoft Dynamics 365 | 114 | 59 | -2.9 | 60 |
| Claude Code Action | 125 | 89 | -2.9 | 26 |
| OpenAI | 147 | 98 | -2.5 | 60 |
| Google Workspace | 124 | 81 | -2.4 | 46 |
| Notion | 206 | 165 | -2.4 | 47 |
| Prometheus | 19 | 40 | 1.2 | 13 |
| Airbnb | 81 | 89 | 1.2 | 1 |
| Ramp | 58 | 83 | 1.3 | 16 |
| Wayfair | 39 | 55 | 1.3 | 30 |
| Square Subscriptions | 6 | 29 | 1.5 | 0 |
| Salesforce Sales Cloud | 178 | 174 | 1.7 | 79 |
| Grafana OSS | 2 | 35 | 1.7 | 7 |
| Shopify | 17 | 45 | 1.9 | 21 |
| Apple | 111 | 131 | 1.9 | 56 |
| QuickBooks | 82 | 112 | 2.8 | 45 |
gpt-5-6 put OpenAI's own products first about half as often
ChatGPT Work was the top brand in 84 answers from gpt-5-5 and in 39 answers from gpt-5-6, on the same prompts. OpenAI went from 58 to 35. Together, OpenAI's products went from 142 top spots to 74, a 48% drop. Over the same pairs, gpt-5-6 named ChatGPT Work 34% less often anywhere in the answer.
The earlier update moved the other way. gpt-5-3 named OpenAI 225 times where gpt-5-2 had named it 108 times. A model's treatment of its own company's products is not fixed. It can change with each version.
gpt-5-6 put OpenAI's own products first about half as often
- ChatGPT Work, gpt-5-584
- ChatGPT Work, gpt-5-639
- OpenAI, gpt-5-558
- OpenAI, gpt-5-635
gpt-5-6 added more reservations to the brands it named
A reservation is a condition, fallback, comparison, budget limit or risk warning attached to a recommended brand. We measured it per sentence that names a brand.
On the same prompts, 20.3% of brand sentences from gpt-5-5 carried a reservation and 23.0% from gpt-5-6. Sentences with a hedged tone nearly doubled, from 2.7% to 5.2%. Positive sentences also rose, from 48.1% to 53.2%. So gpt-5-6 was warmer and more conditional at the same time. Google AI Mode stayed at about 8% over the same dates.
- Brand sentences with a reservation
- 20.3% → 23.0%Brand sentences with a reservation
- Brand sentences with a hedged tone
- 2.7% → 5.2%Brand sentences with a hedged tone
- Positive brand sentences
- 48.1% → 53.2%Positive brand sentences
- Reservations on Google AI Mode, same dates
- 8.2% → 8.1%Reservations on Google AI Mode, same dates
After gpt-5-6, brand lists changed more in 40 of 43 markets
Brand churn is the share of brands that did not come back on the next answer, which is one minus shortlist overlap. We compared it across the update with the same model re-asked, market by market, for the 43 markets with at least 20 pairs of each kind.
Churn rose in 40 of the 43 markets, by 7.0 points on average. Reddit and Community Marketing Services rose the most, from 71.4% to 88.0%. Only Payroll and HRIS Software, Small Business Accounting Software and AI Meeting Assistant & Transcription Tools did not rise. Market samples are 20 to 79 prompts, so read single markets as direction, not precise size.
After gpt-5-6, brand lists changed more in 40 of 43 markets
Churn across update (%) · 10 results
- Reddit and Community Marketing Services88%
- VC & Angel Investor Databases83.2%
- Home Spa and Wellness Products78.9%
- AI Search Visibility Analytics Tools73%
- Cloud Infrastructure Management and Security70.4%
- LLM Agent Frameworks and Tooling69.4%
- Corporate Learning Management Systems (LMS/LXP)66.9%
- Full-Stack Observability Platforms66.7%
- Professional Networking and Career Platforms63.9%
- Crypto Trading Platforms and Wallets62.5%
Explore the data (10 rows)
| Reddit and Community Marketing Services | 49 | 88 | 71.4 | 16.6 |
| Cloud Infrastructure Management and Security | 24 | 70.4 | 55.9 | 14.5 |
| Professional Networking and Career Platforms | 28 | 63.9 | 49.6 | 14.3 |
| Crypto Trading Platforms and Wallets | 21 | 62.5 | 50 | 12.5 |
| Corporate Learning Management Systems (LMS/LXP) | 23 | 66.9 | 54.6 | 12.3 |
| Full-Stack Observability Platforms | 44 | 66.7 | 55.1 | 11.6 |
| VC & Angel Investor Databases | 20 | 83.2 | 71.8 | 11.4 |
| AI Search Visibility Analytics Tools | 46 | 73 | 61.6 | 11.4 |
| Home Spa and Wellness Products | 48 | 78.9 | 67.7 | 11.2 |
| LLM Agent Frameworks and Tooling | 21 | 69.4 | 58.5 | 10.9 |
Over two months, ordinary drift was as large as an update
The middle update, gpt-5-3 to gpt-5-5, came after a four-week gap in our ChatGPT records, so the closest pairs are about two months apart. Across it, 30.1% of pairs kept the same top brand. The same model, gpt-5-5, asked two months apart, kept it in 31.8%. Google AI Mode kept it in 37.5% and 37.7% over the same two spans.
At this distance the update adds less than 2 points. The answers have already drifted so far that the update is hard to see. This is why we measure updates on asks a few days apart.
Over two months, ordinary drift was as large as an update
- ChatGPT, gpt-5-3 then gpt-5-530.1%
- ChatGPT, gpt-5-5 both times31.8%
- Google AI Mode, same span as row 137.5%
- Google AI Mode, same span as row 237.7%
What we excluded and why
We used only answers whose recorded model version we could read. In the first update window, 0.1% of ChatGPT answers had no version and were left out. From August 21, 2026, most ChatGPT Search answers stopped carrying a version, so the gpt-5-6 window ends on August 20.
Some days had incomplete or inconsistent brand detection in our records, or only partial collection. The first window uses February 13 to 26 before the update and March 7 to 9 and 17 to 20 after it, and skips the days in between for that reason. April 2026 is not used at all.
Two more limits. Before May 2026 our records label the two engines only as ChatGPT and Google, and we treat the Google records as Google AI Mode. And the gpt-5-3 to gpt-5-5 update coincided with a change in how we collected answers, so we report it only as a two-month comparison. Every figure is an observed cut of Parse data over the stated windows, not a live reading.
- First-window ChatGPT answers with no model version, excluded
- 0.1%First-window ChatGPT answers with no model version, excluded
- Last day with model versions on ChatGPT Search
- Aug 20Last day with model versions on ChatGPT Search
- Gap in ChatGPT records before gpt-5-5
- 28 daysGap in ChatGPT records before gpt-5-5
The GEO takeaway: check model updates before explaining a change
A model update can change your top-brand position in 1 more answer out of 10 with no change on your side. It can cut the number of brands per answer by a sixth. And it can do this without moving your citations.
Keep a log of engine model updates next to your AI visibility numbers. Judge every change against the noise floor of the same model, re-asked. Track the top brand, the full shortlist and citations as three separate numbers. And re-baseline after each update instead of comparing across it.
How we measured
We analyzed 376,975 same-prompt answer pairs from ChatGPT, ChatGPT Search and Google AI Mode between February 13 and August 28, 2026, around the gpt-5-2 to gpt-5-3, gpt-5-3 to gpt-5-5 and gpt-5-5 to gpt-5-6 model updates, with an unchanged control model and same-model noise floors.
- Kept the top brand across gpt-5-2 to gpt-5-3
- 38.6% vs 48.2%Kept the top brand across gpt-5-2 to gpt-5-3versus an unchanged control model
- Kept the top brand across gpt-5-5 to gpt-5-6
- 42.6% vs 50.0%Kept the top brand across gpt-5-5 to gpt-5-6versus the same model re-asked
- Google AI Mode on the same date
- +2.7 ptsGoogle AI Mode on the same dateno drop where nothing changed
- Brands per answer after gpt-5-6
- −16%Brands per answer after gpt-5-6
Get the data
Sources
- Search Engine Journal: AI recommendations change with nearly every query, SparkToro study (January 2026) · accessed October 6, 2026
- Evertune: ChatGPT gets pickier, GPT-5.4 mini recommends 37% fewer brands (March 2026) · accessed October 6, 2026
- Writesonic: GPT-5.5 Instant citation study · accessed October 6, 2026
- Search Engine Land: what three months of AI visibility tracking data reveals · accessed October 6, 2026