Which AI engine changes its answers the most?
Nearly a tie. Ask the same question again and ChatGPT Search keeps 40.5% of its brand list while Google AI Mode keeps 41.5% — and the less stable engine changes month to month.
By Dimitry Apollonsky · August 29, 2026 · 10 min read
Contents
- Neither engine keeps its brand list: overlap is about 41% on both
- The #1 pick changed about 45% of the time on both engines
- An identical brand list came back 5.6% of the time
- ChatGPT Search replaced the entire list 1.3 times as often
- The typical repeat answer keeps a quarter to half of the list
- About 86% of repeat answers add at least one new brand
- The most volatile engine flipped month to month
- Waiting longer barely lowers overlap
- Tech brand lists are the steadiest; apparel and real estate change most
- A volatile prompt is volatile on both engines
- One-brand answers are the least repeatable
- What this study counts, and what it leaves out
- The GEO takeaway
- Get the data
- Sources
- Related research
We analyzed 663,398 consecutive same-prompt, same-engine answer pairs to 17,039 organic prompts asked on both ChatGPT Search and Google AI Mode from May 24 through August 22, 2026.
Neither engine keeps its brand list: overlap is about 41% on both
An answer's brand list is the set of distinct resolved brands it names directly. Overlap is the share of brands the two answers agree on: brands named in both answers divided by all brands named in either. On consecutive answers to the same prompt, mean overlap was 40.5% on ChatGPT Search (334,028 pairs) and 41.5% on Google AI Mode (329,370 pairs).
The gap between engines is 1.0 points. The gap between either engine and a stable answer is nearly 60 points. Which engine changes most is the wrong question: on average, roughly 3 of every 5 brands across the two answers appear in only one of them, on either engine.
Takeaway
The #1 pick changed about 45% of the time on both engines
The #1 pick is the single brand an answer places at recommendation position 1. On pairs where both answers had exactly one #1 pick, it stayed the same in 54.7% of 266,788 ChatGPT Search pairs and 55.0% of 268,453 Google AI Mode pairs. It changed 45.3% and 45.0% of the time.
This matches our earlier pooled study, which found the top recommendation changed in 43.6% of consecutive pairs over a shorter window. The engines are indistinguishable on this measure: 0.27 points apart.
An identical brand list came back 5.6% of the time
Across both engines, 37,025 of 663,398 repeat answers returned exactly the same brand list: 5.1% on ChatGPT Search and 6.0% on Google AI Mode. That is about 1 identical list in every 18 repeat answers.
SparkToro's 2,961-run study found identical brand lists in fewer than 1% of repeat runs, and the same list in the same order in fewer than 0.1%. Our bar is looser — we compare unordered sets of resolved brands, which merges spelling and naming variants — and the lists still differ 94.4% of the time.
Takeaway
ChatGPT Search replaced the entire list 1.3 times as often
A full swap is a pair whose two answers share no brands at all. ChatGPT Search produced a full swap in 9.3% of pairs (31,194 of 334,028); Google AI Mode in 7.1% (23,434 of 329,370).
This is where the engines differ most. The average overlap is close, but ChatGPT Search reaches the extreme — an answer with zero carryover — 1.3 times as often. Google AI Mode also returned identical lists slightly more often (6.0% vs 5.1%). By tail behavior, ChatGPT Search is the engine that changes its answers the most, by a modest margin.
The typical repeat answer keeps a quarter to half of the list
The most common outcome on both engines is partial overlap: 33.0% of ChatGPT Search pairs and 35.2% of Google AI Mode pairs landed at 25–49% overlap, and about 26% of pairs on each engine landed at 50–74%.
Full agreement and full swap are both tails. The middle of the distribution is an answer that keeps a recognizable core and rotates the rest.
| No shared brands (0%) | 9.34 | 7.11 |
| Identical list (100%) | 5.13 | 6.04 |
| 75–99% | 7.15 | 6.27 |
| 50–74% | 26.06 | 26.4 |
| 25–49% | 33.02 | 35.16 |
| 1–24% | 19.31 | 19.01 |
About 86% of repeat answers add at least one new brand
A repeat answer added at least one brand its predecessor did not name in 85.7% of ChatGPT Search pairs and 85.4% of Google AI Mode pairs. On average, 47.2% of the brands in a ChatGPT Search repeat answer were new, and 45.5% on Google AI Mode.
Answers averaged 5.9 brands on ChatGPT Search and 5.7 on Google AI Mode, so a typical repeat answer carries about 2.8 and 2.6 brands, respectively, that the previous answer did not mention. The change is not brands dropping out of a fixed list; it is a rotating set of near-equal candidate brands.
The most volatile engine flipped month to month
On a fixed panel of 13,640 prompts with repeat answers on both engines in all three months, ChatGPT Search was clearly less stable in June (36.1% overlap vs 42.0% for Google AI Mode — almost 6 points). In July the ranking flipped: ChatGPT Search 42.0%, Google AI Mode 40.8%. In August (through the 22nd) they tied: 42.0% vs 42.1%.
A one-month engine-stability ranking would have named a different winner each month. Studies that pick a most-volatile engine from a single window are measuring that window, not the engine.
Takeaway
Waiting longer barely lowers overlap
Pairs whose two answers were 0–1 days apart overlapped 39.2% on ChatGPT Search and 43.5% on Google AI Mode. At 8–14 days apart, both engines sat near 37% (37.0% and 37.0%). Pairs 4–7 days apart — the most common spacing — overlapped 41.0% and 41.0%.
If answers drifted over time, longer gaps would score much lower than same-day repeats. They barely do. Most of the change happens between any two runs, immediately; Google AI Mode's same-day advantage (4.3 points) is the one place the engines clearly separate, and it fades within a week.
Tech brand lists are the steadiest; apparel and real estate change most
Across 28 industry groups with at least 2,000 pairs per engine and 100 prompts, Information Technology prompts had the steadiest brand lists (46.1% overlap) and Clothing and Apparel the least steady (36.3%), with Real Estate (37.2%) and Consumer Goods (37.5%) close behind.
Google AI Mode was the steadier engine in 17 of the 28 groups and ChatGPT Search in 11. The industry you compete in moves overlap by up to 10 points — a bigger effect than which engine is asked.
| Information Technology | 46.12 | 47.66 | 44.55 | 26,578 |
| Apps | 45.88 | 45.24 | 46.5 | 5,073 |
| Software | 45.42 | 46.68 | 44.14 | 31,637 |
| Gaming | 45.3 | 45.06 | 45.54 | 5,507 |
| Internet Services | 45.08 | 46.43 | 43.75 | 11,464 |
| Media and Entertainment | 44.9 | 46.53 | 43.16 | 10,388 |
| Hardware | 44.57 | 45.86 | 43.21 | 13,240 |
| Consumer Electronics | 44.51 | 46.18 | 42.66 | 10,047 |
| Data and Analytics | 43.34 | 42.55 | 44.15 | 13,147 |
| Sales and Marketing | 43.1 | 42.54 | 43.63 | 9,399 |
| Financial Services | 42.91 | 43.22 | 42.61 | 36,124 |
| Administrative Services | 42.78 | 42.01 | 43.56 | 8,043 |
| Content and Publishing | 42.48 | 42.48 | 42.48 | 5,969 |
| Sports | 42.09 | 42.25 | 41.92 | 9,537 |
| Travel and Tourism | 42 | 42.39 | 41.63 | 6,700 |
| Community and Lifestyle | 41.94 | 40.38 | 43.54 | 9,188 |
| Health Care | 41.73 | 40.95 | 42.51 | 16,636 |
| Education | 41.57 | 41.44 | 41.71 | 9,628 |
| Artificial Intelligence | 40.78 | 41.43 | 40.12 | 18,793 |
| Food and Beverage | 40.65 | 40.08 | 41.25 | 12,081 |
| Transportation | 40.06 | 39.29 | 40.87 | 8,277 |
| Professional Services | 39.81 | 38.19 | 41.43 | 17,338 |
| Design | 39.62 | 38.91 | 40.36 | 4,593 |
| Commerce and Shopping | 38.98 | 37.7 | 40.29 | 23,137 |
| Manufacturing | 38.47 | 37.59 | 39.39 | 5,376 |
| Consumer Goods | 37.54 | 37.62 | 37.46 | 13,528 |
| Real Estate | 37.24 | 35.93 | 38.62 | 11,469 |
| Clothing and Apparel | 36.33 | 35.92 | 36.79 | 4,845 |
A volatile prompt is volatile on both engines
For the 14,101 prompts with at least 10 pairs on each engine, per-prompt overlap on ChatGPT Search correlates 0.61 with per-prompt overlap on Google AI Mode. 70.8% of these prompts sit on the same side of the median on both engines: 35.4% volatile on both, 35.4% steady on both.
Instability follows the question more than the engine. A question with many near-equal candidate brands is unstable on every engine; switching engines will not settle it.
One-brand answers are the least repeatable
When the smaller answer in a pair named just one brand, overlap averaged 31.8% on ChatGPT Search and 28.6% on Google AI Mode. When both answers named 10 or more brands, overlap rose to 43.1% and 46.4%.
A short answer looks decisive but is the least likely to come back the same way. Long brand lists keep a stable core and rotate the rest; a one-brand answer has the least carryover of any answer size.
| 10+ brands | 43.06 | 46.41 |
| 6–9 brands | 41.64 | 44.34 |
| 4–5 brands | 41.78 | 42.94 |
| 2–3 brands | 38.2 | 38.57 |
| 1 brand | 31.83 | 28.62 |
What this study counts, and what it leaves out
A pair is two immediately adjacent answers to the same organic prompt on the same engine, kept when both name at least one resolved brand; every cross-engine number uses the 17,039 prompts (of 17,560 with pairs) that had pairs on both engines. We computed the headline two independent ways — an array-intersection build and a row-level rebuild of the same pairs — and they agree to four decimals. Removing the matched-prompt restriction moves overlap by at most 0.04 points. Excluding pairs where either answer named a single brand moves it by under 1.5 points (40.9% and 42.4%).
The #1 pick is defined in both answers for 79.9% of ChatGPT Search pairs and 81.5% of Google AI Mode pairs; the rest have no single position-1 brand. Answers are observed on a regular cadence, so a pair is a repeat across runs, not a controlled same-minute re-ask. Brand strings are resolved to root brands, which merges spelling variants but can split unresolved names. The window ends August 22, 2026 because the brand-naming layer of the data thins after that date.
The GEO takeaway
Both engines rewrite most of the brand list between runs, the #1 pick changes about 45% of the time, and the monthly most-volatile ranking did not survive the quarter. There is no stable engine to optimize for and no volatile engine to ignore.
Measure your brand's appearance rate: the share of many repeated answers that name you, per engine, over a window of a fixed length. A brand in 8 of 10 runs is visible; a brand in the latest screenshot may not be. Track both engines — brands that drop out of one answer usually come back, and the rotation is how competitors enter answers. Rerun the comparison each quarter with this article's method before believing any engine-stability claim, including ours.
Get the data
Sources
- SparkToro: AIs are highly inconsistent when recommending brands or products · accessed 2026-08-29
- Search Engine Journal: AI recommendations change with nearly every query · accessed 2026-08-29
- Conductor: Why intent type predicts AI output consistency · accessed 2026-08-29
- Schulte, Bleeker, Kaufmann: Don't measure once — measuring visibility in AI search (GEO) · accessed 2026-08-29