Does AI change its mind when comparing brands?
Sometimes. On the same prompt and criterion, one AI engine picked a different brand winner in 383 of 2,438 repeated matchups, or 15.71%.
By Dimitry Apollonsky · July 16, 2026 · 9 min read
Contents
- AI changes the winner in 15.71% of repeated matched comparisons
- Prompt matching cuts the apparent reversal rate from 27.21% to 15.71%
- Most winner changes appear once, but 29 repeat
- Google AI Overviews changes winners more than Google AI Mode
- Ease of use is the least stable comparison criterion
- The current-engine gap narrows on 127 identical matchups
- Jira Software and Linear have the most recorded winner changes
- The method excludes ambiguous answers and criteria outside the six groups
- What marketers should do
- Get the data
- Sources
- Related research
We analyzed 240,719 explicit brand-comparison statements across 95,068 AI answers, 13,933 organic prompts, and 18,313 cleaned brand names on ChatGPT, ChatGPT Search, Google AI Overviews, and Google AI Mode from October 19, 2025 through July 9, 2026.
AI changes the winner in 15.71% of repeated matched comparisons
A repeated matchup is one brand pair compared on the same prompt, criterion, and engine in at least two answers. An answer-level verdict is the one unambiguous winner in one answer. One engine picked both brands as the winner in separate answers for 383 of 2,438 repeated matchups, or 15.71%.
The matched set covers 1,091 prompts, 1,507 cleaned brand names, and 6,777 answer-level verdicts. One answer is not enough to establish a stable comparison winner.
Takeaway
Prompt matching cuts the apparent reversal rate from 27.21% to 15.71%
Grouping repeated comparisons without holding the prompt constant produces 887 winner changes across 3,260 matchups, or 27.21%. Holding the prompt, engine, brand pair, and criterion constant produces 383 across 2,438, or 15.71%.
BrightEdge compares brand sets across engines. SparkToro measures complete recommendation lists and order. Conductor measures brand-list overlap and lead-brand stability. This study measures the narrower decision of which brand wins one fixed comparison.
Takeaway
Most winner changes appear once, but 29 repeat
Of the 383 reversing matchups, 354 have one answer for the less frequent winner. In 29 matchups, each brand wins at least twice.
A single reversal can be an isolated answer. Repeated wins for both brands identify comparisons that need closer monitoring, but they do not explain why the decision changed.
Google AI Overviews changes winners more than Google AI Mode
Google AI Overviews changes the winner in 192 of 949 repeated matchups, or 20.23%.
ChatGPT is at 23 of 138, or 16.67%.
ChatGPT Search is at 132 of 930, or 14.19%.
Google AI Mode is at 36 of 421, or 8.55%.
The four engine samples cover different time windows and matchup mixes. The rates show where repeated checks matter in this observed cut; they are not a synchronized engine experiment.
Ease of use is the least stable comparison criterion
Ease-of-use comparisons change the winner in 127 of 625 repeated matchups, or 20.32%. Price is at 16.35%, features 13.99%, security 11.11%, performance 10.64%, and support 8.93%.
A single overall rate hides meaningful differences by criterion. Competitive monitoring should preserve what the brands were compared on, not only which names appeared.
The current-engine gap narrows on 127 identical matchups
We restricted ChatGPT Search and
Google AI Mode to the same 127 prompt, pair, and criterion matchups.
ChatGPT Search changes the winner in 13, or 10.24%, while
Google AI Mode changes it in nine, or 7.09%.
The matched set removes a different matchup mix as the full explanation. Both rates still show that a repeated answer can reverse a fixed brand comparison.
Jira Software and Linear have the most recorded winner changes
Jira Software and
Linear change winners eight times across 63 repeated matchups and 216 answer-level verdicts.
Asana and
Trello change seven times across 13 matchups.
DraftKings and
FanDuel,
Shopify and WooCommerce, and
HelloFresh and
EveryPlate each change five times.
The table requires at least 10 repeated matchups and 30 answer-level verdicts. A high count identifies comparisons to investigate; it does not show that either winner was factually correct.
| 8 | 63 | 12.7 | 216 | 39 | 4 | |
| 7 | 13 | 53.85 | 42 | 8 | 3 | |
| 5 | 20 | 25 | 57 | 12 | 4 | |
| 5 | 13 | 38.46 | 44 | 8 | 3 | |
| 5 | 12 | 41.67 | 30 | 7 | 4 | |
| Leadpages and Instapage | 4 | 10 | 40 | 31 | 2 | 4 |
| Shortcut and | 3 | 35 | 8.57 | 112 | 23 | 3 |
| 3 | 18 | 16.67 | 56 | 16 | 3 | |
| 3 | 15 | 20 | 39 | 13 | 3 | |
| Make and Zapier | 3 | 14 | 21.43 | 38 | 11 | 3 |
| 2 | 13 | 15.38 | 37 | 5 | 3 | |
| Istio and Linkerd | 1 | 18 | 5.56 | 89 | 4 | 4 |
| Selenium and | 1 | 18 | 5.56 | 78 | 11 | 3 |
| 1 | 14 | 7.14 | 41 | 8 | 4 | |
| 1 | 14 | 7.14 | 70 | 8 | 3 |
The method excludes ambiguous answers and criteria outside the six groups
The source cut contains 240,719 explicit comparison statements. We resolved decisive comparisons to cleaned brand names and grouped related labels into the six criteria used in this report. Of 29,278 grouped answer comparisons on those six criteria, 385 name both brands as the winner and are excluded.
Keeping only higher-confidence comparisons produces 366 reversals across 2,367 matchups, or 15.46%. Keeping every original criterion label separate produces 173 across 2,176, or 7.95%. The criterion grouping is a significant analysis choice, so quarterly reruns must preserve it or report the change.
What marketers should do
Rerun the same comparison prompt on the same engine. Record the criterion and answer-level winner. Treat a single result as one observed answer, not a durable competitive verdict.
Investigate comparisons that reverse repeatedly. Check whether the answers use different facts or definitions before changing positioning, content, or competitive claims.
Takeaway
Get the data
Sources
- BrightEdge: ChatGPT vs. Google AI: 62% brand recommendation disagreement · accessed 2026-07-16
- SparkToro: AIs are highly inconsistent when recommending brands or products · accessed 2026-07-16
- Conductor: AI recommendation consistency analysis · accessed 2026-07-16