Parse
Work with usPricing
Sign inCheck your brand
Research/Does AI change its mind when comparing brands?

Does AI change its mind when comparing brands?

Sometimes. On the same prompt and criterion, one AI engine picked a different brand winner in 383 of 2,438 repeated matchups, or 15.71%.

By Dimitry Apollonsky · July 16, 2026 · 9 min read

15.71%
of repeated matched comparisons changed the winner
383 of 2,438
▸Contents
  • AI changes the winner in 15.71% of repeated matched comparisons
  • Prompt matching cuts the apparent reversal rate from 27.21% to 15.71%
  • Most winner changes appear once, but 29 repeat
  • Google AI Overviews changes winners more than Google AI Mode
  • Ease of use is the least stable comparison criterion
  • The current-engine gap narrows on 127 identical matchups
  • Jira Software and Linear have the most recorded winner changes
  • The method excludes ambiguous answers and criteria outside the six groups
  • What marketers should do
  • Get the data
  • Sources
  • Related research
Contents
  • AI changes the winner in 15.71% of repeated matched comparisons
  • Prompt matching cuts the apparent reversal rate from 27.21% to 15.71%
  • Most winner changes appear once, but 29 repeat
  • Google AI Overviews changes winners more than Google AI Mode
  • Ease of use is the least stable comparison criterion
  • The current-engine gap narrows on 127 identical matchups
  • Jira Software and Linear have the most recorded winner changes
  • The method excludes ambiguous answers and criteria outside the six groups
  • What marketers should do
  • Get the data
  • Sources
  • Related research

We analyzed 240,719 explicit brand-comparison statements across 95,068 AI answers, 13,933 organic prompts, and 18,313 cleaned brand names on ChatGPT, ChatGPT Search, Google AI Overviews, and Google AI Mode from October 19, 2025 through July 9, 2026.

15.71%
of repeated matched comparisons changed the winner
383
winner changes
2,438
repeated matched comparisons
1,507
cleaned brand names in the matched set

AI changes the winner in 15.71% of repeated matched comparisons

A repeated matchup is one brand pair compared on the same prompt, criterion, and engine in at least two answers. An answer-level verdict is the one unambiguous winner in one answer. One engine picked both brands as the winner in separate answers for 383 of 2,438 repeated matchups, or 15.71%.

The matched set covers 1,091 prompts, 1,507 cleaned brand names, and 6,777 answer-level verdicts. One answer is not enough to establish a stable comparison winner.

15.71%
of repeated matched comparisons changed the winner
383 of 2,438

Takeaway

Rerun the same comparison before treating one winner as a stable AI recommendation.

Prompt matching cuts the apparent reversal rate from 27.21% to 15.71%

Grouping repeated comparisons without holding the prompt constant produces 887 winner changes across 3,260 matchups, or 27.21%. Holding the prompt, engine, brand pair, and criterion constant produces 383 across 2,438, or 15.71%.

BrightEdge compares brand sets across engines. SparkToro measures complete recommendation lists and order. Conductor measures brand-list overlap and lead-brand stability. This study measures the narrower decision of which brand wins one fixed comparison.

27.21%
without matching the prompt
887 of 3,260 matchups
15.71%
with the prompt matched
383 of 2,438 matchups

Takeaway

A consistency rate is only comparable when the prompt, engine, brand pair, and criterion are held constant.

Most winner changes appear once, but 29 repeat

Of the 383 reversing matchups, 354 have one answer for the less frequent winner. In 29 matchups, each brand wins at least twice.

A single reversal can be an isolated answer. Repeated wins for both brands identify comparisons that need closer monitoring, but they do not explain why the decision changed.

354
matchups with one answer for the less frequent winner
of 383 reversals
29
matchups where each brand wins at least twice
of 383 reversals

Google AI Overviews changes winners more than Google AI Mode

Google logoGoogle AI Overviews changes the winner in 192 of 949 repeated matchups, or 20.23%. ChatGPT logoChatGPT is at 23 of 138, or 16.67%. ChatGPT logoChatGPT Search is at 132 of 930, or 14.19%. Google logoGoogle AI Mode is at 36 of 421, or 8.55%.

The four engine samples cover different time windows and matchup mixes. The rates show where repeated checks matter in this observed cut; they are not a synchronized engine experiment.

Winner-change rate by engine
  • Google logoGoogle AI Overviews20.23%
  • ChatGPT logoChatGPT16.67%
  • ChatGPT logoChatGPT Search14.19%
  • Google logoGoogle AI Mode8.55%
Share of repeated matched comparisons that name both brands as the winner across separate answers.

Ease of use is the least stable comparison criterion

Ease-of-use comparisons change the winner in 127 of 625 repeated matchups, or 20.32%. Price is at 16.35%, features 13.99%, security 11.11%, performance 10.64%, and support 8.93%.

A single overall rate hides meaningful differences by criterion. Competitive monitoring should preserve what the brands were compared on, not only which names appeared.

Winner-change rate by comparison criterion
  • Ease of use20.32%
  • Price16.35%
  • Features13.99%
  • Security11.11%
  • Performance10.64%
  • Support8.93%
Related comparison labels from the answers are grouped into the six criteria used in this report.

The current-engine gap narrows on 127 identical matchups

We restricted ChatGPT logoChatGPT Search and Google logoGoogle AI Mode to the same 127 prompt, pair, and criterion matchups. ChatGPT logoChatGPT Search changes the winner in 13, or 10.24%, while Google logoGoogle AI Mode changes it in nine, or 7.09%.

The matched set removes a different matchup mix as the full explanation. Both rates still show that a repeated answer can reverse a fixed brand comparison.

10.24%
ChatGPT Search
13 of 127 identical matchups
7.09%
Google AI Mode
9 of 127 identical matchups

Jira Software and Linear have the most recorded winner changes

Jira Software logoJira Software and Linear logoLinear change winners eight times across 63 repeated matchups and 216 answer-level verdicts. Asana logoAsana and Trello logoTrello change seven times across 13 matchups. DraftKings logoDraftKings and FanDuel logoFanDuel, Shopify logoShopify and WooCommerce, and HelloFresh logoHelloFresh and EveryPlate logoEveryPlate each change five times.

The table requires at least 10 repeated matchups and 30 answer-level verdicts. A high count identifies comparisons to investigate; it does not show that either winner was factually correct.

Cleaned brand pairs by recorded winner changes
At least 10 repeated matchups and 30 answer-level verdicts.
Jira Software logoJira Software and Linear logoLinear86312.7216394
Asana logoAsana and Trello logoTrello71353.854283
DraftKings logoDraftKings and FanDuel logoFanDuel5202557124
Shopify logoShopify and WooCommerce51338.464483
HelloFresh logoHelloFresh and EveryPlate logoEveryPlate51241.673074
Leadpages and Instapage410403124
Shortcut and Jira Software logoJira Software3358.57112233
Asana logoAsana and ClickUp logoClickUp31816.6756163
Jira Software logoJira Software and Asana logoAsana3152039133
Make and Zapier31421.4338113
Pulley logoPulley and Carta logoCarta21315.383753
Istio and Linkerd1185.568944
Selenium and Playwright logoPlaywright1185.5678113
Coinbase logoCoinbase and Kraken logoKraken1147.144184
Datadog logoDatadog and SigNoz1147.147083

The method excludes ambiguous answers and criteria outside the six groups

The source cut contains 240,719 explicit comparison statements. We resolved decisive comparisons to cleaned brand names and grouped related labels into the six criteria used in this report. Of 29,278 grouped answer comparisons on those six criteria, 385 name both brands as the winner and are excluded.

Keeping only higher-confidence comparisons produces 366 reversals across 2,367 matchups, or 15.46%. Keeping every original criterion label separate produces 173 across 2,176, or 7.95%. The criterion grouping is a significant analysis choice, so quarterly reruns must preserve it or report the change.

240,719
explicit comparison statements
385
ambiguous answer comparisons excluded
15.46%
higher-confidence check
366 of 2,367 matchups
7.95%
exact-label check
173 of 2,176 matchups

What marketers should do

Rerun the same comparison prompt on the same engine. Record the criterion and answer-level winner. Treat a single result as one observed answer, not a durable competitive verdict.

Investigate comparisons that reverse repeatedly. Check whether the answers use different facts or definitions before changing positioning, content, or competitive claims.

Takeaway

Measure comparison consistency at the prompt, engine, brand-pair, and criterion level.

Get the data

Dataset CSVThe metrics behind every figure in this report.

Sources

  1. BrightEdge: ChatGPT vs. Google AI: 62% brand recommendation disagreement · accessed 2026-07-16
  2. SparkToro: AIs are highly inconsistent when recommending brands or products · accessed 2026-07-16
  3. Conductor: AI recommendation consistency analysis · accessed 2026-07-16

Related research

How AI picks a winner in head-to-head comparisons
When AI compares two brands on more than one thing, it picks a different winner about half the time. There is no single winner, only a winner per axis.
AI citation volatility by industry: one-shot checks miss the signal
Ask the same AI question again and the cited sources usually change. Across repeated runs, ChatGPT answers shared only about 21% of their cited sources, and every industry showed high churn.
ChatGPT vs Google: same winner, different shortlist
The popular line is that AI engines disagree about who wins. They don't. Across 1,655 buyer categories ChatGPT and Google AI Overviews pick the same #1 brand 93% of the time. What they disagree about is everyone else on the list.
ChatGPT Search vs Google AI Mode: same brands, different sources
The two newest AI search surfaces do not read the same web. On the same 18,206 prompts, ChatGPT Search and Google AI Mode share only 6.2% of cited sources per prompt, while sharing 26.9% of named brands.
Does AI recommend the same brand when you ask again?
Only about six in ten times. The top recommendation stayed the same in 90,817 of 161,023 consecutive same-prompt, same-engine answer pairs, or 56.40%.

About this research

Dimitry Apollonsky

Founder, Parse

I built Parse to track where AI answers really come from: the sources they cite and the brands they name. DM me on LinkedIn to talk shop.

Track whether your brand wins the same comparison consistently.

Run a free check against live AI answers — no account needed.

Parse

Parse indexes AI recommendations so brands know where they stand.

Products

  • Brands
  • Markets
  • Integrations
  • Work with us
  • Pricing
  • MCP

Resources

  • Research
  • Methodology
  • Blog

© 2026 Parse. All rights reserved.

LegalPrivacy PolicyTerms of Service