In Parse's 185,723 head-to-head brand comparisons, AI rarely produced one overall winner. When a comparison covered two recurring dimensions, such as price and ease of use, the winning brand changed by dimension 50.7% of the time. With three or more dimensions, it changed 73.3%. Teams should measure who wins each decision factor, not only the final recommendation.
AI compares dimensions instead of choosing one winner
Parse records, for each AI answer, the explicit comparisons the model makes between two brands: the subject, the object, the dimension being judged (the "axis"), and which brand the model called superior. Across the panel that is 185,723 comparison verdicts. The raw axis labels are messy (the model writes "pricing," "cost," "fees," "value for money" for what is really one idea), so we mapped them to six recurring dimensions and kept only high-confidence verdicts where the model named a clear winner. That leaves 51,790 decisive verdicts spread across 29,592 distinct brand matchups.
The headline is what happens when the same pair is judged on more than one dimension. Take every matchup the model decided on at least two of the six dimensions: in 50.7% of them, brand A won at least one dimension and brand B won at least one other. Raise the bar to matchups decided on three or more dimensions, and the split rate climbs to 73.3%. The more thoroughly AI compares two brands, the less likely it is to end up with a single winner.
- Across 185,723 head-to-head comparisons AI made between named brands, the winner depends on which dimension you ask about. On matchups judged on two dimensions, the winner differed by dimension 50.7% of the time; on three or more, 73.3% (Parse, four engines, October 2025 to June 2026).
- AI is most willing to declare a winner on price: 79.1% of price comparisons end in a clear verdict. It is least willing on features, where only 56.7% resolve and the rest end in "it depends" or "equal."
- The split-verdict behavior is consistent across engines: ChatGPT (51.6%), ChatGPT Search (50.0%), and Google AI Overviews (49.0%) all split on roughly half of two-dimension matchups.
- A recurring shape: the premium or incumbent brand tends to win ease of use, the challenger wins price, and the more configurable option wins features. Shopify beats WooCommerce on ease of use while losing on price and customization; 1Password beats Bitwarden on usability while losing on price.
- "Winning the comparison" is not one outcome. A competitor who "beats you in ChatGPT" almost certainly beats you on one axis while you win another, and that is invisible to any tool that reports a single win or loss.
How we measured the split
We used Parse's comparison-fact layer, which extracts every explicit two-brand comparison the model states inside an answer. Each fact carries the subject brand, the object brand, the axis the model used (its own words), the direction of the verdict (subject better, subject worse, equal, tradeoff, or unclear), and a confidence score. The window is October 2025 to June 2026 and spans all four surfaces Parse has monitored in that period: the earlier ChatGPT and Google AI Overviews collections and the current ChatGPT Search and Google AI Mode collections.
To make axes comparable across the model's loose vocabulary, we mapped each verdict to one of six recurring dimensions: price (price, cost, fees, value), ease of use (usability, learning curve, setup, complexity), performance (speed, latency, reliability, accuracy), features (functionality, customization, integrations, flexibility), support, and security. We kept only verdicts with confidence at least 0.7 and a clear winner (the model said one brand was better, not "equal" or "a tradeoff"). For each matchup and dimension we took the brand the model judged superior more often, requiring a strict majority so a tie counts as undecided. A matchup is "split" when, across the dimensions it was decided on, each brand won at least one. Everything is aggregate across Parse's monitored panel, with no single customer identifiable.
A note on the bar we set. Requiring a strict per-dimension majority is conservative: it throws out matchups where the model went back and forth within a single dimension. Tightening the rule further (at least two verdicts behind every decided dimension) moves the two-dimension split rate from 50.7% to 55.5%, so the finding does not depend on thin evidence.
Where AI commits, and where it hedges
Not all dimensions are equal. The model is far readier to name a winner on some than on others. Of all the price comparisons in the panel, 79.1% ended in a clear verdict; only about one in five came back "equal" or "a tradeoff." Features sat at the opposite end: just 56.7% resolved, and a quarter ended in "equal," with another 15% flagged as an explicit tradeoff.
| Dimension | Comparisons end in a clear winner | Comparisons end "equal" or "tradeoff" |
|---|---|---|
| Price | 79.1% | 19.9% |
| Ease of use | 75.1% | 24.1% |
| Performance | 73.1% | 25.1% |
| Support | 69.1% | 29.4% |
| Security | 63.7% | 34.0% |
| Features | 56.7% | 39.4% |
The pattern is intuitive once you see it. Price is close to objective: one product costs less, and the model will say so. Features are where "better" depends on what you need, so the model hedges, calls it a tradeoff, or declares a tie. This is the mechanism behind the split. The dimensions AI commits on (price, ease of use) tend to favor different brands than the dimensions it weighs alongside them, so a matchup decided on several axes accumulates winners on both sides.
If you want to see how AI engines describe your own brand, run a free brand check — it takes a minute.
The shape of a split, in brands you know
The clearest examples are common buyer comparisons. Shopify versus WooCommerce was the most-compared pair in the panel, and each brand won different dimensions.
| Dimension | Winner AI names | The model's reasoning, in its words |
|---|---|---|
| Ease of use | Shopify | "Shopify is fully managed with no server maintenance or security worries, whereas WooCommerce requires the owner to handle updates." |
| Performance | Shopify | "WooCommerce requires optimization work to match Shopify's speed." |
| Support | Shopify | Managed hosting and PCI compliance handled by the platform. |
| Features | WooCommerce | "WooCommerce provides superior flexibility to modify any aspect of the site." |
| Price | WooCommerce | "Shopify has a higher initial cost; WooCommerce is a free plugin, pay for hosting." |
Ask "which is better" and there is no answer. Ask "which is easier" and it is Shopify every time; ask "which is cheaper and more customizable" and it is WooCommerce every time. The same shape recurs across categories and is remarkably stable for the brands buyers compare most:
- 1Password vs Bitwarden. 1Password wins ease of use and features; Bitwarden wins price decisively (20 verdicts to none) and edges security.
- Coinbase vs Kraken. Coinbase wins ease of use; Kraken wins on fees (27 verdicts to none) and security.
- HubSpot vs Salesforce. HubSpot wins ease of use and price; Salesforce wins features.
- Cypress vs Playwright. Cypress wins ease of use; Playwright wins performance (13 to 2), support, and features.
There is a soft regularity here worth naming, even though it is a tendency and not a law: the premium or incumbent brand tends to take ease of use and support, the challenger takes price, and the more open or configurable option takes features. It is not a clean rule (across matchups decided on both price and ease of use, the same brand wins both about as often as not), so do not treat it as one. But it explains why so many real matchups split: AI is not weighing one quality, it is weighing several that rarely point the same way.
The same behavior on every engine
A reasonable worry is that this is an artifact of one model's style. It is not. We computed the two-dimension split rate separately per engine, and three of the four land within three points of each other.
| Engine | Matchups decided on 2+ dimensions | Split rate |
|---|---|---|
| ChatGPT | 374 | 51.6% |
| ChatGPT Search | 418 | 50.0% |
| Google AI Overviews | 1,656 | 49.0% |
| Google AI Mode | 169 | 29.0% |
ChatGPT, ChatGPT Search, and Google AI Overviews all split on about half of their multi-dimension matchups. Google AI Mode is the outlier at 29%, but on a much smaller base: AI Mode writes terser comparisons and reaches two decided dimensions on a given pair far less often, so it has fewer chances to split. The consistency across the other three says the behavior is a property of how these systems reason about a comparison, not a quirk of one model's prose.
What this changes about competitor analysis
If you track your brand's AI visibility as a single comparison result, you are measuring the wrong thing. The first step is knowing who you are even being compared against: Parse's data on which competitors AI pairs your brand with maps the rivals that show up alongside you. "We lose to a competitor in ChatGPT" is almost never the whole story, because the model does not render a whole-story verdict. The useful questions are per axis. Which dimensions do you win, and which do you lose? A brand that loses on price but wins on ease of use is in a completely different position from one that loses on both, and the two demand opposite responses: the first defends and amplifies its usability story, the second has a real product gap.
You do not have to be better overall to change the answer because the model often does not choose one overall winner. You need stronger evidence on a specific dimension. If a rival consistently wins "ease of use" because of review-site claims, improve the documentation, reviews, and comparisons that support that dimension. Measure each dimension separately.
None of this makes AI comparisons soft. On price, the model commits four times out of five, and on the dimensions where it commits, it is consistent enough that you can read a real position off the data. The point is that the position is multidimensional. The brands that manage their AI presence well are the ones that know exactly which axes they own, which they are contesting, and which they have quietly conceded, rather than the ones chasing a single number that was never going to resolve.
How Parse measures head-to-head comparisons
Parse records the rival, dimension, and verdict in each AI comparison. The public index spans more than 4.7 million AI responses, 603,000 brands, and 57 million citations. Parse's Brand Lookup shows which dimensions each brand wins against a competitor, preserving detail that a single win-loss number hides. For related reading, see a brand mention is not an AI recommendation, what words AI uses to describe brands, and why one AI visibility score is misleading.
Does AI pick a single winner when it compares two brands?
Usually not. In Parse's data, when AI judged a brand matchup on two of six recurring dimensions, the preferred brand changed by dimension 50.7% of the time. With three or more dimensions, it changed 73.3% of the time. AI often identifies a winner for each factor without choosing one overall winner.
Which dimension is AI most willing to declare a winner on?
Price. Across the panel, 79.1% of price comparisons ended in a clear verdict, because price is close to objective. Features is the opposite: only 56.7% resolved, and the rest came back "equal" or "a tradeoff," because "better features" depends on what the buyer needs.
Is this split behavior the same across ChatGPT and Google?
Largely yes. ChatGPT (51.6%), ChatGPT Search (50.0%), and Google AI Overviews (49.0%) all split on about half of their two-dimension matchups. Google AI Mode splits less often (29%), but it writes terser comparisons and reaches two decided dimensions on a pair far less frequently, on a much smaller base.
What does the split mean for competitor analysis?
It means you should track AI comparisons per axis, not as a single win or loss. A competitor who "beats you" almost certainly beats you on one dimension while you win another. Knowing which axes you own, contest, and have conceded tells you where to defend and where you have a real product gap.
How can a brand change the verdict AI gives in a comparison?
Improve evidence for the specific dimension the brand is losing. For an "ease of use" result, that may mean clearer documentation, current reviews, and credible comparisons. Re-run the same prompts and measure whether that dimension changes.