When does AI pick the smaller brand over the leader?
Often. In 44% of decisive head-to-head verdicts inside AI answers, the win went to the lower-ranked brand. On cost, the smaller brand won 58% of the time.
By Dimitry Apollonsky · August 29, 2026 · 9 min read
Contents
- AI handed the win to the lower-ranked brand in 44% of verdicts
- 75,644 verdicts cover 30,618 brand pairs
- Cost is the challenger's axis: upsets reach 58%
- Market position and integrations stay with the leader
- A bigger rank gap barely protects the leader
- Top-1,000 leaders lost 45% of verdicts to brands outside the top 10,000
- Google AI Mode upset most; ChatGPT upset least
- OpenObserve beat Datadog on cost in 115 of 116 verdicts
- In 4 of 10 contested matchups the challenger held the majority
- Two of three comparison claims named a winner
- What we excluded and why
- The GEO takeaway
- Get the data
- Sources
- Related research
We analyzed 418,813 high-confidence head-to-head comparison claims inside AI answers on ChatGPT, ChatGPT Search, Google AI Overviews, and Google AI Mode from October 2025 through July 19, 2026, and scored the 75,644 distinct verdicts where both brands hold a distinct rank on the public Parse index.
AI handed the win to the lower-ranked brand in 44% of verdicts
A verdict is one distinct comparison inside an AI answer that names a winning brand and a losing brand on a stated axis, such as cost or ease of use. We joined both brands in each verdict to their rank on the public Parse index. The leader is the pair's better-ranked brand, the challenger is the worse-ranked brand, and an upset is a verdict the challenger wins. The challenger won 33,556 of 75,644 verdicts, or 44.4%.
Smaller here means less visible in AI answers, measured by index rank. It does not mean smaller by revenue or headcount. Even with that definition, the result is close to an even contest: the brand that AI mentions less often still wins more than 4 in 10 of the head-to-head calls AI makes.
Takeaway
75,644 verdicts cover 30,618 brand pairs
The corpus started as 418,813 comparison claims read from AI answers at high extraction confidence. 199,624 of them named a winner and a loser on an axis. After mapping both brands unambiguously to the Parse index, requiring two distinct brands with two distinct ranks, and removing repeated claims within one answer, 75,644 verdicts remained.
Those verdicts span 30,618 brand pairs, 19,847 brands, and 47,979 answers. The upset result is not driven by one category or one well-known matchup.
Cost is the challenger's axis: upsets reach 58%
Each verdict carries an axis. We grouped the raw axis labels with Parse's canonical comparison-axis vocabulary. On cost, the challenger won 4,711 of 8,166 verdicts, or 57.7%. That is the only axis where the smaller brand wins more often than the leader.
Performance (47.4%), specialization (47.3%), and ease of use (47.3%) sit close to an even split. Functionality (42.5%), quality (41.0%), and reliability (40.7%) lean toward the leader.
Takeaway
Market position and integrations stay with the leader
The two axes with the lowest upset rates are the two that restate the leader's position. On market position, the challenger won 398 of 1,358 verdicts, or 29.3%. On integrations, 769 of 2,575, or 29.9%. Both sit roughly 15 percentage points below the 44.4% overall rate.
The pattern is consistent: AI gives popularity and ecosystem breadth to the visible leader, and gives price and focus to the challenger.
A bigger rank gap barely protects the leader
We bucketed verdicts by the rank-gap ratio: the challenger's index rank divided by the leader's. Near-peers (ratio under 2x) produced upsets 45.7% of the time. A 10x to 100x gap lowered that to 42.0%. A gap of 100x or more moved it back up to 47.6%.
The whole range spans less than 6 percentage points. Being far more visible than a rival does not buy the leader a proportionally safer verdict.
Takeaway
Top-1,000 leaders lost 45% of verdicts to brands outside the top 10,000
We isolated the matchups with the most lopsided ranks: the leader ranks in the index top 1,000 and the challenger ranks outside the top 10,000. There were 6,022 such verdicts. The challenger won 2,687 of them, or 44.6%.
That matches the overall upset rate almost exactly. In these verdicts, the written outcome tracks the stated axis rather than the visibility order, even at the widest rank gaps we can measure.
Google AI Mode upset most; ChatGPT upset least
Google AI Mode handed 49.3% of its verdicts to the challenger, the highest of the four engines. ChatGPT was lowest at 40.6%, with ChatGPT Search at 42.0% and Google AI Overviews at 46.1%.
The two Google engines sit above the two OpenAI engines. The legacy pair (ChatGPT, Google AI Overviews) was observed mostly from October 2025 through April 2026, and the current pair (ChatGPT Search, Google AI Mode) from May 24, 2026, so this is an engine comparison across adjacent collection windows, not a controlled same-day test.
OpenObserve beat Datadog on cost in 115 of 116 verdicts
Upsets are not random noise. Some challengers win the same axis against the same leader almost every time. OpenObserve, ranked #2,171 on the index, beat Datadog, ranked #21, on cost in 115 of 116 verdicts. SigNoz beat Datadog on cost in 86 of 87. FanDuel beat DraftKings on ease of use in 45 of 50.
The table lists the ten highest-volume challenger streaks: pairs with at least 15 verdicts on one named axis where the challenger won at least 80% of them. Observability tools dominate the list because cost comparisons in that category are frequent and one-sided.
| Datadog (#21) | OpenObserve (#2,171) | Cost | 116 | 115 | 99.1 |
| Datadog (#21) | SigNoz (#2,169) | Cost | 87 | 86 | 98.9 |
| Splunk (#757) | Grafana Loki (#2,550) | Cost | 53 | 48 | 90.6 |
| DraftKings (#3) | FanDuel (#22) | Ease of use | 50 | 45 | 90 |
| Splunk (#757) | OpenObserve (#2,171) | Cost | 49 | 49 | 100 |
| Datadog (#21) | Grafana Loki (#2,550) | Cost | 47 | 45 | 95.7 |
| Ethereum (#2) | Solana (#35) | Cost | 44 | 38 | 86.4 |
| Datadog (#21) | New Relic (#294) | Cost | 43 | 41 | 95.3 |
| Datadog (#21) | Coralogix (#3,290) | Cost | 41 | 40 | 97.6 |
| Shopify (#15) | Woocommerce (#540) | Functionality | 34 | 28 | 82.4 |
Takeaway
In 4 of 10 contested matchups the challenger held the majority
We grouped verdicts into pair-axis cells: one brand pair judged on one axis. Among the 2,042 cells with at least 5 verdicts, the challenger won the majority of verdicts in 811, or 39.7%. Another 78 cells split exactly even.
Upsets are therefore not scattered one-off calls. In a large minority of repeated matchups, the lower-ranked brand is the usual winner, not the occasional one.
Two of three comparison claims named a winner
Verdicts exist because AI answers commit to them. Of the 418,813 high-confidence comparison claims in the window, 279,482, or 66.7%, declared one brand better than the other. 20.6% judged the brands equal, 10.1% stated a tradeoff, and 2.7% were unclear.
The upset analysis covers the decisive two-thirds. The equal and tradeoff claims are a real part of how AI compares brands, and our related study on head-to-head verdicts covers how often multi-axis comparisons split.
What we excluded and why
Of the 199,624 claims that named a winner and a loser, 85,880 rows, or 43.0%, survived the strict mapping: both brand names resolved unambiguously to Parse index brands with two distinct ranks. Ambiguous names, unindexed brands, self-pairs, and equal-rank pairs were dropped rather than guessed. Removing repeated claims within one answer cut 10,236 more rows, leaving 75,644 verdicts.
47.2% of verdicts carry an axis label outside the 11 named axis families; they count in every total but are absent from the axis chart. As a check, an independent recomputation from the comparison-direction field, on the row grain and a wider 119,531-row mapping, gives an upset rate of 47.8%, in the same range as the published 44.4%. Index rank comes from a single all-time snapshot computed July 13, 2026, not from a per-verdict date, and it measures AI visibility, not company size.
The GEO takeaway
The verdict layer is contestable even where the visibility ranking is not. A brand ranked thousands of places below its rival still won 44% of direct comparisons, and the upset rate barely moved as the rank gap grew. What moved it was the axis: cost upsets reached 57.7%, while market position stayed with the leader at a 29.3% upset rate.
If you are the challenger, publish concrete price and ease-of-use evidence for the exact matchups buyers ask about, because those are the axes AI already hands to smaller brands. If you are the leader, your mention volume does not defend the comparison; your integration and ecosystem evidence does. Either way, find the head-to-head questions in your category and read the verdicts, axis by axis.
Get the data
Sources
- Brandlight: The AI Search Shakeup — Why Challenger Brands Outperform $75B Giants · accessed 2026-08-29
- MaxAEO: Does ChatGPT Favor Big Brands? Evidence and 2026 Tests · accessed 2026-08-29
- Search Engine Land: Which brands are vanishing from AI search? · accessed 2026-08-29
- Kantar: Beyond Visibility — Brand building with GEO and AI Search · accessed 2026-08-29