Does question wording change which brands AI names?
Yes, a lot. Two phrasings of the same buyer question shared 11.7% of named brands on average. The identical phrasing re-asked one to three days later shared 38.0%. Nearly half of phrasing pairs shared zero brands.
By Dimitry Apollonsky · August 29, 2026 · 9 min read
Contents
- Rewording a question cuts brand overlap to a third of the rerun level
- 45% of phrasing pairs shared zero brands
- The whole distribution shifts, not just the average
- The first-named brand matched in 13% of phrasing pairs
- On 5% of question-days every phrasing led with the same brand
- One phrasing shows about a third of the brand universe
- ChatGPT Search holds slightly more of its list across wordings than Google AI Mode
- Scenario-style phrasings diverge most
- How much the wording differs barely matters
- The effect holds question by question
- What we excluded and why
- The GEO takeaway
- Get the data
- Sources
- Related research
We analyzed 120,749 brand-naming answers to 1,677 phrasings of 340 tracked buyer questions on ChatGPT Search and Google AI Mode, July 1 through August 26, 2026: 497,084 same-day pairs of differently phrased answers, compared against 106,823 same-phrasing rerun pairs.
Rewording a question cuts brand overlap to a third of the rerun level
A phrasing pair is two answers from the same engine on the same day to two different phrasings of the same underlying buyer question. A rerun pair is two answers to the identical phrasing on the same engine, one to three days apart. Brand overlap is the share of brands the two answers have in common, out of all brands either answer names (the Jaccard index over distinct named brands).
Across 497,084 phrasing pairs, mean brand overlap was 11.7%. Across 106,823 rerun pairs, it was 38.0%. Re-asking the identical question already changes most of the list; changing the wording cuts the remaining agreement to about a third of that level. The rerun number is the noise floor, so the wording effect is real and large, not a measurement artifact.
Takeaway
45% of phrasing pairs shared zero brands
In 45.3% of phrasing pairs, the two answers to the same buyer question had not a single brand in common. For rerun pairs the zero-overlap rate was 8.9%.
Complete disagreement is the modal outcome for reworded questions, five times as common as for reruns. Identical brand lists were near-impossible in both directions: 0.11% of phrasing pairs and 4.9% of rerun pairs matched exactly.
The whole distribution shifts, not just the average
Bucketing every pair by brand overlap shows two different shapes. Phrasing pairs pile up at the bottom: 84.4% of them overlap 25% or less. Rerun pairs center in the middle: 39.4% of them overlap between 25% and 50%.
The headline averages are not driven by a few extreme questions. Rewording moves the entire distribution toward zero.
| Zero | 45.3% | 8.9% |
| Up to 25% | 39.0% | 26.9% |
| Identical lists | 0.11% | 4.9% |
| 75% to under 100% | 0.26% | 3.6% |
| 50% to 75% | 2.2% | 16.3% |
| 25% to 50% | 13.1% | 39.4% |
The first-named brand matched in 13% of phrasing pairs
The first-named brand is the brand that appears earliest in the answer text. Two phrasings of the same question opened with the same brand in 13.1% of pairs. The identical phrasing re-asked opened with the same brand in 49.3% of pairs.
The top of the list is even more wording-sensitive than the list as a whole. Whoever leads the answer for one phrasing usually does not lead it for the next.
On 5% of question-days every phrasing led with the same brand
We grouped answers by question, engine, and day: 19,495 question-engine-days had two or more phrasings answered. On 5.3% of them, every phrasing opened with the same first-named brand. Restricted to the 15,360 question-engine-days with three or more phrasings, the all-agree rate fell to 2.5%.
There is rarely a single brand that "the answer" to a buyer question opens with. Which brand leads depends on how the question is put.
One phrasing shows about a third of the brand universe
The average answer named 6.4 distinct brands. Across all phrasings of the same question answered by the same engine on the same day (6.0 phrasings on average), the engine named 26.5 distinct brands. A single answer covered 34.1% of that day's brand universe for its question, on average.
The brands an engine associates with a buyer question form a pool roughly four times larger than any one answer shows. Each phrasing draws a different sample from that pool.
Takeaway
ChatGPT Search holds slightly more of its list across wordings than Google AI Mode
Mean brand overlap on phrasing pairs was 13.2% on ChatGPT Search (247,803 pairs) and 10.2% on Google AI Mode (249,281 pairs). Zero-overlap rates were 43.1% and 47.5%.
The rerun baselines ran the other way: 37.0% on ChatGPT Search and 39.1% on Google AI Mode. Google AI Mode repeats itself slightly more but reacts to wording slightly more. The wording effect is large on both engines, so this is a difference of degree, not behavior.
Scenario-style phrasings diverge most
We split phrasings by whether they describe the asker's situation in the first person ("we are looking to switch providers") or ask neutrally ("what are the top alternatives"). Pairs of two neutral phrasings overlapped 16.86% on average (47,134 pairs). Mixed pairs overlapped 12.97% (107,235 pairs). Pairs of two first-person phrasings overlapped 10.53% (342,715 pairs), with a 48.48% zero-overlap rate.
Adding situational detail steers the engine harder than syntax does. Two buyers describing their own version of the same problem get more different brand lists than two buyers asking the textbook question.
How much the wording differs barely matters
We compared each pair's word-count gap. Pairs of phrasings within five words of each other overlapped 12.0% (283,351 pairs). A gap of 6 to 15 words gave 11.2% (195,207 pairs). A gap of 16 words or more gave 11.0% (18,526 pairs).
The penalty comes from rewording at all, not from how far the surface form drifts. Even near-identical-length rephrasings of the same question landed near the 11.7% average, far below the 38.0% rerun level.
The effect holds question by question
Averaging within each question first: 270 questions had at least 20 phrasing pairs. Their question-level mean brand overlap was 13.9%, with a median of 12.0%, close to the pooled 11.7% headline. Only 7 of the 270 questions reached the 38% rerun level; 58 averaged below 5%.
The headline is not a blend of stable questions and a few chaotic ones. For 97% of questions, wording moved the brand list more than re-asking did.
What we excluded and why
Answers naming zero brands were excluded from both sides: 19,112 of 139,861 results in the window, or 13.7%. We kept one answer per phrasing, engine, and day (the latest), so same-day retries never count as pairs. Phrasing pairs compare answers from the same day; rerun pairs span one to three days, and 97.7% of them span exactly one day, where overlap was 38.3%, so the small time gap does not explain the difference.
The headline was computed three ways: the pair-level mean (11.7%), the pooled micro-average weighting large brand lists more (10.4%), and the question-level mean (13.9%). All three sit far below every rerun estimate. Phrasing variants are generated to restate the same underlying buyer question; they can shift emphasis or add situational detail, which is the behavior buyers exhibit and the phenomenon measured here. This is measured in the Parse index over the stated window, on brand identity only; it does not measure whether any answer was correct.
The GEO takeaway
Any single prompt is a weak measurement instrument. Two phrasings of the same buyer question shared 11.7% of named brands, the leading brand matched 13.1% of the time, and one answer showed 34.1% of the brands the engine tied to that question that day. A rank tracked on one phrasing will swing for reasons that have nothing to do with your visibility.
Track the question, not a phrasing: monitor a spread of wordings, including first-person scenario versions, which diverge most. Judge visibility by your share of the union of brands across phrasings, not by presence in one answer. And treat the 38.0% rerun overlap as the ceiling: even a perfect wording panel re-measures against an engine that changes its own answer from day to day.
Takeaway
Get the data
Sources
- SparkToro: AIs are highly inconsistent when recommending brands or products · accessed 2026-08-29
- Search Engine Land: AI recommendation lists repeat less than 1% of the time · accessed 2026-08-29
- Search Engine Journal: Does prompt variance impact brand mentions? (Peec AI study) · accessed 2026-08-29
- Salinas & Morstatter: The butterfly effect of altering prompts (arXiv) · accessed 2026-08-29