Does AI keep criticizing the same brand?
Only about one in four times. When a ranked brand was criticized in either of two adjacent answers to the same prompt on the same engine, it was criticized both times in 21,346 of 82,007 comparisons.
AI repeated criticism in 26% of affected same-brand comparisons
A specific criticism appeared in both adjacent answers in 21,346 of 82,007 comparisons where either answer criticized the ranked brand, or 26.0%. The same brand had to be ranked and reviewed in both answers to the same prompt on the same engine.
SparkToro found that complete brand lists almost never repeat and said its study did not measure how each brand was described. This study adds that language layer. Treat one criticism as a review signal, not a stable brand assessment. Repeat the same prompt on the same engine before prioritizing a response.
Takeaway
Nearly three in four criticisms appeared in only one answer
Criticism appeared only in the earlier answer in 29,252 comparisons, or 35.7%; only in the later answer in 31,409, or 38.3%; and in both in 21,346, or 26.0%. The three states account for all 82,007 affected comparisons.
A single snapshot can both surface a criticism that disappears and miss one that appears on the next run. Store answer history instead of overwriting the prior state.
- Only in the later answer38.3% (31,409 of 82,007)
- Only in the earlier answer35.7% (29,252 of 82,007)
- In both answers26.0% (21,346 of 82,007)
ChatGPT Search repeated criticism more often than Google AI Mode
ChatGPT Search repeated criticism in 18,649 of 67,922 affected comparisons, or 27.5%. Google AI Mode did so in 2,697 of 14,085, or 19.1%. The ChatGPT Search rate was 1.4 times as high.
Use separate engine baselines. A blended persistence rate hides an 8.3-point difference and does not show that either engine's criticism is more accurate.
- ChatGPT Search27.5% (18,649 of 67,922)
- Google AI Mode19.1% (2,697 of 14,085)
Takeaway
Criticism repeated more often below first place
The repeat rate was 22.9% when the brand ranked first in both answers, 27.6% when it ranked second or third in both, and 29.9% when it ranked fourth or lower in both. When recommendation position changed, the rate was 22.1%.
Compare criticism at similar recommendation positions. The pattern does not prove rank caused the language, but rank changes the baseline by as much as 7.0 points.
- Fourth or lower in both29.9% (7,936 of 26,527)
- Second or third in both27.6% (5,431 of 19,651)
- First in both22.9% (2,015 of 8,796)
- Position changed22.1% (5,964 of 27,033)
The answer gap changed the repeat rate by only 1.3 points
Criticism repeated in 4,612 of 18,416 affected comparisons one to three days apart, or 25.0%; 15,327 of 58,117 four to seven days apart, or 26.4%; and 1,407 of 5,474 eight to fourteen days apart, or 25.7%.
The headline is not a simple artifact of one answer interval. The observed schedule does not support a claim about how persistence changes beyond fourteen days.
- One to three days25.0% (4,612 of 18,416)
- Four to seven days26.4% (15,327 of 58,117)
- Eight to fourteen days25.7% (1,407 of 5,474)
Industry persistence ranged from 19.8% to 30.4% in larger cuts
Among displayed industries with at least 500 affected comparisons, Apps recorded 235 of 772, or 30.4%. Health Care recorded 426 of 1,414, or 30.1%. Transportation recorded 113 of 571, or 19.8%. Three other industries complete the comparison.
Use an industry baseline to prioritize repeat checks. The breakdowns describe the observed prompt corpus and do not show that industry caused the difference.
- Apps30.4% (235 of 772)
- Health Care30.1% (426 of 1,414)
- Financial Services28.8% (1,637 of 5,687)
- Software27.1% (1,843 of 6,795)
- Real Estate22.5% (141 of 626)
- Transportation19.8% (113 of 571)
Atlassian and Datadog had the most recurring criticism
Atlassian had 437 recurring cases among 1,256 affected comparisons across 181 prompts, or 34.8%. Datadog had 355 of 1,092 across 120 prompts, or 32.5%. Eight more cleaned brand names complete the leaderboard.
Use the table as a review-volume list, not a brand-quality or accuracy ranking. Start with the underlying prompts and claims for high-volume brands.
Takeaway
Stricter panel checks did not reverse the result
The main rate was 21,346 of 82,007, or 26.0%. Keeping the same extraction contract in both answers returned 13,332 of 46,795, or 28.5%. Keeping the same recommendation position returned 9,323 of 32,126, or 29.0%. Limiting the gap to seven days returned 19,939 of 76,533, or 26.1%. Exact criticism wording repeated in only 625 of 21,346 recurring cases, or 2.9%.
The conclusion survives narrower panels, but the headline measures the presence of a specific criticism field, not identical wording or human-rated semantic equivalence. The method excludes 16,828 ambiguous prompt-and-engine timestamps, brands missing from either answer, unranked mentions, unreviewed language, and unresolved product identities. It does not test truth, source support, causation, or buyer opinion.
What marketers should do
Criticism repeated in 21,346 of 82,007 affected comparisons. It appeared in only one answer in 60,661 cases, or 74.0%, and the engine rates differed by 8.3 points.
Repeat priority prompts on each engine. Store the brand, prompt, recommendation position, and specific criticism for every answer. Separate one-run criticism from recurring criticism. Review source support and factual accuracy before changing messaging. Rerun the fixed panel next quarter before calling a difference movement.
Takeaway
How we measured
In the Parse index, we analyzed 1,627,442 reviewed brand statements across 741,493 AI answers, producing 550,093 matched same-brand comparisons covering 45,143 brands and 18,142 organic prompts on ChatGPT Search and Google AI Mode from May 24 through August 19, 2026.
Get the data
Sources
These are the pages this study used.
- SparkToro: AI brand recommendation consistency research · accessed August 23, 2026
- BrightEdge: Negative brand sentiment by engine · accessed August 23, 2026
- Semrush: AI Visibility Brand Performance Reports · accessed August 23, 2026
- Carnegie Mellon: Estimating LLM consistency · accessed August 23, 2026