What happens when ChatGPT searches the web?
The brand answer changes. Across 43,757 matched answer pairs, ChatGPT's searched answer named the same first brand as its no-search answer only 35% of the time. Two searched runs to the same prompt agree 46% of the time.
By Dimitry Apollonsky · August 29, 2026 · 10 min read
Contents
- Searching changed ChatGPT's first-named brand in 65% of matched pairs
- Two searched runs of the same prompt agree far more often
- The searched and no-search brand lists share 28% of their brands
- Searching trims the brand list
- The brands searching removes are widely known
- The brands searching adds skew newer and more specialized
- ChatGPT Search almost always runs exactly one search: the prompt itself
- When ChatGPT reformulates, it compresses the question into keywords
- No-search answers concentrate in May and June
- Google AI Mode almost never answers without citing
- What we counted and what we excluded
- The GEO takeaway
- Get the data
- Sources
- Related research
We analyzed 452,990 ChatGPT Search answers to 19,792 tracked prompts, carrying 2,118,975 citations, from May 24 through August 28, 2026. From them we built 43,757 matched pairs — a no-search answer and the nearest searched answer to the same prompt within 14 days — plus 90,858 rerun pairs of two searched answers as a baseline.
Searching changed ChatGPT's first-named brand in 65% of matched pairs
A no-search answer is a ChatGPT Search answer that cited no sources; a searched answer cited at least one. A matched pair is a no-search answer and the nearest searched answer to the same prompt within 14 days. The first-named brand is the brand that appears earliest in the answer text.
In 38,997 matched pairs where both answers named at least one brand, the two answers had the same first-named brand 34.5% of the time. In the other 65.5%, searching the web changed which brand ChatGPT named first. The average gap between the two answers was 4.5 days.
Takeaway
Two searched runs of the same prompt agree far more often
ChatGPT's answers vary between runs even with nothing else changed, so a raw 34.5% is not evidence by itself. As a baseline we built rerun pairs: two consecutive searched answers to the same prompts, restricted to before July 6, 2026 — the calendar weeks where nearly all no-search answers occur.
Rerun pairs kept the same first-named brand 45.9% of the time (81,859 pairs). Matched pairs kept it 34.5% of the time. The 11.3-point difference is our estimate of the retrieval effect after accounting for normal run-to-run change.
Takeaway
The searched and no-search brand lists share 28% of their brands
For each pair we measured brand-list overlap: the share of brands named in either answer that appear in both (the Jaccard index). Matched pairs averaged 27.9% overlap; rerun pairs averaged 36.1%.
In 92.3% of matched pairs the searched answer named at least one brand that the no-search answer did not (rerun baseline: 88.2%). Retrieval does not only reorder a fixed list. It swaps brands in and out.
Searching trims the brand list
Within matched pairs, the no-search answer named 5.9 brands on average and the searched answer named 5.45. The searched answer named fewer brands than its no-search partner in 45.7% of pairs where both named brands; the rerun baseline is 38.7%.
Rerun pairs show no trim at all (5.3 vs 5.3 brands). A searched answer is built around the pages it retrieved, and that appears to narrow the list slightly rather than widen it.
The brands searching removes are widely known
For every brand appearing in at least 300 matched pairs we computed an appearance ratio: pairs where the brand appears in the searched answer divided by pairs where it appears in the no-search answer. A ratio below 1 means searching removes the brand.
The lowest ratios belong to widely known brands: Brandwatch 0.21, Upwork 0.35, PayPal 0.38, Trello 0.41, Zapier 0.46, Amazon 0.48, and OpenAI itself at 0.50. These brands are strong in the model's stored knowledge — what it can produce with no retrieval — but appear less often when the answer is built from retrieved pages.
Takeaway
The brands searching adds skew newer and more specialized
The highest appearance ratios mostly belong to newer or more specialized brands: Peec AI 4.5, Grafana Cloud 1.7, SigNoz 1.6, Rippling 1.3, Atlassian 1.2, Honeycomb 1.2, Dynatrace 1.2, and Linear 1.18.
Peec AI is the clearest case: it appeared in 296 searched answers but only 65 no-search answers across its 313 pairs. A brand too new or too small to sit in the model's stored knowledge can still enter answers whenever ChatGPT retrieves pages that name it.
Takeaway
ChatGPT Search almost always runs exactly one search: the prompt itself
97.9% of answers recorded a single executed search, and in 100.0% of answers the first recorded search matched the prompt text once casing and punctuation are ignored. ChatGPT Search does not routinely plan multi-step research for these questions; it forwards the question and writes from what comes back.
8,075 answers ran two searches, 1,086 ran three, and 323 ran four or more.
Takeaway
When ChatGPT reformulates, it compresses the question into keywords
7,953 answers, or 1.8%, ran an extra search whose text genuinely differed from the prompt. In those answers, prompts averaged 13.6 words and the reformulated queries 5.0 words; 95.5% of reformulated queries were shorter than the prompt.
17.83% of reformulated queries contain the word "best". Almost none add a year (2.55%), "review" (0.03%), or "reddit" (0.00%). The reformulations look like short commercial keyword queries, not research plans.
No-search answers concentrate in May and June
45,429 answers, or 10.0% of the window, cited nothing. That share is not stable: 44.0% of May answers cited nothing, 28.1% in June, 1.9% in July, and 0.81% in August. The drop reflects how the answers were collected, not a change in ChatGPT's behavior.
This is why the study compares matched pairs rather than raw populations, and why the rerun baseline is restricted to before July 6. In the steady state since July, ChatGPT Search cites sources in about 99% of answers.
Google AI Mode almost never answers without citing
Over the same window, 1.7% of Google AI Mode answers cited nothing (7,709 of 445,163), against 10.0% on ChatGPT Search. When both engines cite, Google AI Mode attaches 15.3 citations per answer and ChatGPT Search 5.20.
Both engines are retrieval-first in the steady state. The difference is how much of the retrieved web each engine exposes per answer.
What we counted and what we excluded
Searched means the answer cited at least one source; retrieval that produced no citations is invisible to us and counts as no-search. Of 45,429 no-search answers, 205 had no searched answer to the same prompt and 1,467 more had none within 14 days; both groups were excluded. Brand sets use canonical brand identities, so product-name variants of one brand count once.
The headline held two independent ways: recomputing first-named-brand agreement from a separate storage path gave 34.8% against the primary 34.5%. It is also stable across time gaps: matched-pair agreement was 34.6% at gaps under 2 days, 35.1% at 2 to 7 days, and 32.3% at 7 to 14 days, against rerun-pair agreement of 43.1%, 46.8%, and 45.0% in the same buckets. All figures are measured in the Parse index over the stated window.
The GEO takeaway
Retrieval decides which brands enter the answer. When ChatGPT searched, it dropped brands it names freely from stored knowledge — Amazon, PayPal, Trello, Zapier — and pulled in brands the retrieved pages named, including brands too new for its stored knowledge to know.
Two moves follow. First, measure your brand's AI visibility with retrieval on, because that is the production behavior: about 99% of ChatGPT Search answers cite sources in the steady state. Second, invest in pages that get retrieved for your buyers' questions as the buyers word them — ChatGPT usually runs one search, the prompt itself, and the retrieved pages decide which brands it names.
Get the data
Sources
- OpenAI: Introducing ChatGPT search · accessed 2026-08-29
- Search Engine Land: ChatGPT search officially launches · accessed 2026-08-29
- Mallen et al.: When Not to Trust Language Models — Investigating Effectiveness of Parametric and Non-Parametric Memories · accessed 2026-08-29
- Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks · accessed 2026-08-29