What happened
Semrush and Kevin Indig ran 100 prompts through GPT-5.2 twice — once in Instant mode (minimal reasoning), once in Thinking mode (high reasoning) — covering 20 buyer journeys across B2B SaaS, finance, consumer tech, and health/lifestyle. The headline: only 25.6% of cited domains overlapped between modes for the same prompt. Roughly three-quarters of the sources ChatGPT cites change when the reasoning dial moves.
Thinking mode works harder across the board. It cited sources in 68% of responses versus 50% for Instant, averaged 4.5 citations per response versus 2.6, and fired 4.6x more sub-queries — 1,130 total web searches across the test versus 245 in Instant mode. It behaves less like autocomplete and more like a researcher fanning out and cross-checking.
And it reads different material. Reddit's citation share fell from 15% in Instant mode to 7% in Thinking mode, with UGC and review sites dropping from 14.3% to 6%. Government and academic sources quadrupled, from 1.9% to 8.8%, and official documentation rose from 12.4% to 17.5%. Brand citations held roughly steady — 62.4% versus 60.6%.
Why this matters
You're fighting on two battlefields, and most stores only track one. Instant mode leans on community sentiment — Reddit threads, review sites, the fast consensus. Thinking mode acts like a fact-checker: it fans out into sub-questions, cross-references, and reaches for pages that read like documentation. A shopper doing serious research on a considered purchase is far more likely to be in the second mode, and that's where your spec pages, buying guides, and policy pages either hold up under scrutiny or vanish.
The stable brand-citation share is the quiet good news. Brands weren't squeezed out of deep-research answers — the bar just changed. Factual, verifiable, well-structured pages earn Thinking-mode citations; clever ones don't. That's a content brief, not a mystery.
Keep the skepticism proportionate: 100 prompts is a small sample, it's one model version, and OpenAI revises retrieval constantly. Trust the direction — deeper reasoning shifts citations toward authoritative sources and fans out into sub-queries — and don't build a strategy on the second decimal place.
What to do about it
Track your AI visibility in both modes
Next time you spot-check whether ChatGPT recommends your store, run the same buyer prompt in Instant and in Thinking mode and log both results. With a 25.6% domain overlap, one check tells you almost nothing about the other. Two columns in the same spreadsheet you already keep.
Build one documentation-grade page per money category
Thinking mode reaches for pages that look like official documentation: spec tables, sourced claims, dates, clear structure. Take your highest-revenue category and give it one page that would survive a fact-checker — that's the format deep reasoning cites at 17.5%.
Cover the sub-questions, not just the head term
Thinking mode fires 4.6x more sub-queries, spanning comparison and validation questions around the main prompt. Map the follow-up questions a careful buyer asks about your category — versus alternatives, sizing, returns, compatibility — and make sure a page answers each one directly.
Don't abandon the community lane
Reddit still carries 15% of Instant-mode citations, and Instant is where quick answers get formed. Keep earning genuine review and community presence — just stop treating it as your whole AI-visibility strategy, because it fades to 7% the moment the buyer thinks harder.