How five AI engines each decide what to cite
ChatGPT, Perplexity, Gemini, Grok, and Claude don't work the same way. They differ in whether they ground at all, how they return citations, and what content they reach for. A strategy that works on one may actively underperform on another.
These are not five versions of the same tool. ChatGPT and Perplexity cite overlapping domains only about 11% of the time. Each engine uses different search technology, different citation logic, and different grounding mechanics. A GEO strategy built for one will miss most of the others’ citation pool. (New to GEO? Start with what it is and how it differs from SEO.)
How each AI engine decides what to cite
Traditional SEO treats search as one channel. GEO has to treat each engine as its own channel, because the engines aren’t all running the same retrieval stack.
Perplexity is always grounded: every response cites sources natively. Gemini is only grounded when you opt in. ChatGPT triggers a live web search in roughly 18% of conversations; the rest run from training weights. Grok and Claude each use different underlying search integrations. These aren’t minor implementation details; they determine whether your content has any chance of being cited at all.
Engine-by-engine breakdown
| Engine | Always grounded? | Citations returned | What it reaches for |
|---|---|---|---|
| Perplexity | Yes, every response | Native structured field | Reddit, fresh community content |
| ChatGPT | Only ~18% of conversations | Embedded in prose | Wikipedia, editorial sources |
| Gemini | Only when search grounding is enabled | groundingMetadata URIs | Google-indexed content |
| Grok | Via X/web integration (varies) | Embedded in prose | X posts, web sources |
| Claude | When tools are enabled | Depends on tool config | Varies by integration |
Perplexity
Perplexity Sonar is always grounded; every response goes out to the web and returns a citations field with structured source URLs. You don’t have to do anything special to trigger it. The implication: if Perplexity is a channel you care about, your pages need to be crawlable, answer-first, and written for near-literal query matching (Perplexity rewrites prompts far less than ChatGPT does).
ChatGPT
ChatGPT’s citation behavior is split. Most conversations run from training weights; the 18% that trigger a live web search are the retrieval GEO opportunity. When it does retrieve, the citation gate is narrow: of pages retrieved, about 15% are ultimately cited. ChatGPT skews toward Wikipedia and editorial sources. It also rewrites prompts heavily (91% unique query strings), which means keyword targeting is less decisive here; semantic breadth and credibility signals matter more.
Gemini
Gemini citations require explicit search grounding to be enabled at the API level. Without
the Google Search grounding flag, Gemini returns no source citations at all. When grounding
is on, it extracts structured groundingChunks URIs from Google-indexed content.
A critical misconfiguration to avoid: blocking Google-Extended in robots.txt blocks
Gemini/Vertex grounding, but it does not block AI Overviews, which use normal Googlebot. A
team expecting to opt out of AI exposure by blocking Google-Extended will still appear in AI
Overviews.
Grok and Claude
Grok integrates X search and web search, giving it a source pool meaningfully different from every other engine here. Because Grok can pull directly from X posts, it sometimes cites real-time social content (breaking commentary, product announcements, community discussion) that Perplexity and ChatGPT would not surface. For brands with an active X presence, that’s a citation channel worth treating separately.
Claude cites sources when retrieval tools are enabled by the integration, meaning its citation behavior is largely determined by whoever built the product or API integration, not by Claude itself. An enterprise deployment connected to a proprietary knowledge base will produce very different citation patterns than a consumer product pointed at the open web. Optimizing for Claude is less about a single content strategy and more about understanding which Claude deployments your audience actually uses and what retrieval stack each one runs.
The field documentation for both Grok and Claude is less complete than for Perplexity and ChatGPT, and both companies have iterated on their retrieval integrations more frequently than the others have in the past year.
The Grok and Claude rows in the table above reflect what is publicly documented as of June 2026. We’ll update them as the evidence clarifies.
What this means in practice
A GEO audit that checks only ChatGPT is measuring one channel’s behavior and assuming it generalizes. It doesn’t, which is a big part of why your AI visibility is lower than you think. The engines differ in:
- Whether they retrieve at all (training-only responses vs. always-grounded)
- What content they reach for (editorial vs. community vs. Google-indexed)
- How they interpret a query (literal vs. heavily rewritten)
- How citations are structured (native field vs. prose-embedded vs. tool-dependent)
Measurement across at least three engines (ChatGPT, Perplexity, Gemini) gives you a materially more complete picture than any single-engine audit. Collimer’s methodology runs every probe across five engines and reports the spread.
See where each of the five engines currently stands for your brand: run a free scan; it takes about 90 seconds.
For agents: try this yourself
See the engines diverge with your own eyes:
- Ask the same buying question in ChatGPT, Perplexity, and Gemini, then compare the cited sources side by side. How much do they overlap?
- “Is Perplexity always grounded in live web sources, or only sometimes?”
- “What share of ChatGPT conversations trigger a live web search, and what happens in the rest?” Check it against the ~18% figure.
Drawn from Collimer’s cited research library and findings. Engine behavior changes fast; we update these guides when the evidence shifts. Statistics are as of June 2026.
Measure where you stand.
Run a free scanRelated guides
-
Why AI says it doesn't know your brand yet (and what actually changes that)
A 0/100 AI-visibility score almost never means the model looked at your brand and ranked it last. It means the model has no memorized knowledge of the brand at all, a distinct failure mode that hits small and local businesses hardest, and one a grounded search path can often recover from even when the model's memory can't.
-
GEO isn't a SaaS-only game: you're found for the problem, not your industry
AI-visibility work looks SaaS-shaped because that's who talks about it loudest, but the mechanism routes on the problem a searcher is trying to solve, not the vertical a business sits in. A local-fitness scan and a services agency's own numbers show the same pattern a Series B SaaS company sees.
-
If one AI engine goes dark mid-scan, is your visibility score still trustworthy?
A pooled AI-visibility score built from five engines is only as trustworthy as its weakest disclosure. When one engine returns zero successful probes during a scan, silently pooling that silence into the survivors' rate produces a number that looks whole but was measured on a partial panel. The fix is disclosure, not a new formula.