Skip to main content
CANLAH AI
Try
SEO/GEO 2026-08-30 · 9 min read

How AI Engines Choose Their Sources: ChatGPT vs Perplexity vs Gemini vs AI Overviews

The four engines that decide your AI visibility do not pick sources the same way. A mechanism-level comparison of ChatGPT Search, Perplexity, Gemini and Google AI Overviews — with one concrete action for each.

Four dark panels of identical text rows, each lit with a different number of indigo citation marks, the last with none

TL;DR

"Optimizing for AI search" is not one job. Google AI Overviews composes from the Google index it already has. ChatGPT Search runs a live retrieval pass with its own crawler. Perplexity indexes independently and cites more densely than anyone. Gemini sits on Google's index but is gated by a separate crawler control. The same page can be quoted by one and ignored by the other three — which is why you pick an engine to fight before you pick a tactic.

Most GEO advice treats "AI engines" as a single audience. It is not. Four systems currently decide whether your brand shows up in an AI answer, and they disagree with each other constantly — about which sources are worth reading, how many to cite, and how fresh a page has to be before it counts.

If you have ever seen your brand named confidently by Perplexity and completely absent from Google AI Overviews on the same question, this is why. Below is what each engine actually does to select sources, and the one change that moves the needle for each.

The four engines side by side

  Google AI Overviews ChatGPT Search Perplexity Gemini
Where sources come from The existing Google search index Live retrieval, own crawler plus third-party search Its own independent index Google index, plus grounding at answer time
Crawler to allow Googlebot OAI-SearchBot, ChatGPT-User PerplexityBot, Perplexity-User Googlebot, gated by Google-Extended
Citation density Low — a handful of links, often collapsed Moderate — inline, one per claim High — numbered sources on nearly every sentence Low to moderate, varies by surface
Freshness weight Follows normal search freshness signals High on time-sensitive queries Highest — recency is a visible ranking factor Moderate
Reads the full page? Often composes from title, meta and first screen Usually fetches the page Fetches and extracts passages Mixed
Blocking it costs you All Google traffic — never block ChatGPT citations only Perplexity citations only Gemini grounding, not search rank

Read the last row carefully, because it is the one most sites get wrong. Google-Extended and GPTBot are commonly disallowed by legal or IT teams who see "AI crawler" and reach for the block. That decision does not protect search rankings — it removes you from the answer layer while leaving your rankings exactly where they were.

Google AI Overviews: it already decided, before you asked

AI Overviews does not go looking for new sources. It composes an answer from pages Google has already indexed and already ranks well for related queries. There is no separate "AI Overviews index" to get into.

The practical consequence is that Overviews frequently assembles an answer from titles, meta descriptions and the first screen of a page, without the full document ever being weighed. A page whose answer is buried in section four is a page Overviews will summarize badly, or skip in favour of a competitor who said it in line one.

The one action: put a 40–60 word direct answer in the first screen of every page that targets a question. Not a hook, not a brand promise — the literal answer, phrased so it survives being lifted out of context. Everything else on the page can stay as it is.

More on how this surface has reshaped click behaviour in how AI Overviews are changing SEO in 2026.

ChatGPT Search: retrieval happens at question time

ChatGPT with search behaves differently from the model answering out of memory. When it searches, it runs a live retrieval pass, fetches candidate pages, and writes with those pages open. Two crawlers matter: OAI-SearchBot, which builds the search surface, and ChatGPT-User, which fetches a page in response to a specific user request.

Because retrieval is live, a page published this week can be cited this week. But because retrieval is also selective, it tends to favour pages that resolve the question completely over pages that partially cover ten questions. Long, unfocused pillar pages underperform here relative to how they perform in classic SEO.

The one action: check that OAI-SearchBot and ChatGPT-User are not disallowed in robots.txt. This is a two-minute check that silently decides whether the rest of your work is even visible. Then make sure each key page answers exactly one buyer question end to end.

If your brand is missing here specifically, the four usual causes are laid out in why your business doesn't show up in ChatGPT.

Perplexity: the most generous citer, and the most demanding

Perplexity runs its own index through PerplexityBot and cites more densely than any other engine — numbered sources attached to individual sentences rather than a link cluster at the end. For a brand trying to earn AI visibility, this is the highest-leverage surface available: more citation slots per answer means more chances to be one of them.

The trade is that Perplexity weights recency heavily and extracts at passage level rather than page level. It is not asking "is this a good page?" so much as "is this specific paragraph the best available statement of this specific fact?" Pages that are structurally scannable — clear headings, self-contained paragraphs, one claim per block — get extracted. Walls of prose do not.

The one action: restructure your top pages so every paragraph stands alone as a quotable unit. If a paragraph only makes sense after reading the two before it, it will not be extracted.

The full playbook is in how to get your brand cited by Perplexity AI.

Gemini: Google's index, a different gate

Gemini draws on Google's index, which means the SEO work you have already done carries over. The complication is Google-Extended: a separate control that governs whether your content can be used for Gemini's grounding and training, independent of whether Googlebot can crawl you for search.

This is the single most common own-goal in the whole category. A site can rank on page one of Google and be structurally invisible to Gemini, because someone added one line to robots.txt two years ago. Nothing in Search Console will flag it.

The one action: open your robots.txt and confirm Google-Extended is not disallowed. If it is, that was a decision someone made — make sure it is still the decision you want.

What all four reward

The mechanisms differ, but three signals pay off across every engine:

  • Answer-first structure. The claim before the context, in the first lines, phrased to survive extraction.
  • Corroboration off your own site. Engines weight what other sources say about you. A claim that exists only on your homepage is a claim with one witness.
  • Explicit entity facts. What you are, where you operate, who you serve — stated plainly and consistently, and marked up with structured data so it does not have to be inferred.

Note what is absent from that list: keyword density, word count, and llms.txt. The last one draws a disproportionate amount of attention for something no major engine currently uses to decide citations — see what llms.txt actually does.

Which engine should you fight first?

Not all four at once. Pick by where your buyers actually are:

  • High-volume consumer categories — restaurants, retail, local services: start with AI Overviews. It sits on the searches your buyers already run.
  • Considered B2B purchases — software, agencies, professional services: start with ChatGPT. Buyers use it to build shortlists before they ever search.
  • Research-heavy or technical categories: start with Perplexity. Its users are comparing, and its citation density gives you the most room to appear.
  • Already strong in Google organic: check Google-Extended first. You may be one line away from Gemini visibility you have already paid for.

Why checking once tells you nothing

A caution before you go and test any of this by hand. Asking an engine the same question twice returns an identical answer less than 1% of the time, and across repeated runs of the same prompt, the set of brands named overlaps only about 45–59%.

So a single check proves nothing, and a month-over-month comparison built on one ask per month is measuring noise. Any claim that a change "improved AI visibility" needs repeated sampling across engines under a fixed protocol, or it is a story rather than a measurement.

Frequently asked questions

If I optimize for one engine, do the others improve too?

Partially. Answer-first structure and off-site corroboration help everywhere. Crawler access, freshness handling and passage extraction are engine-specific and do not transfer.

Does blocking AI crawlers protect my content?

It reduces the chance your content is used, and it guarantees you are not cited. Those are the same setting. Decide which one you care about more.

Is Google-Extended the same as blocking Googlebot?

No. Google-Extended governs Gemini grounding and training. Blocking it does not affect search rankings — and allowing it does not improve them.

How long before changes show up in AI answers?

Live-retrieval engines can reflect a change within days. Index-dependent surfaces move on crawl and index cycles, so weeks. Neither is instant, and neither is reliably observable without repeated sampling.

Do I need a different page for each engine?

No. One well-structured page can serve all four. What differs is crawler access and which pages you prioritise — not the content itself.

Where to start

Two things, in order. Check your robots.txt for OAI-SearchBot, PerplexityBot and Google-Extended — that is fifteen minutes and it decides whether anything else matters. Then measure where you actually stand across engines before you change a single page, so you have a baseline to compare against.

You can run that baseline for free: our AI-visibility snapshot probes multiple engines against real buyer questions and returns the records — screenshots and timestamps included — so the result is something you can re-verify rather than take on faith. The wider practice it belongs to is described on our GEO service page.

FREE WHITEPAPER

Marketing in the Agent Era

All 13 chapters public — no email wall. Includes an original dataset on the agent-readiness of 50 cross-border DTC brands.

Read it free →