Canlah Research · Open case study · CC BY 4.0
Same Question, Seven Engines: what each AI search surface actually cited
We asked seven search/answer engines the same commercial question — in English and Chinese, on one night — captured every page they cited, classified each one with a frozen codebook, and cross-checked what we observed against each engine's own documentation.
Probed 2026-08-29 · v1.0 published 2026-08-30 · v1.1 (logged-out ChatGPT replication) 2026-08-31 · Singapore egress
Four numbers
-
0
pages cited by 5 or more of the 7 engines (the best-covered pages reached 4)
-
0.171 / 0.091
mean pairwise top-10 domain Jaccard (Chinese / English)
-
116
citation records (80 raw URLs, 71 unique pages after normalization, all fetch-attempted and archived)
-
53 / 3 / 0
of 56 documented-mechanism claims after adversarial verification: confirmed / softened / refuted
Seven engines, seven sourcing appetites
The dominant source shape observed in this run, and the working hypothesis to test first if you target that surface (hypotheses, not conclusions — see limits below).
| Engine | Dominant source shape (this run) | Hypothesis to test first |
|---|---|---|
| Google organic SERP | vendor service pages (13/18) | an exact-match, locally-anchored service page may be the primary entry asset |
| Google AI Mode | service pages + entity/Places material | entity-layer work (business profile, reviews, locality) may matter more than page-shape work |
| Gemini + Search grounding (API) | geospatial/GIS pages (16/16 — query misread) | explicit disambiguation may be a precondition for appearing at all |
| OpenAI web_search (API) | pre-made “Best X” listicles (6/9) | a methodology-backed ranked list may be the highest-leverage asset |
| ChatGPT consumer (logged in) | deep clusters from 6 vendor domains (17 citations) | a deep, internally consistent page cluster may outperform a single optimized page |
| Exa (neural search API) | vendor capability pages (18/20) | full-sentence service descriptions may be the unit of optimization |
| Tavily (RAG search API) | widest mix incl. directories, media, social | breadth of third-party presence may matter more than any single owned asset |
What this run found
- No shared source layer: of 71 unique pages, none was cited by 5 or more engines; the three best-covered pages reached 4 each.
- Sources diverged, brands recurred: engines cited largely different URLs while naming overlapping vendor sets. One hypothesis consistent with the data: cross-engine presence travels through a second-order mention network — many publishers repeatedly naming the same entities — not through any single high-ranking page.
- Self-published “Best X” listicles were taken up by exactly three surfaces: OpenAI web_search (3/9), ChatGPT consumer (6/17) and Google organic (4/18 — top-10 entries, never #1); zero uptake by Exa, Tavily, AI Mode or Gemini. No engine visibly filtered the conflict of interest: every self-listicle ranks its publisher first, and was cited anyway.
- Structured data correlated with citation and guarantees nothing: 17/17 ChatGPT-cited and 16/18 Google-cited records carried JSON-LD — but AI Mode and Gemini also cited schema-free pages, bot-challenged pages and one 404.
- Word-sense disambiguation decided one engine entirely: Gemini resolved “GEO” to geospatial/GIS in both languages (16/16 sources), including freshly updated pages — freshness did not rescue the lost word-sense.
v1.1 replication: logged-out ChatGPT is a different engine
Two days later we re-ran both prompts logged out (incognito, no account, web search on): 3 Chinese and 5 English answers. Every logged-out answer opened with a map and 11–15 rated business cards — a Places-style layer the logged-in probe never showed, and in 4 of 8 runs those cards included geotechnical/geospatial firms (Gemini's word-sense leak, now inside ChatGPT's entity layer). The prose shortlists were far more stable than the cited pages had been: run-to-run brand Jaccard 0.70 (EN) / 0.54 (CN). Anyone measuring “ChatGPT visibility” should state which surface they measured.
Limits — read before quoting any number
- n = 1 query × 1 night × 1 locale × 2 languages. A forensic case study with a rerunnable method — not a statistical survey; AI answers vary run to run, so every number is one snapshot.
- Traditional engines contribute result lists; answer engines contribute citation lists — related, not identical, comparison units.
- Alignment between documented mechanisms and this run's output is stated as consistency, never causation.
- The logged-in ChatGPT probe could not rule out personalization; the v1.1 logged-out replication partially retires that caveat (English side), but is itself limited — repeated submissions inside one conversation.
Disclosure
Canlah Research is the research arm of Canlah AI Pte. Ltd., a Singapore GEO vendor — we are a participant in the market this study measures. In this run, “Canlah AI” was named by two engines and canlah.ai was cited by one (the logged-in probe carrying the personalization caveat; in the v1.1 logged-out replication: 3/5 English runs, 0/3 Chinese). It did not appear in the other engines. We publish the losses with the wins — and a fixed classifier bug — with every label, matrix and verification verdict in the open repository, so every derived statistic can be recomputed without trusting us.
Rerun it for your category
The method holds for any category: pick one question your buyers actually ask, ask it on every engine they actually use, capture the same night, fetch and classify every cited page with the frozen codebook in the repository (method/probes.md, scripts/classify.py), then read which page shapes each engine took up.
Cite as
Canlah Research (2026). Same Question, Seven Engines: What Each AI Search Engine Cited — and Working Hypotheses for Targeting Each. https://canlah.ai/lab/seven-engines/