Skip to main content

GEO Playbook · General Edition · Chapter 1 (2 of 17)

How AI picks its answers: what buyers ask, and which pages ChatGPT and Google's AI cite

AI reads pages live, the two engines cite different pages, and three things decide whether you get named

1.1 How buyers ask, and what AI reads

What you'll do in this section: understand where the paragraph AI gives the buyer comes from. AI fetches a batch of pages live, reads only the text visible on those pages and does not execute JavaScript. After reading this, you'll know why the later chapters fix the door first, and why the main battlefield is not your website.

Example (dental) Buyers no longer scroll through ten blue links. They open ChatGPT and ask directly: "新加坡哪家做隐形矫正好?" (which clinic in Singapore is best for clear aligner treatment?) AI replies with a paragraph that names a few clinics.

receives the question

reads only

writes from this

names a few businesses

Buyer asks a question

Fetches a batch of pages live

Visible text on the page

Replies in one paragraph

Names per answer: single digits, own tests only

Figure: AI does not answer from memory. It fetches pages live and reads only visible text; the businesses the answer names make up its total names

How many names an answer holds is not something someone else's industry report can tell you. Trust only the number you count when you measure your own set of questions. On the same set of questions, some answers name 3 businesses, some name 8, and some name none at all (they only explain the basics and tell you to "consult a professional").

After the buyer has a name

  1. ①See the namesThe answer names a few businesses; about 1% of visits click the links in the summary (US sample)
  2. ②Search againThey skip the links and search the name directly
  3. ③Go to the websiteCheck price, process, credentials
  4. ④Check profilesReviews, directories, business profiles
  5. ⑤Get in touchThis is where the sale happens
Figure: Buyers rarely click the links in the answer. Once they have a name they go back and check it themselves, so pick "who to buy from" questions, not educational questions

Where citations come from: the main battlefield is not your website

earned media (others writing about you)84%Muck Rack
of which news coverage27%included within the 84
paid placements and advertorials0.3%Muck Rack
Figure: Most AI citations come from pages other people write about you; paid placements and advertorials are close to zero

a minority

the vast majority

only covers

main battlefield

not

The pages AI fetches live

Your website

Pages other people write about you

Fixed prices, process, credentials, org facts

Media, associations, gov registers, vendor certs

Paid listings, advertorials

Figure: Your website only covers the cells third parties cannot fill; the main off-site battlefield is media and institutions, not paid listings

What AI reads: visible HTML, no JavaScript

Do not execute JS (zero evidence of execution in a large sample)Render pages (exceptions)
GPTBot, OAI-SearchBot, ChatGPT-UserGemini: fully rendered through Google's infrastructure
ClaudeBot, Claude-UserApplebot: a browser-style crawler
PerplexityBot, Bytespider—
Meta-ExternalAgent—
Price written only in JS: not a single character comes throughPrice written only in JS: it comes through
Figure: Mainstream AI crawlers do not execute JavaScript; only two exceptions render pages

Three things later chapters take as given

  1. Before answering, AI fetches pages live and reads only visible HTML: it does not read JSON-LD, and mainstream retrieval crawlers do not execute JavaScript (the exceptions are Gemini and Applebot).
  2. Most citations come from pages other people write about you (earned media, 84%); paid placements and advertorials make up 0.3%. Your website only covers the cells third parties cannot fill.
  3. The order is door × identity × shelf (1.3), and progress is measured only with the three rulers plus the noise band from same-week retests (1.3, → General Edition 7.1 What a retest produces, and the rulers). The other premises (citation density by page type, the correlation coefficient for video, the share of price questions that trigger AI Overviews, annual opportunity value) are given in the section where each is used.

1.2 Two legs: ChatGPT looks for the source, AI Mode for second-hand summaries

What you'll do in this section: tell apart which pages each leg looks for and which ruler measures each one. After reading this, you'll have one more habit: before you write any sentence, ask "which leg is this sentence for?"

ChatGPT legAI Mode leg
Looks for the source: government pages, official pricing pages, product pagesLooks for second-hand summaries: peer articles, lists, videos, Google cards
Takes the source facts and writes its own conclusionCites pages other people have already summarised
Feeds on own sentences: you as the subject, exact amount and unitFeeds on market sentences: range + variables + source and date
Figure: The two legs do not look for the same kind of page, and they want different kinds of sentences

In our measurements, the direction was the same in all five industries. For what each leg cites and in what share, industry by industry, see → General Edition A.3 Page-type measurements (1): two legs, seven rules, the nine tier-A types.

Example (B2B SaaS) Official pricing pages make up about 53% of ChatGPT's citations, plus about 7 citations of help centres, and then ChatGPT writes its own conclusion. For AI Mode, YouTube, Reddit and LinkedIn together make up 40%, and the rest are peers' "vs" pages and lists.

Three rulers

The figure above answers "which pages to write, which sentences to write"; the one below answers "what to measure with". In the page-type tests, the AI Mode leg is used only to see which kinds of page it cites. It is not used as a monthly ruler.

measures

measures

only uses

can it fetch you

calibrates once a month

calibrates

not added

not added

not added

OpenAI model API

Model-API visibility baseline · OpenAI leg

Gemini model API

Gemini model leg

AI Overviews and AI Mode

Search Console generative AI report impressions

check Googlebot

ChatGPT web version

web control leg, not used for acceptance

one combined visibility score

Figure: Three systems, three rulers: each measures its own thing, and they are never added together
  • In reports, the ruler legs may only be called this: "Model-API visibility baseline · OpenAI leg" and "Gemini model leg", printed in the footer of every report. Never write "what users see in ChatGPT" or "real user visibility". Citation distributions on the API and on the web version show a measured, systematic skew. It is not noise, and running more rounds does not smooth it out. For how large the skew is and how to run the web control leg, see → General Edition 4.4 The web control leg and the frozen baseline (the book's only full spec).
  • The naming rule is not about wording: get the name wrong, and a whole quarter's off-site effort gets aimed at the API leg's top 20, while buyers may be seeing a different set of pages.
  • Gemini API ≠ AI Overviews ≠ AI Mode. They are three different products, with different fetch paths, different ways of generating answers and different source pools. Passing off Gemini model leg numbers as visibility in AI Overviews means measuring A and using it to sign off B.
  • The Search Console ruler has impressions only, with no times named and no position. It is already included in total impressions, so it must not be added to total impressions.
  • For why you check Googlebot rather than Google-Extended, see → General Edition 2.4 What each crawler is for, and judging the door leg by leg (read-only).

1.3 Three gates and three paths

What you'll do in this section: memorise the order of door × identity × shelf and the acceptance ruler for each stage, and know that only three paths can move your ranking. After reading this, you'll be able to tag every action on your plate ①, ② or ③ and delete whatever you cannot tag. You'll also be able to judge for yourself whether you are winning or losing this month.

Three gates: the first two are switches, only the third adds points

no ×0

yes ×1

no ×0

yes ×1

adds points month by month

acceptance

ruler

ruler

ruler

Door: can the crawler read the body text?

everything after is multiplied by 0

Identity: is this business identified correctly?

credit goes to others, or facts are wrong

Shelf · placement: you are on those pages

Shelf · extraction: answer copies your sentence

crawler hit table + view source with JS off

factual errors (count)

on the list and cited X/20

seat count

Figure: The door and identity are ×0/1 switches; only the shelf adds points month by month. Each stage has its own acceptance ruler
  • Why the order cannot be swapped: the typical result of reversing it is spending three months on the shelf (writing pages, sending outreach letters), then finding on day 91 that the door was closed the whole time. Everything produced in the first 90 days goes back to zero and has to be redone.
  • Why the shelf is split into two stages: when you are listed as a source but not written into the answer, adding pages does not help; you need to change the sentences. Without the "extraction" stage, the people doing the work will just keep adding pages.

Example (e-commerce) A category page with numbers but no sentences most often gets stuck at the "extraction" stage: you are on the page, but the sentence the answer copies into its body is not yours.

What counts as a win on each of the three rulers

RulerWhat it measuresWhat counts as a win
Seat countHow many times you are named on the frozen questionsThe increase is larger than your own measured noise band, and the direction holds for two months running
On the list and cited X/20Of the 20 third-party pages AI most often cites for this set of questions, the number that include your name and were actually cited this monthThe baseline is usually 0–2; +1 a month, ≥ +3 at 90 days
Factual errors (count)How many things AI gets wrong about your address, price, business status and credentialsDown to 0, with before-and-after screenshots for each one

Only three paths move your ranking

page count n to n+1

the set grows

precondition, otherwise ② is always 0

confidence it is the same business

multiplied by

multiplied by

① Put your words onto pages already cited

how many of the fetched pages mention you

② Get your own page into the candidate set

Door is open

③ Make the words about you consistent

words about you there: specific and consistent

times named

Figure: Times named is the product of two factors, so only three paths can move it; there is no fourth
PathWhat the action looks likeWhich stage it lands in
①Update letters, correction letters, bylined contributions, listings, platform profiles, official registersShelf · placement. The candidate set stays the same; the shortest causal chain
②Door + high-density page type + checkable numbers + sources + a visible update dateDoor + shelf · placement. The whole page that joins the set is about you
③Unique spelling, /facts, person pages, six-trace word-for-word alignment, correctionsIdentity. Adds no pages; raises the confidence that "these pages are about the same business"
  • The shelf · extraction stage does not need another path. It depends on how the sentence is written: checkable numbers, no adjectives, a source.
  • The feed channel that sends product data to the shopping shelf does not go through the door, so count it as a variant of ②. Flag this difference separately. If you don't, you get one of two mistakes: "the door is broken, so the feed must be useless too", or "the feed is on, so there's no need to fix the door". For how to align product pages with the feed, see → General Edition 5.20 Entity anchor pages: three subtypes, and the product detail page (pt09).

How to tell you are losing

risen

not risen

not risen, for two months running

next month

risen

cleared, same direction for two months

not cleared

three months running

On the list X/20: risen?

On the list and cited: risen too?

Has the seat increase cleared the noise band?

effort goes to containers that are not cited

put effort only into sources cited this month

counts as a win

within normal variation; not-moved triage

nine in ten: door not truly open, or ruler wrong

Figure: Pages on the list can rise while you are still losing. Check cited first, then whether seats have cleared the noise band
  • This is what typical losing looks like: three months have passed, 8 pages have been written, 40 outreach letters have gone out off-site, and pages on the list have risen from 9/20 to 13/20. Every table is filled in, but the seat count wobbles inside the noise band, and nobody can say whether it has moved at all, let alone why it hasn't.
  • "The door is not really open": you think crawlers are let through, but the WAF is actually blocking them, or the price exists only in JS. "The ruler measured wrong": using one sampling gap as the threshold for a different sampling gap, adding passes partway through, or describing an API-leg reading as "what users see".
  • The invisible way to lose is harder to spot. If you watch only "on the list", the whole loop can close out normally while rankings do not move at all, and after three months not one number can answer "did AI actually cite the pages we got our words onto, or not?". For how to triage when nothing has moved, see → General Edition 7.3 Not-moved triage and next month's three points.

Two disciplines

  1. Tagging discipline: every action must be tagged ①, ② or ③. Delete any action you cannot tag, however much it looks like marketing work. Writing proposals, building competitor analysis tables, rewriting the brand positioning statement, posting social media image posts: none of these can be tagged, so none of them belong in this playbook.
  2. Confidence discipline: low-confidence items still go into the playbook, with "not measured" at the end of the line; never leave something out just because there is no evidence. But the mechanism must be spelled out. An action you cannot tag ①, ② or ③ is not "not measured"; it "does not hold". Never mix the two up.

1.4 Three kinds of things not to do

What you'll do in this section: check your current plan against the three figures below. Delete the actions that do not move rankings, replace the ones that backfire, and fix the wrong criteria that would leave you unable to find the cause. The third kind is the most expensive: it has you put another three months into the wrong direction.

Things that do not move rankings

Do not doDo instead
Write llms.txtDo not write it
Write JSON-LD page by page and check it word for wordSet up a template once when building the site; leave it off the per-page checklist
Treat sameAs as the main entity leverUse a visible-HTML dl with one field per line, plus clickable links
Stack FAQPage schema, add self-rated scoresQuestion-form H2s + a short Q&A at the foot of the page
Chase view counts, buy views, do flashy editsProduce reference material that can be cited
Make paid listings, advertorials and directory submissions the main lineearned media + institutions, government, associations
Keep a six-category ledger of citation sourcesAssign tasks by the eight page-type categories
Notarise screenshots and run permutation tests during the baseline periodSave the raw answers to disk + the noise band
Put pure symptom questions into the question poolUse the symptom + option + price three-part form
Build a separate page for every phrasing of the questionFold the variants into H2s on the same page
Figure: Ten things that do not move rankings, each with a replacement to do instead
  • llms.txt: take it out of the prediction model and citation frequency is predicted more accurately, not less; it adds noise, not signal. Google's AI optimisation guidance also says outright that AI Overviews / AI Mode do not need it. It is the biggest fake GEO lever of 2026.
  • Page-by-page JSON-LD: after JSON-LD was added, AI Overviews (AIO) citations changed by −4.6%. That was the only significant result, and it was negative; the other two platforms showed no significant change. For how to seal the template once, see → General Edition 2.7 Gates 3 and 4: indexing paths, and JSON-LD sealed once.
  • sameAs: live fetching does not read JSON-LD (1.1), so however many registration numbers, association pages or Wikidata QIDs you chain together, the benefit on the path "AI reads this page live" is zero. Write the schema anyway, but note it as "indirect, Google leg only".
  • FAQPage and self-rated scores: the former is built for rich results and does not feed into AI citations; the latter, on the strictly regulated side, also runs into the ban on testimonials and ratings (Statute text in the example industry).
  • View counts: views, likes and subscriber counts are almost uncorrelated with citation frequency. For the kind of long video to make instead, see → General Edition 6.7 Long videos narrated by the named expert.
  • Paid listings: they are the shortest bar in the bar chart in 1.1. Put the same spend into earned media and institutions, government and associations instead, and the difference is an order of magnitude.
  • The six-category ledger: the six buckets do not produce a single action. Effort allocation reads the page-type column, and its eight categories are: comparison / listicle articles / category directories / profiles / government registers / review sites / communities / own website. The categories are not a ledger; they are the basis for assigning tasks.
  • Notarised screenshots, permutation tests: notarisation is valuable for comparing "after" with "before", and during the baseline period there is only "before". At n ≤ 10, a permutation test does not change any decision.
  • Pure symptom questions: for pure symptom questions, AI cites public-health websites and encyclopaedias and almost never names a practitioner.
  • Phrasing variants: a new page only dilutes the signal for the same intent cluster; see → General Edition 5.5 The page unit: one intent cluster, one page.
  • The samples and sources behind each point are in → General Edition A.1 Evidence for mechanics, the door and identity and → General Edition A.2 Evidence for picking targets, writing pages, off-site and retests.

Things that make it worse

Do not do (it backfires)Do instead
Append an allow group to the end of robots.txtCopy the wildcard group's Disallow lines, one by one, into every named group
Rewrite an old pageOnly expand it: leave the URL, the H1 and the paragraphs carrying ranking keywords untouched
Edit Wikipedia's body text to add yourselfOnly add checkable sources to an existing entry; do not add your organisation's name
Create a Wikidata entry with no third-party materialCreate one only once you have a news report, a bylined journal article or an association announcement
On a regulated side, actively ask for reviews, buy a list spot or write a "from" priceCheck 5.2–5.3 for your side first
Send out material that names peersKeep it for your own eyes only
Change the question, the engine or the number of passes partway throughLeave what's frozen alone
Figure: Seven things that make it worse, and what to do in their place

Things that hide the cause

Wrong criterionWhat to check instead
curl with a spoofed crawler UA, check what comes backUse a browser UA to check content; use a crawler UA only to check the status code
Treat IndexNow as a switch for the ChatGPT legPaste the URL and have ChatGPT quote the price line back word for word
Check whether GPTBot is blockedCheck OAI-SearchBot
To judge AI Overviews, check Google-ExtendedCheck Googlebot and nosnippet
Add the two legs together or combine them into one scoreKeep the three rows side by side; do not add them, do not use one to confirm another
Describe an API-leg result as what users seeQualify it as "on the model API"
Cite "69% of AI crawlers do not execute JS"Zero execution evidence; the only exceptions are Gemini and Applebot
Let bare category words make up most of the question poolLocation ≥60%, price words ≥30%, bare words ≤10%
Run it once and declare zero visibilityA single run on a single engine, n ≤ 3, is directional evidence only
Skip triage because pages on the list have risenIf seats have not cleared the noise band, always start triage at layer 1
Figure: Ten wrong criteria that leave you unable to find the cause, and what to check instead

Back to contents · GEO Playbook: General Edition

This chapter is published under a CC BY 4.0 licence · © Canlah AI. To republish or adapt it, credit “Canlah AI · GEO Playbook” and link to this page.

A condensed version for AI assistants is on GitHub, and the Markdown version of this chapter can go straight to an AI assistant. The quick guide and full-book downloads are in the downloads section. The measurements behind the numbers in this book are on the dataset page (CC BY 4.0).