# Chapter 1 · How AI picks its answers: what buyers ask, and which pages ChatGPT and Google's AI cite

> GEO Playbook · General Edition v1.0 · Canlah AI · CC BY 4.0 · Web page: https://canlah.ai/playbook/how-ai-answers/
> Markdown edition for AI assistants, same content as the web page. Figures are code blocks (wireframe / mermaid / bars / steps / split); "→" links open the matching section on the web, and the same URL with .md is its Markdown edition.

## 1.1 How buyers ask, and what AI reads

**What you'll do in this section**: understand where the paragraph AI gives the buyer comes from. AI fetches a batch of pages live, reads only the text visible on those pages and does not execute JavaScript. After reading this, you'll know why the later chapters fix the door first, and why the main battlefield is not your website.

> **Example** (dental) Buyers no longer scroll through ten blue links. They open ChatGPT and ask directly: "新加坡哪家做隐形矫正好？" (which clinic in Singapore is best for clear aligner treatment?) AI replies with a paragraph that names a few clinics.

```mermaid Figure: AI does not answer from memory. It fetches pages live and reads only visible text; the businesses the answer names make up its total names
flowchart LR
  q["Buyer asks a question"] -->|receives the question| f["Fetches a batch of pages live"]
  f -->|reads only| v["Visible text on the page"]:::hl
  v -->|writes from this| a["Replies in one paragraph"]
  a -->|names a few businesses| s["Names per answer: single digits, own tests only"]
```

How many names an answer holds is not something someone else's industry report can tell you. Trust only the number you count when you measure your own set of questions. On the same set of questions, some answers name 3 businesses, some name 8, and some name none at all (they only explain the basics and tell you to "consult a professional").

### After the buyer has a name

```steps Figure: Buyers rarely click the links in the answer. Once they have a name they go back and check it themselves, so pick "who to buy from" questions, not educational questions
① | See the names | The answer names a few businesses; about 1% of visits click the links in the summary (US sample)
② | Search again | They skip the links and search the name directly
③ | Go to the website | Check price, process, credentials
④ | Check profiles | Reviews, directories, business profiles
⑤ | Get in touch | This is where the sale happens
```

- Educational questions often produce answers that name no business at all. Even if the answer cites you, the buyer has no name to go back and search for.
- Steps ②–④ affect the conversion rate, not exposure: a wrong address, price or business status stated by AI is what the buyer runs into at these steps. This is one reason identity comes before writing pages; see （→ 通用版 3.1 为什么排在写页之前 · 两小时清单）.
- The sample and source for "about 1%" are in （→ 通用版 A.1 机制、门与人这一层的证据）.

### Where citations come from: the main battlefield is not your website

```bars id=earned-share Figure: Most AI citations come from pages other people write about you; paid placements and advertorials are close to zero
unit: %
earned media (others writing about you) | 84 | Muck Rack
of which news coverage | 27 | included within the 84
paid placements and advertorials | 0.3 | Muck Rack
```

```mermaid Figure: Your website only covers the cells third parties cannot fill; the main off-site battlefield is media and institutions, not paid listings
flowchart LR
  ai["The pages AI fetches live"] -->|a minority| own["Your website"]
  ai -->|the vast majority| ext["Pages other people write about you"]
  own -->|only covers| cells["Fixed prices, process, credentials, org facts"]:::hl
  ext -->|main battlefield| main["Media, associations, gov registers, vendor certs"]
  ext -->|not| paid["Paid listings, advertorials"]:::warn
```

- Besides covering these cells, your website has one more job: supplying original facts that other people can copy.
- Even when the buyer searches your brand name, owned content accounts for only 23% of AI citations (Omniscient; its denominator is different from the bar chart above, so the two numbers must not be compared side by side).
- So when you record the shelf, what you record is not "am I there" but **which containers currently hold the seats for this question**. For how to split off-site effort, see （→ 通用版 6.1 火力方向与两个不能并排的数）.
- The samples and how the numbers were counted in both studies are in （→ 通用版 A.1 机制、门与人这一层的证据）.

### What AI reads: visible HTML, no JavaScript

```split Figure: Mainstream AI crawlers do not execute JavaScript; only two exceptions render pages
Do not execute JS (zero evidence of execution in a large sample) || Render pages (exceptions)
GPTBot, OAI-SearchBot, ChatGPT-User || Gemini: fully rendered through Google's infrastructure
ClaudeBot, Claude-User || Applebot: a browser-style crawler
PerplexityBot, Bytespider || —
Meta-ExternalAgent || —
Price written only in JS: not a single character comes through || Price written only in JS: it comes through
```

- **This exception is useful: if the Google leg has seats and the ChatGPT leg has zero, check CSR first, not your choice of questions.** For how to check, see （→ 通用版 2.9 门层排查与本章验收）.
- Structured data that is not visible on the page is not read either. When the five main systems (ChatGPT / Claude / Perplexity / Gemini / AI Mode) fetch live, not one of them reads JSON-LD; hidden Microdata and RDFa are ignored too, and every test price was taken from visible HTML. This only shows that "schema does nothing on the live-fetch path". It does not follow that "schema is useless". For where schema belongs and how far to build it out, see （→ 通用版 2.7 闸三闸四：收录通路与 JSON-LD 一次封版）.
- The samples and dates of both tests, and how many prices each system retrieved, are in （→ 通用版 A.1 机制、门与人这一层的证据）.

### Three things later chapters take as given

1. Before answering, AI fetches pages live and reads only visible HTML: it does not read JSON-LD, and mainstream retrieval crawlers do not execute JavaScript (the exceptions are Gemini and Applebot).
2. Most citations come from pages other people write about you (earned media, 84%); paid placements and advertorials make up 0.3%. Your website only covers the cells third parties cannot fill.
3. The order is door × identity × shelf (1.3), and progress is measured only with the three rulers plus the noise band from same-week retests (1.3, （→ 通用版 7.1 复测的产出与量具）). The other premises (citation density by page type, the correlation coefficient for video, the share of price questions that trigger AI Overviews, annual opportunity value) are given in the section where each is used.

## 1.2 Two legs: ChatGPT looks for the source, AI Mode for second-hand summaries

**What you'll do in this section**: tell apart which pages each leg looks for and which ruler measures each one. After reading this, you'll have one more habit: before you write any sentence, ask "which leg is this sentence for?"

```split Figure: The two legs do not look for the same kind of page, and they want different kinds of sentences
ChatGPT leg || AI Mode leg
Looks for the source: government pages, official pricing pages, product pages || Looks for second-hand summaries: peer articles, lists, videos, Google cards
Takes the source facts and writes its own conclusion || Cites pages other people have already summarised
Feeds on own sentences: you as the subject, exact amount and unit || Feeds on market sentences: range + variables + source and date
```

**In our measurements, the direction was the same in all five industries.** For what each leg cites and in what share, industry by industry, see （→ 通用版 A.3 页型实测（一）：两条腿、七条规则、档 A 九型）.

> **Example** (B2B SaaS) Official pricing pages make up about 53% of ChatGPT's citations, plus about 7 citations of help centres, and then ChatGPT writes its own conclusion. For AI Mode, YouTube, Reddit and LinkedIn together make up 40%, and the rest are peers' "vs" pages and lists.

- The last row of the figure is the direct takeaway for writing pages. ChatGPT also uses market sentences when it writes an overview, so write both kinds of sentence on every page, each in its own place; the spec is in （→ 通用版 5.6 两条腿，两种句子）.
- **In regulated industries, market sentences are written differently, and some may not be written at all.** First check the quick reference for your side in （→ 通用版 5.2 三侧速查（一）：先判身份；价格、促销与赠送、结果数字）. Do not copy sentence patterns straight from this figure.
- Local questions draw on business-profile sources, and IndexNow is not a switch for the ChatGPT leg either; both points are covered in （→ 通用版 2.7 闸三闸四：收录通路与 JSON-LD 一次封版）.

### Three rulers

The figure above answers "which pages to write, which sentences to write"; the one below answers "what to measure with". In the page-type tests, the AI Mode leg is used only to see which kinds of page it cites. It is not used as a monthly ruler.

```mermaid id=legs-rulers Figure: Three systems, three rulers: each measures its own thing, and they are never added together
flowchart LR
  p1["OpenAI model API"] -->|measures| r1["Model-API visibility baseline · OpenAI leg"]
  p2["Gemini model API"] -->|measures| r2["Gemini model leg"]
  p3["AI Overviews and AI Mode"] -->|only uses| r3["Search Console generative AI report impressions"]:::hl
  p3 -->|can it fetch you| gb["check Googlebot"]
  web["ChatGPT web version"] -->|calibrates once a month| wl["web control leg, not used for acceptance"]
  wl -->|calibrates| r1
  r1 -->|not added| x["one combined visibility score"]:::warn
  r2 -->|not added| x
  r3 -->|not added| x
```

- **In reports, the ruler legs may only be called this**: "Model-API visibility baseline · OpenAI leg" and "Gemini model leg", printed in the footer of every report. **Never write "what users see in ChatGPT" or "real user visibility".** Citation distributions on the API and on the web version show a measured, systematic skew. It is not noise, and running more rounds does not smooth it out. For how large the skew is and how to run the web control leg, see （→ 通用版 4.4 网页对照腿与冻结基线（全书唯一完整规格））.
- **The naming rule is not about wording**: get the name wrong, and a whole quarter's off-site effort gets aimed at the API leg's top 20, while buyers may be seeing a different set of pages.
- **Gemini API ≠ AI Overviews ≠ AI Mode.** They are three different products, with different fetch paths, different ways of generating answers and different source pools. Passing off Gemini model leg numbers as visibility in AI Overviews means measuring A and using it to sign off B.
- The Search Console ruler has impressions only, with no times named and no position. It is already included in total impressions, so it must not be added to total impressions.
- For why you check `Googlebot` rather than `Google-Extended`, see （→ 通用版 2.4 爬虫用途表与逐腿判门（只读））.

## 1.3 Three gates and three paths

**What you'll do in this section**: memorise the order of door × identity × shelf and the acceptance ruler for each stage, and know that only three paths can move your ranking. After reading this, you'll be able to tag every action on your plate ①, ② or ③ and delete whatever you cannot tag. You'll also be able to judge for yourself whether you are winning or losing this month.

### Three gates: the first two are switches, only the third adds points

```mermaid id=three-gates Figure: The door and identity are ×0/1 switches; only the shelf adds points month by month. Each stage has its own acceptance ruler
flowchart LR
  door{"Door: can the crawler read the body text?"} -->|no ×0| zero["everything after is multiplied by 0"]:::warn
  door -->|yes ×1| who{"Identity: is this business identified correctly?"}
  who -->|no ×0| wrong["credit goes to others, or facts are wrong"]:::warn
  who -->|yes ×1| put["Shelf · placement: you are on those pages"]:::hl
  put -->|adds points month by month| pick["Shelf · extraction: answer copies your sentence"]:::hl
  door -->|acceptance| g1["crawler hit table + view source with JS off"]
  who -->|ruler| g2["factual errors (count)"]
  put -->|ruler| g3["on the list and cited X/20"]
  pick -->|ruler| g4["seat count"]
```

- **Why the order cannot be swapped**: the typical result of reversing it is spending three months on the shelf (writing pages, sending outreach letters), then finding on day 91 that the door was closed the whole time. Everything produced in the first 90 days goes back to zero and has to be redone.
- **Why the shelf is split into two stages**: when you are listed as a source but not written into the answer, adding pages does not help; you need to change the sentences. Without the "extraction" stage, the people doing the work will just keep adding pages.

> **Example** (e-commerce) A category page with numbers but no sentences most often gets stuck at the "extraction" stage: you are on the page, but the sentence the answer copies into its body is not yours.

### What counts as a win on each of the three rulers

| Ruler | What it measures | What counts as a win |
|---|---|---|
| Seat count | How many times you are named on the frozen questions | The increase is larger than your own measured noise band, and the direction holds for two months running |
| On the list and cited X/20 | Of the 20 third-party pages AI most often cites for this set of questions, the number that include your name and were actually cited this month | The baseline is usually 0–2; +1 a month, ≥ +3 at 90 days |
| Factual errors (count) | How many things AI gets wrong about your address, price, business status and credentials | Down to 0, with before-and-after screenshots for each one |

- The three rows sit side by side: never merge them and never add them together. The full definitions and how to run them each month are in （→ 通用版 7.1 复测的产出与量具）; how to measure the noise band is in （→ 通用版 7.2 噪声带、页级信号与每月十步）.
- **If you watch only one number, watch the factual error count.** AI copies factual errors from the source pages it retrieves, so when you fix the source, both legs change with it. Of all the rulers it is the least sensitive to the difference between the API leg and the web leg, and the only one in the whole set that is not contaminated by the sampling-leg skew.

### Only three paths move your ranking

```mermaid Figure: Times named is the product of two factors, so only three paths can move it; there is no fourth
flowchart LR
  r1["① Put your words onto pages already cited"] -->|page count n to n+1| n["how many of the fetched pages mention you"]
  r2["② Get your own page into the candidate set"] -->|the set grows| n
  gate["Door is open"] -->|precondition, otherwise ② is always 0| r2
  r3["③ Make the words about you consistent"] -->|confidence it is the same business| q["words about you there: specific and consistent"]
  n -->|multiplied by| s["times named"]:::hl
  q -->|multiplied by| s
```

| Path | What the action looks like | Which stage it lands in |
|---|---|---|
| ① | Update letters, correction letters, bylined contributions, listings, platform profiles, official registers | Shelf · placement. The candidate set stays the same; **the shortest causal chain** |
| ② | Door + high-density page type + checkable numbers + sources + a visible update date | Door + shelf · placement. The whole page that joins the set is about you |
| ③ | Unique spelling, `/facts`, person pages, six-trace word-for-word alignment, corrections | Identity. Adds no pages; raises the confidence that "these pages are about the same business" |

- The shelf · extraction stage does not need another path. It depends on how the sentence is written: checkable numbers, no adjectives, a source.
- The feed channel that sends product data to the shopping shelf does not go through the door, so count it as a variant of ②. Flag this difference separately. If you don't, you get one of two mistakes: "the door is broken, so the feed must be useless too", or "the feed is on, so there's no need to fix the door". For how to align product pages with the feed, see （→ 通用版 5.20 实体锚点页：三个子型与商品详情页（pt09））.

### How to tell you are losing

```mermaid Figure: Pages on the list can rise while you are still losing. Check cited first, then whether seats have cleared the noise band
flowchart LR
  a{"On the list X/20: risen?"} -->|risen| b{"On the list and cited: risen too?"}
  a -->|not risen| e{"Has the seat increase cleared the noise band?"}
  b -->|not risen, for two months running| c["effort goes to containers that are not cited"]:::warn
  c -->|next month| d["put effort only into sources cited this month"]:::hl
  b -->|risen| e
  e -->|cleared, same direction for two months| win["counts as a win"]
  e -->|not cleared| f["within normal variation; not-moved triage"]
  f -->|three months running| g["nine in ten: door not truly open, or ruler wrong"]:::warn
```

- **This is what typical losing looks like**: three months have passed, 8 pages have been written, 40 outreach letters have gone out off-site, and pages on the list have risen from 9/20 to 13/20. Every table is filled in, but the seat count wobbles inside the noise band, and nobody can say whether it has moved at all, let alone why it hasn't.
- "The door is not really open": you think crawlers are let through, but the WAF is actually blocking them, or the price exists only in JS. "The ruler measured wrong": using one sampling gap as the threshold for a different sampling gap, adding passes partway through, or describing an API-leg reading as "what users see".
- **The invisible way to lose** is harder to spot. If you watch only "on the list", the whole loop can close out normally while rankings do not move at all, and after three months not one number can answer "did AI actually cite the pages we got our words onto, or not?". For how to triage when nothing has moved, see （→ 通用版 7.3 没动分诊与下月三个点）.

### Two disciplines

1. **Tagging discipline**: every action must be tagged ①, ② or ③. **Delete any action you cannot tag**, however much it looks like marketing work. Writing proposals, building competitor analysis tables, rewriting the brand positioning statement, posting social media image posts: none of these can be tagged, so none of them belong in this playbook.
2. **Confidence discipline**: low-confidence items still go into the playbook, with "**not measured**" at the end of the line; never leave something out just because there is no evidence. But the mechanism must be spelled out. An action you cannot tag ①, ② or ③ is not "not measured"; it "**does not hold**". Never mix the two up.

## 1.4 Three kinds of things not to do

**What you'll do in this section**: check your current plan against the three figures below. Delete the actions that do not move rankings, replace the ones that backfire, and fix the wrong criteria that would leave you unable to find the cause. The third kind is the most expensive: it has you put another three months into the wrong direction.

### Things that do not move rankings

```split Figure: Ten things that do not move rankings, each with a replacement to do instead
Do not do || Do instead
Write llms.txt || Do not write it
Write JSON-LD page by page and check it word for word || Set up a template once when building the site; leave it off the per-page checklist
Treat sameAs as the main entity lever || Use a visible-HTML dl with one field per line, plus clickable links
Stack FAQPage schema, add self-rated scores || Question-form H2s + a short Q&A at the foot of the page
Chase view counts, buy views, do flashy edits || Produce reference material that can be cited
Make paid listings, advertorials and directory submissions the main line || earned media + institutions, government, associations
Keep a six-category ledger of citation sources || Assign tasks by the eight page-type categories
Notarise screenshots and run permutation tests during the baseline period || Save the raw answers to disk + the noise band
Put pure symptom questions into the question pool || Use the symptom + option + price three-part form
Build a separate page for every phrasing of the question || Fold the variants into H2s on the same page
```

- **llms.txt**: take it out of the prediction model and citation frequency is predicted more accurately, not less; it adds noise, not signal. Google's AI optimisation guidance also says outright that AI Overviews / AI Mode do not need it. It is the biggest fake GEO lever of 2026.
- **Page-by-page JSON-LD**: after JSON-LD was added, AI Overviews (AIO) citations changed by −4.6%. That was the only significant result, and it was negative; the other two platforms showed no significant change. For how to seal the template once, see （→ 通用版 2.7 闸三闸四：收录通路与 JSON-LD 一次封版）.
- **sameAs**: live fetching does not read JSON-LD (1.1), so however many registration numbers, association pages or Wikidata QIDs you chain together, the benefit on the path "AI reads this page live" is zero. Write the schema anyway, but note it as "indirect, Google leg only".
- **FAQPage and self-rated scores**: the former is built for rich results and does not feed into AI citations; the latter, on the strictly regulated side, also runs into the ban on testimonials and ratings (Statute text in the example industry).
- **View counts**: views, likes and subscriber counts are almost uncorrelated with citation frequency. For the kind of long video to make instead, see （→ 通用版 6.7 本人口述长视频）.
- **Paid listings**: they are the shortest bar in the bar chart in 1.1. Put the same spend into earned media and institutions, government and associations instead, and the difference is an order of magnitude.
- **The six-category ledger**: the six buckets do not produce a single action. Effort allocation reads the page-type column, and its eight categories are: comparison / listicle articles / category directories / profiles / government registers / review sites / communities / own website. **The categories are not a ledger; they are the basis for assigning tasks.**
- **Notarised screenshots, permutation tests**: notarisation is valuable for comparing "after" with "before", and during the baseline period there is only "before". At n ≤ 10, a permutation test does not change any decision.
- **Pure symptom questions**: for pure symptom questions, AI cites public-health websites and encyclopaedias and almost never names a practitioner.
- **Phrasing variants**: a new page only dilutes the signal for the same intent cluster; see （→ 通用版 5.5 页面单位：一个意图簇一页）.
- The samples and sources behind each point are in （→ 通用版 A.1 机制、门与人这一层的证据） and （→ 通用版 A.2 选点、写页、站外与复测的证据）.

### Things that make it worse

```split Figure: Seven things that make it worse, and what to do in their place
Do not do (it backfires) || Do instead
Append an allow group to the end of robots.txt || Copy the wildcard group's Disallow lines, one by one, into every named group
Rewrite an old page || Only expand it: leave the URL, the H1 and the paragraphs carrying ranking keywords untouched
Edit Wikipedia's body text to add yourself || Only add checkable sources to an existing entry; do not add your organisation's name
Create a Wikidata entry with no third-party material || Create one only once you have a news report, a bylined journal article or an association announcement
On a regulated side, actively ask for reviews, buy a list spot or write a "from" price || Check 5.2–5.3 for your side first
Send out material that names peers || Keep it for your own eyes only
Change the question, the engine or the number of passes partway through || Leave what's frozen alone
```

- **robots**: the matching rule is "the most specific group wins, and every other group is ignored". Append a group directly and the named crawlers break away from the wildcard group completely: all the protection you had on the admin area, shopping cart and site search stops applying to them. For the merge steps, see （→ 通用版 2.6 改门：nosnippet、robots 四步合并、WAF 白名单）.
- **Rewriting an old page**: an old page has history, links and existing rankings. The expansion cap is relaxed to double the original length. Before editing, export the page's top 10 queries from Search Console, circle the matching paragraphs one by one, and do not delete them. See （→ 通用版 5.5 页面单位：一个意图簇一页）.
- **Wikipedia**: self-referencing edits get reverted and leave a trace in the edit history, so the risk outweighs the benefit. See （→ 通用版 3.4 第三方资质分级、Wikidata 与五处商家档案）.
- **Wikidata**: it is not "no threshold". Admission requires that the subject "can be described using serious and publicly available reference material". Create an entry without that and the community will nominate it for deletion. See （→ 通用版 3.4 第三方资质分级、Wikidata 与五处商家档案）.
- **The two regulated-side rows** are not an efficiency issue; they are a **compliance issue**, and "disclosing that it is paid" is not a release from liability. **Comparison material that names peers is never sent out.** For whether and how each cell may be written, check （→ 通用版 5.2 三侧速查（一）：先判身份；价格、促销与赠送、结果数字） and （→ 通用版 5.3 三侧速查（二）：证言与评价、比较、榜单、头衔、外链、FAQ 与图注） for your side. For the stop-and-escalate and sign-off rules when you are unsure, see （→ 通用版 0.4 合规句的三档标签与停笔规则）: **a written sign-off resolves only the three stop-and-escalate situations; it never turns something banned as Statute text into something publishable.**
- **Changing the ruler**: change the question, the engine or the number of passes, and the before and after numbers can no longer be compared; see （→ 通用版 7.5 量具纪律：换尺子就不可比）.

### Things that hide the cause

```split Figure: Ten wrong criteria that leave you unable to find the cause, and what to check instead
Wrong criterion || What to check instead
curl with a spoofed crawler UA, check what comes back || Use a browser UA to check content; use a crawler UA only to check the status code
Treat IndexNow as a switch for the ChatGPT leg || Paste the URL and have ChatGPT quote the price line back word for word
Check whether GPTBot is blocked || Check OAI-SearchBot
To judge AI Overviews, check Google-Extended || Check Googlebot and nosnippet
Add the two legs together or combine them into one score || Keep the three rows side by side; do not add them, do not use one to confirm another
Describe an API-leg result as what users see || Qualify it as "on the model API"
Cite "69% of AI crawlers do not execute JS" || Zero execution evidence; the only exceptions are Gemini and Applebot
Let bare category words make up most of the question pool || Location ≥60%, price words ≥30%, bare words ≤10%
Run it once and declare zero visibility || A single run on a single engine, n ≤ 3, is directional evidence only
Skip triage because pages on the list have risen || If seats have not cleared the noise band, always start triage at layer 1
```

- **Why each of the first four rows is wrong** is explained in one place only: the curl row in （→ 通用版 2.3 闸二：实抓的四身份（只读））, where it is downgraded to a troubleshooting tool, not a criterion; IndexNow in （→ 通用版 2.7 闸三闸四：收录通路与 JSON-LD 一次封版）; GPTBot and Google-Extended in （→ 通用版 2.4 爬虫用途表与逐腿判门（只读））.
- **Adding them together**: the three rows measure three different systems, with different denominators and different failure modes. Add them up and any one of them getting worse can be hidden by another (1.2).
- **The API leg**: citation distributions on the API and on the web version have a systematic skew, so wherever you only have API-leg evidence, always qualify the statement as "on the model API" (1.2).
- **69%**: this figure is second-hand (reprocessed), and the actual finding is stronger; see 1.1.
- **Bare category words**: the wording of the question pool can produce a fake zero visibility. For how the ratios are fixed, see （→ 通用版 4.2 题池：三十句从哪来、怎么配、怎么签）; the comparison data behind this is in （→ 通用版 A.2 选点、写页、站外与复测的证据）.
- **Pages on the list have risen**: pages on the list are the output of your own work, so +1 a month is close to the norm. Whether they have risen has nothing to do with whether you need triage; see （→ 通用版 7.3 没动分诊与下月三个点）.