1.1 How buyers ask, and what AI reads
What you'll do in this section: understand where the paragraph AI gives the buyer comes from. AI fetches a batch of pages live, reads only the text visible on those pages and does not execute JavaScript. After reading this, you'll know why the later chapters fix the door first, and why the main battlefield is not your website.
Example (dental) Buyers no longer scroll through ten blue links. They open ChatGPT and ask directly: "新加坡哪家做隐形矫正好?" (which clinic in Singapore is best for clear aligner treatment?) AI replies with a paragraph that names a few clinics.
How many names an answer holds is not something someone else's industry report can tell you. Trust only the number you count when you measure your own set of questions. On the same set of questions, some answers name 3 businesses, some name 8, and some name none at all (they only explain the basics and tell you to "consult a professional").
After the buyer has a name
- ①See the namesThe answer names a few businesses; about 1% of visits click the links in the summary (US sample)
- ②Search againThey skip the links and search the name directly
- ③Go to the websiteCheck price, process, credentials
- ④Check profilesReviews, directories, business profiles
- ⑤Get in touchThis is where the sale happens
- Educational questions often produce answers that name no business at all. Even if the answer cites you, the buyer has no name to go back and search for.
- Steps ②–④ affect the conversion rate, not exposure: a wrong address, price or business status stated by AI is what the buyer runs into at these steps. This is one reason identity comes before writing pages; see → General Edition 3.1 Why this comes before writing pages · The two-hour checklist.
- The sample and source for "about 1%" are in → General Edition A.1 Evidence for mechanics, the door and identity.
Where citations come from: the main battlefield is not your website
- Besides covering these cells, your website has one more job: supplying original facts that other people can copy.
- Even when the buyer searches your brand name, owned content accounts for only 23% of AI citations (Omniscient; its denominator is different from the bar chart above, so the two numbers must not be compared side by side).
- So when you record the shelf, what you record is not "am I there" but which containers currently hold the seats for this question. For how to split off-site effort, see → General Edition 6.1 Where the effort goes, and two numbers that must not sit side by side.
- The samples and how the numbers were counted in both studies are in → General Edition A.1 Evidence for mechanics, the door and identity.
What AI reads: visible HTML, no JavaScript
- This exception is useful: if the Google leg has seats and the ChatGPT leg has zero, check CSR first, not your choice of questions. For how to check, see → General Edition 2.9 Door-layer troubleshooting and this chapter's acceptance checks.
- Structured data that is not visible on the page is not read either. When the five main systems (ChatGPT / Claude / Perplexity / Gemini / AI Mode) fetch live, not one of them reads JSON-LD; hidden Microdata and RDFa are ignored too, and every test price was taken from visible HTML. This only shows that "schema does nothing on the live-fetch path". It does not follow that "schema is useless". For where schema belongs and how far to build it out, see → General Edition 2.7 Gates 3 and 4: indexing paths, and JSON-LD sealed once.
- The samples and dates of both tests, and how many prices each system retrieved, are in → General Edition A.1 Evidence for mechanics, the door and identity.
Three things later chapters take as given
- Before answering, AI fetches pages live and reads only visible HTML: it does not read JSON-LD, and mainstream retrieval crawlers do not execute JavaScript (the exceptions are Gemini and Applebot).
- Most citations come from pages other people write about you (earned media, 84%); paid placements and advertorials make up 0.3%. Your website only covers the cells third parties cannot fill.
- The order is door × identity × shelf (1.3), and progress is measured only with the three rulers plus the noise band from same-week retests (1.3, → General Edition 7.1 What a retest produces, and the rulers). The other premises (citation density by page type, the correlation coefficient for video, the share of price questions that trigger AI Overviews, annual opportunity value) are given in the section where each is used.
1.2 Two legs: ChatGPT looks for the source, AI Mode for second-hand summaries
What you'll do in this section: tell apart which pages each leg looks for and which ruler measures each one. After reading this, you'll have one more habit: before you write any sentence, ask "which leg is this sentence for?"
In our measurements, the direction was the same in all five industries. For what each leg cites and in what share, industry by industry, see → General Edition A.3 Page-type measurements (1): two legs, seven rules, the nine tier-A types.
Example (B2B SaaS) Official pricing pages make up about 53% of ChatGPT's citations, plus about 7 citations of help centres, and then ChatGPT writes its own conclusion. For AI Mode, YouTube, Reddit and LinkedIn together make up 40%, and the rest are peers' "vs" pages and lists.
- The last row of the figure is the direct takeaway for writing pages. ChatGPT also uses market sentences when it writes an overview, so write both kinds of sentence on every page, each in its own place; the spec is in → General Edition 5.6 Two legs, two kinds of sentence.
- In regulated industries, market sentences are written differently, and some may not be written at all. First check the quick reference for your side in → General Edition 5.2 Quick reference by side (1): identify the advertiser first; prices, promotions and freebies, result numbers. Do not copy sentence patterns straight from this figure.
- Local questions draw on business-profile sources, and IndexNow is not a switch for the ChatGPT leg either; both points are covered in → General Edition 2.7 Gates 3 and 4: indexing paths, and JSON-LD sealed once.
Three rulers
The figure above answers "which pages to write, which sentences to write"; the one below answers "what to measure with". In the page-type tests, the AI Mode leg is used only to see which kinds of page it cites. It is not used as a monthly ruler.
- In reports, the ruler legs may only be called this: "Model-API visibility baseline · OpenAI leg" and "Gemini model leg", printed in the footer of every report. Never write "what users see in ChatGPT" or "real user visibility". Citation distributions on the API and on the web version show a measured, systematic skew. It is not noise, and running more rounds does not smooth it out. For how large the skew is and how to run the web control leg, see → General Edition 4.4 The web control leg and the frozen baseline (the book's only full spec).
- The naming rule is not about wording: get the name wrong, and a whole quarter's off-site effort gets aimed at the API leg's top 20, while buyers may be seeing a different set of pages.
- Gemini API ≠ AI Overviews ≠ AI Mode. They are three different products, with different fetch paths, different ways of generating answers and different source pools. Passing off Gemini model leg numbers as visibility in AI Overviews means measuring A and using it to sign off B.
- The Search Console ruler has impressions only, with no times named and no position. It is already included in total impressions, so it must not be added to total impressions.
- For why you check
Googlebotrather thanGoogle-Extended, see → General Edition 2.4 What each crawler is for, and judging the door leg by leg (read-only).
1.3 Three gates and three paths
What you'll do in this section: memorise the order of door × identity × shelf and the acceptance ruler for each stage, and know that only three paths can move your ranking. After reading this, you'll be able to tag every action on your plate ①, ② or ③ and delete whatever you cannot tag. You'll also be able to judge for yourself whether you are winning or losing this month.
Three gates: the first two are switches, only the third adds points
- Why the order cannot be swapped: the typical result of reversing it is spending three months on the shelf (writing pages, sending outreach letters), then finding on day 91 that the door was closed the whole time. Everything produced in the first 90 days goes back to zero and has to be redone.
- Why the shelf is split into two stages: when you are listed as a source but not written into the answer, adding pages does not help; you need to change the sentences. Without the "extraction" stage, the people doing the work will just keep adding pages.
Example (e-commerce) A category page with numbers but no sentences most often gets stuck at the "extraction" stage: you are on the page, but the sentence the answer copies into its body is not yours.
What counts as a win on each of the three rulers
| Ruler | What it measures | What counts as a win |
|---|---|---|
| Seat count | How many times you are named on the frozen questions | The increase is larger than your own measured noise band, and the direction holds for two months running |
| On the list and cited X/20 | Of the 20 third-party pages AI most often cites for this set of questions, the number that include your name and were actually cited this month | The baseline is usually 0–2; +1 a month, ≥ +3 at 90 days |
| Factual errors (count) | How many things AI gets wrong about your address, price, business status and credentials | Down to 0, with before-and-after screenshots for each one |
- The three rows sit side by side: never merge them and never add them together. The full definitions and how to run them each month are in → General Edition 7.1 What a retest produces, and the rulers; how to measure the noise band is in → General Edition 7.2 The noise band, the page-level signal and the ten monthly steps.
- If you watch only one number, watch the factual error count. AI copies factual errors from the source pages it retrieves, so when you fix the source, both legs change with it. Of all the rulers it is the least sensitive to the difference between the API leg and the web leg, and the only one in the whole set that is not contaminated by the sampling-leg skew.
Only three paths move your ranking
| Path | What the action looks like | Which stage it lands in |
|---|---|---|
| ① | Update letters, correction letters, bylined contributions, listings, platform profiles, official registers | Shelf · placement. The candidate set stays the same; the shortest causal chain |
| ② | Door + high-density page type + checkable numbers + sources + a visible update date | Door + shelf · placement. The whole page that joins the set is about you |
| ③ | Unique spelling, /facts, person pages, six-trace word-for-word alignment, corrections | Identity. Adds no pages; raises the confidence that "these pages are about the same business" |
- The shelf · extraction stage does not need another path. It depends on how the sentence is written: checkable numbers, no adjectives, a source.
- The feed channel that sends product data to the shopping shelf does not go through the door, so count it as a variant of ②. Flag this difference separately. If you don't, you get one of two mistakes: "the door is broken, so the feed must be useless too", or "the feed is on, so there's no need to fix the door". For how to align product pages with the feed, see → General Edition 5.20 Entity anchor pages: three subtypes, and the product detail page (pt09).
How to tell you are losing
- This is what typical losing looks like: three months have passed, 8 pages have been written, 40 outreach letters have gone out off-site, and pages on the list have risen from 9/20 to 13/20. Every table is filled in, but the seat count wobbles inside the noise band, and nobody can say whether it has moved at all, let alone why it hasn't.
- "The door is not really open": you think crawlers are let through, but the WAF is actually blocking them, or the price exists only in JS. "The ruler measured wrong": using one sampling gap as the threshold for a different sampling gap, adding passes partway through, or describing an API-leg reading as "what users see".
- The invisible way to lose is harder to spot. If you watch only "on the list", the whole loop can close out normally while rankings do not move at all, and after three months not one number can answer "did AI actually cite the pages we got our words onto, or not?". For how to triage when nothing has moved, see → General Edition 7.3 Not-moved triage and next month's three points.
Two disciplines
- Tagging discipline: every action must be tagged ①, ② or ③. Delete any action you cannot tag, however much it looks like marketing work. Writing proposals, building competitor analysis tables, rewriting the brand positioning statement, posting social media image posts: none of these can be tagged, so none of them belong in this playbook.
- Confidence discipline: low-confidence items still go into the playbook, with "not measured" at the end of the line; never leave something out just because there is no evidence. But the mechanism must be spelled out. An action you cannot tag ①, ② or ③ is not "not measured"; it "does not hold". Never mix the two up.
1.4 Three kinds of things not to do
What you'll do in this section: check your current plan against the three figures below. Delete the actions that do not move rankings, replace the ones that backfire, and fix the wrong criteria that would leave you unable to find the cause. The third kind is the most expensive: it has you put another three months into the wrong direction.
Things that do not move rankings
- llms.txt: take it out of the prediction model and citation frequency is predicted more accurately, not less; it adds noise, not signal. Google's AI optimisation guidance also says outright that AI Overviews / AI Mode do not need it. It is the biggest fake GEO lever of 2026.
- Page-by-page JSON-LD: after JSON-LD was added, AI Overviews (AIO) citations changed by −4.6%. That was the only significant result, and it was negative; the other two platforms showed no significant change. For how to seal the template once, see → General Edition 2.7 Gates 3 and 4: indexing paths, and JSON-LD sealed once.
- sameAs: live fetching does not read JSON-LD (1.1), so however many registration numbers, association pages or Wikidata QIDs you chain together, the benefit on the path "AI reads this page live" is zero. Write the schema anyway, but note it as "indirect, Google leg only".
- FAQPage and self-rated scores: the former is built for rich results and does not feed into AI citations; the latter, on the strictly regulated side, also runs into the ban on testimonials and ratings (Statute text in the example industry).
- View counts: views, likes and subscriber counts are almost uncorrelated with citation frequency. For the kind of long video to make instead, see → General Edition 6.7 Long videos narrated by the named expert.
- Paid listings: they are the shortest bar in the bar chart in 1.1. Put the same spend into earned media and institutions, government and associations instead, and the difference is an order of magnitude.
- The six-category ledger: the six buckets do not produce a single action. Effort allocation reads the page-type column, and its eight categories are: comparison / listicle articles / category directories / profiles / government registers / review sites / communities / own website. The categories are not a ledger; they are the basis for assigning tasks.
- Notarised screenshots, permutation tests: notarisation is valuable for comparing "after" with "before", and during the baseline period there is only "before". At n ≤ 10, a permutation test does not change any decision.
- Pure symptom questions: for pure symptom questions, AI cites public-health websites and encyclopaedias and almost never names a practitioner.
- Phrasing variants: a new page only dilutes the signal for the same intent cluster; see → General Edition 5.5 The page unit: one intent cluster, one page.
- The samples and sources behind each point are in → General Edition A.1 Evidence for mechanics, the door and identity and → General Edition A.2 Evidence for picking targets, writing pages, off-site and retests.
Things that make it worse
- robots: the matching rule is "the most specific group wins, and every other group is ignored". Append a group directly and the named crawlers break away from the wildcard group completely: all the protection you had on the admin area, shopping cart and site search stops applying to them. For the merge steps, see → General Edition 2.6 Fixing the door: nosnippet, the four-step robots merge, the WAF allowlist.
- Rewriting an old page: an old page has history, links and existing rankings. The expansion cap is relaxed to double the original length. Before editing, export the page's top 10 queries from Search Console, circle the matching paragraphs one by one, and do not delete them. See → General Edition 5.5 The page unit: one intent cluster, one page.
- Wikipedia: self-referencing edits get reverted and leave a trace in the edit history, so the risk outweighs the benefit. See → General Edition 3.4 Third-party credential tiers, Wikidata and the five business profiles.
- Wikidata: it is not "no threshold". Admission requires that the subject "can be described using serious and publicly available reference material". Create an entry without that and the community will nominate it for deletion. See → General Edition 3.4 Third-party credential tiers, Wikidata and the five business profiles.
- The two regulated-side rows are not an efficiency issue; they are a compliance issue, and "disclosing that it is paid" is not a release from liability. Comparison material that names peers is never sent out. For whether and how each cell may be written, check → General Edition 5.2 Quick reference by side (1): identify the advertiser first; prices, promotions and freebies, result numbers and → General Edition 5.3 Quick reference by side (2): testimonials and reviews, comparisons, lists, titles, outbound links, FAQ and captions for your side. For the stop-and-escalate and sign-off rules when you are unsure, see → General Edition 0.4 The three labels for compliance sentences, and the stop-and-escalate rule: a written sign-off resolves only the three stop-and-escalate situations; it never turns something banned as Statute text into something publishable.
- Changing the ruler: change the question, the engine or the number of passes, and the before and after numbers can no longer be compared; see → General Edition 7.5 Ruler discipline: change the ruler and nothing is comparable.
Things that hide the cause
- Why each of the first four rows is wrong is explained in one place only: the curl row in → General Edition 2.3 Gate 2: four identities, fetched live (read-only), where it is downgraded to a troubleshooting tool, not a criterion; IndexNow in → General Edition 2.7 Gates 3 and 4: indexing paths, and JSON-LD sealed once; GPTBot and Google-Extended in → General Edition 2.4 What each crawler is for, and judging the door leg by leg (read-only).
- Adding them together: the three rows measure three different systems, with different denominators and different failure modes. Add them up and any one of them getting worse can be hidden by another (1.2).
- The API leg: citation distributions on the API and on the web version have a systematic skew, so wherever you only have API-leg evidence, always qualify the statement as "on the model API" (1.2).
- 69%: this figure is second-hand (reprocessed), and the actual finding is stronger; see 1.1.
- Bare category words: the wording of the question pool can produce a fake zero visibility. For how the ratios are fixed, see → General Edition 4.2 The question pool: where the 30 questions come from, how they are balanced, how they are signed; the comparison data behind this is in → General Edition A.2 Evidence for picking targets, writing pages, off-site and retests.
- Pages on the list have risen: pages on the list are the output of your own work, so +1 a month is close to the norm. Whether they have risen has nothing to do with whether you need triage; see → General Edition 7.3 Not-moved triage and next month's three points.