# Chapter 2 · Let AI read your website: check first, then fix the settings

> GEO Playbook · General Edition v1.0 · Canlah AI · CC BY 4.0 · Web page: https://canlah.ai/playbook/open-the-door/
> Markdown edition for AI assistants, same content as the web page. Figures are code blocks (wireframe / mermaid / bars / steps / split); "→" links open the matching section on the web, and the same URL with .md is its Markdown edition.

The door is a multiplier: if crawlers cannot get in, however many pages you write afterwards get multiplied by zero. With `nosnippet` switched on, the Google leg is counted as zero for the whole quarter; put the price in JS, and mainstream AI crawlers get not a single word of it. This whole chapter is free and takes half a day to a day, but you do it in two passes. 2.1–2.5 is the D1 read-only audit: get access, record the current state, and change nothing. Only after the baseline has been saved, following （→ 通用版 4.4 网页对照腿与冻结基线（全书唯一完整规格））, do you do 2.6–2.9 and fix the door. For how the three gates — door, identity, shelf — relate to each other, see （→ 通用版 1.3 三个关口与三条路）.

**Hard gate: until the baseline is saved, do not change the door, any profile or the facts page.**

## 2.1 The five gates at a glance, and gate 0: access logs (read-only)

**What you'll do in this section**: first see clearly why the door is done in two passes; on D1, the first move is to send the access request, pull the *Crawler hit table*, and use the four ways of reading it to log problems as to-fix items. When you're done, you'll have a hit table and a to-fix list, and you won't have touched the door.

```mermaid id=read-then-fix Figure: audit read-only first, freeze the baseline, then fix the door — only then are before and after comparable; reverse the order and what you freeze is a shelf you've already changed yourself
flowchart LR
  d1["D1 read-only audit"] -->|change nothing| fz["Freeze baseline 4.1–4.4"]
  fz -->|baseline Top 20, noise band| ld["Baseline saved"]
  ld --> fix["Fix the door, log split day"]
  fix -->|counted against the same denominator| ok["Before and after comparable"]:::hl
  bad["Fix the door before freezing"] --> dirty["Freezes a shelf you've already changed"]
  dirty --> no["Before and after not comparable"]:::warn
```

Why it has to be this way: what the baseline records — page-type mix, the four states per URL, question shape — is all measured on the shelf "while the door is still as it was". Fix the door first and freeze second, and your own changes are already mixed into the before sample; after that, whenever something "went up", you can't tell whether the door did it or it would have risen anyway. So the build order (door first) and the measurement order (freeze first) are two different things — never mix them. On D1 there are only three kinds of thing you can do that do not count as a change: request access, turn on logging, record the current state. The book's week numbers all count from the day you fix the door; see （→ 通用版 0.1 全书一句话与 90 天翻书顺序）.

```steps Figure: go through the five gates in the same order twice — D1 only looks, changes nothing; once the baseline is saved, fix them in this same order
Gate 0 | Access logs | D1 pulls the table; if you can't get it, open a blocker
Gate 1a | Snippet switches | The three nosnippet controls, D1 greps, 2.6 fixes
Gate 1b | robots | D1 records the six lines, 2.6 does the four-step merge
Gate 1c | WAF | D1 checks 403/429, 2.6 adds the allowlist
Gate 2 | Four identities, fetched live | D1 keeps one receipt for each
Gate 3 | Indexing and data | D1 connects the dashboards, 2.7 submits
Gate 4 | JSON-LD | 2.7 sealed in half a day
```

Two ordering rules:

1. **Gate 0 comes before every other gate.** It is the only direct evidence, and the only piece of work where "the doing takes half an hour, but getting access can take days". In your first email, ask for all three kinds of access at once: **domain verification, permission to edit robots, and permission to export access logs.**
2. **Gate 3 connects the free first-party data before you spend money on probes** (see 2.5). Do it the other way round and you'll pay to "discover" something that was already free.

### How to pull gate 0

Faking a UA, a `site:` query, eyeballing it in a browser — all of these are indirect evidence, and each can be wrong in either direction. Without access logs, the whole technical door layer rests on inference.

```steps Figure: four steps to pull the logs; identify a crawler by its IP range only, because anyone can fake a UA string
Pull 30 days | Raw logs | Cloudflare Logpush or the host's access.log
Get IP ranges | Officially published ranges | OpenAI, Anthropic, Perplexity
CIDR filter | Identify by IP | Not by UA string
Produce the hit table | Four columns | Save as crawler_hits_30d.csv
```

OpenAI's three IP ranges are `openai.com/searchbot.json`, `openai.com/gptbot.json` and `openai.com/chatgpt-user.json`; Anthropic and Perplexity each publish their own official ranges. A hit count based on UA strings has scanners and spoofed traffic mixed into it, which invalidates the whole thing as evidence. For the template of the four-column hit table (crawler name, hits in 30 days, top 20 hit URLs, status-code distribution), see （→ 通用版 B.1 开门模板）.

**Two fallbacks when you can't get it**: no export access → open a task the same day and list a **red blocker**, and let gate 2 stand in with criteria b–d for now (see 2.3); the site never kept logs at all → turn on Logpush / access logs the same day, and **the first table only exists after 30 days**.

### How to read the hit table

```mermaid Figure: each reading on the hit table maps to exactly one to-fix item; before the baseline is saved, record only, change nothing
flowchart LR
  t["Crawler hit table"] --> c1["OAI-SearchBot hits = 0"]
  c1 -->|verdict| f1["Door not open, log to-fix 2.6"]:::warn
  t --> c2["403 + 429 over 5% combined"]
  c2 -->|verdict| f2["WAF is blocking, log to-fix 2.6"]
  t --> c3["All 200 but only the homepage hit"]
  c3 -->|verdict| f3["Sitemap and internal links, log to-fix 2.7"]
  t --> c4["ChatGPT-User has hits"]
  c4 -->|verdict| f4["Live-fetch leg is open, log monthly"]:::hl
  t --> c5["Can't get the logs"]
  c5 -->|write only| f5["No direct evidence, fill in after 30 days"]:::warn
```

Three rules outside the figure:

- **At the layer where `OAI-SearchBot` hits are 0, do not write "wrong topic", and do not skip ahead to the "change topic" layer.** If the door isn't open, or the site simply hasn't been discovered, that's a door problem.
- **403 / 429 mean the WAF or a rate rule is blocking, not robots.** Editing robots will not fix it.
- **This table is three things at once**: gate 2's criterion a, the first check in the monthly troubleshooting pass, and the only direct evidence of whether a given leg is actually open. When you can't get the logs, don't say "the door is open"; note the catch-up date on the first line. Saying you have no evidence when you have none is a book-wide rule of how you write things up (see （→ 通用版 D.1 数字纪律：数字怎么标、什么只能自己看）).

For how to fix each reading, see the two diagnostic figures in 2.9.

## 2.2 Gate 1 read-only audit: what to check in nosnippet, robots and the WAF

**What you'll do in this section**: spend 10 minutes grepping the three nosnippet controls, then spend half an hour recording the current state of robots and the WAF in six lines. When you're done, you'll have a *Six-line door read* and 8 screenshots — record only, change nothing; every item you record gets fixed in 2.6.

### Check nosnippet first, then robots

Google's own documentation states it plainly: to restrict how a page's content shows up in Search (**including AI Overviews and AI Mode**), you use `nosnippet`, `data-nosnippet`, `max-snippet` and `noindex` (developers.google.com/search/docs/crawling-indexing/robots-meta-tag). This is **the one switch on the Google leg that can zero you out with a single click**, and CMS templates and SEO plugins often turn it on by default. It takes 10 minutes to check — check it first.

Do this on 8 pages: the homepage, the price page, the `/facts` page, plus 5 template pages (a service page, a practitioner's or author's person page, an old blog post, a list page, a landing page). `curl` each page for **the HTML plus the response headers**.

```split Figure: the three nosnippet controls hide in three places; looking only at the HTML misses the response-header one
Where to check || What a hit looks like
meta robots / googlebot in the HTML || nosnippet, or max-snippet:0 / :20
The HTTP response header X-Robots-Tag || Same as above; you need curl -I to see it
Tag attributes on key parts of the body || data-nosnippet; most often found on price tables, FAQ answers and fact blocks
```

D1's receipt: the results of the three greps (all empty, or which one hit and on which page), one screenshot per page for all 8. **This switch gets turned back on after a redesign, a template change or a plugin update**, so it is also a standing item in the monthly door-layer check (see 2.9).

> ⚠️ Many people treat `Google-Extended` as the AI Overviews switch and spend effort arguing over whether to allow it, while missing this real switch. For why it isn't `Google-Extended`, see the four misconceptions in 2.4.

### robots and WAF: record six lines

```steps Figure: six lines for the read-only audit — miss even one and it doesn't count as checked; lines 2, 4 and 6 each sit on a failure path that can waste the whole quarter
1 | Retrieval UAs | Record each of the 12, allowed or blocked
2 | Wildcard group N | N Disallow lines, copy them out line by line, word for word
3 | Training UAs | All 6, mark every one "unrelated to whether you get cited"
4 | Snippet switches | curl -I the homepage and the price page, check all three places
5 | CDN/WAF | Block AI bots, 429s, allowlist
6 | Access logs | Whether you have them, the date you requested export access
```

The three failure paths are: skip line 2, and merging robots later strips away the protection that was already there (2.6); miss line 4, and `nosnippet` stays on with nobody noticing, so the Google leg counts as zero for the quarter; skip asking about line 6, and gate 0 can't start work the same day, leaving the whole door layer as nothing but inference — which is why this line should be asked as early as possible.

Line 5 needs three things nailed down: ① whether the "Block AI bots" / "Bot Fight Mode" switch is on, off, or doesn't exist; ② whether rate rules hit crawlers with 429s (check against the 429 share in the gate-0 table); ③ whether the IP range for `openai.com/chatgpt-user.json` is on the allowlist.

Where to check each line: for lines 1 and 3, open `https://<domain>/robots.txt` and read it group by group; for line 5, check the CDN / WAF dashboard; for line 6, ask ops and record "has Logpush" / "has `access.log`" / "kept no logs". The six lines take roughly 10, 5, 5, 10, 10 and 5 minutes in turn. For the list and purpose of the 12 retrieval UAs and 6 training UAs, see 2.4; for a blank record sheet, see （→ 通用版 B.1 开门模板）.

> ⚠️ **Allowed in robots does not mean actually let through** — the CDN / WAF layer can still return a straight 403 regardless of robots (**inferred, 85% confidence**). "Allowed" from the audit is only the paper-level first layer; the only evidence of actually being let through is gate 0's *Crawler hit table*.

## 2.3 Gate 2: four identities, fetched live (read-only)

**What you'll do in this section**: stop judging the door by looking at content through `curl -A`, and instead look at the same URL through four paths, then read the differences between the four to tell CSR, WAF and a closed door apart. When you're done, you'll have four separate receipts — never merge them into one.

```split Figure: use curl -A to impersonate a crawler and read the content back, and you can get the verdict wrong in either direction
False negative: you conclude "blocked" || False positive: you conclude "open"
The WAF allows verified crawlers through by IP + rDNS || The site serves a degraded SSR version to non-browser UAs
You fake the UA from your office and get blocked as a spoofed crawler || The real retrieval path actually gets the CSR version
So you go fix something that wasn't broken || So you spend the whole quarter writing pages AI can't read
```

`curl -A "OAI-SearchBot"` only proves how your infrastructure reacts to that string. The kind of pass-through on the left isn't hypothetical: Cloudflare Verified Bots really does allow by IP + rDNS.

### Four identities: one URL as four paths see it

Ranked by how trustworthy each one is; all of them are free.

| Identity | Where you get it | Pass criterion | What it can prove |
|---|---|---|---|
| **a · What the real crawler sees** | Gate 0's access logs, CIDR-filtered against the official IP ranges such as `openai.com/searchbot.json` | This URL was hit by `OAI-SearchBot` within the last 30 days **and returned 200** | **The only direct evidence**, applies across every leg |
| **b · The HTML Bing fetched** | Bing Webmaster Tools → URL Inspection → the source returned by **Live Test** | **You can string-search the source and find the price figure and the body text** | The Bing / Copilot leg; also the free sample closest to "what a non-rendering crawler sees" |
| **c · The HTML Google fetched** | Google Search Console (GSC) → URL Inspection → the HTML under "**Crawled page**" | Same as above, string-searchable | **Only proves Google's side** (Google renders), never use it alone as your conclusion |
| **d · What ChatGPT's live-fetch path sees** | Paste this URL into ChatGPT, verbatim: "Open this link and copy out, word for word, the line on the page that gives the price." | It can **quote it back word for word** (not paraphrase it, not estimate it) | The `ChatGPT-User` live-fetch leg is open |

`curl` is not a fifth identity — it's a troubleshooting tool, and it has to be split into two separate reads: ① fetch the HTML with an **ordinary browser UA** and check whether the body text and price are there — this is the right way to judge CSR, and **CSR has nothing to do with which UA you use**; ② run it again with the `OAI-SearchBot` UA, **look only at the status code, not the content**, to judge the WAF. Screenshot each of the two separately, and write on the receipt "**for troubleshooting only, not a pass criterion**". The half-day quick-start version (（→ 通用版 4.8 两条捷径与本章 checklist）) also judges JS following step ①.

### How to read the differences between the four

```mermaid Figure: the four-identity decision tree — the conclusion comes from the differences between the four, not from ticking each one off
flowchart LR
  a{"How does a look in the logs"} -->|hit and 200| pass["Pass"]:::hl
  a -->|all 0| shut["Door not open, or not discovered"]:::warn
  a -->|403/429 over 5%| waf["WAF or a rate rule is blocking"]
  a -->|no logs| cb{"Is c open and b not?"}
  cb -->|yes| csr["CSR dependency"]:::warn
  cb -->|no| bd{"Is b open? Is d open?"}
  bd -->|both open| low["Pass, degraded"]
  bd -->|b open, d not| waf
  bd -->|any other combination| wait["No verdict yet, fill in a"]
```

Outside the figure, also note:

- **"Pass, degraded" must have "no direct evidence, fill in a after 30 days" noted on the first line.**
- **CSR dependency** = Google's rendering can get it, but a non-rendering crawler cannot. Log it as a to-fix item and open an SSR / prerendering task the same day; push all page work back as a whole, and add no new content pages this month.
- **Never edit content for the WAF cell** — that belongs to the WAF allowlist in 2.6.
- **For the door-not-open cell, never write "demand doesn't exist"**, and never skip ahead to "change topic".

D1 is the audit. On day 7 after fixing the door (2.6), pull all four receipts again and lay them side by side against the D1 set.

## 2.4 What each crawler is for, and judging the door leg by leg (read-only)

**What you'll do in this section**: tell the 12 retrieval crawlers and the 6 training crawlers apart, and know which leg breaks when you block which one; when you hit a platform, multiple subdomains, a login wall or a multilingual site, judge the door leg by leg. When you're done, you'll have a *Per-leg door table*: one row per host, one column per leg — never merge into one cell.

This is the single easiest thing in the whole chapter to get wrong: block the wrong one and nothing happens; allow the wrong one and the whole quarter is wasted.

```split Figure: block a retrieval leg and that leg breaks; Disallow every training leg and not one retrieval leg is affected
Retrieval legs (12) || Training legs (6)
Decide whether you can appear in AI answers at all || Only about content licensing, unrelated to whether you get cited
Block OAI-SearchBot and the ChatGPT leg is zero for the quarter || Block GPTBot and you can still be cited
Block Googlebot and you lose AI Overviews and organic rankings together || Google-Extended has nothing to do with AI Overviews
Allow or not: allow || Allow or not: your own call
```

### 12 retrieval legs: what blocking each one costs

| Crawler | Belongs to | What blocking it costs |
|---|---|---|
| `OAI-SearchBot` | OpenAI | **The ChatGPT leg is zero for the quarter** |
| `ChatGPT-User` | OpenAI | Users who open your page from inside ChatGPT can't read it. ⚠️ **The real switch is in the WAF, not robots** (2.6) |
| `OAI-AdsBot` | OpenAI | Only affects ad scenarios; if you don't run ads, you don't have to allow it |
| `Bingbot` | Microsoft | Copilot breaks outright; one of ChatGPT's eight retrieval sources goes missing |
| `Googlebot` | Google | **You lose AI Overviews / AI Mode and organic rankings together**: all three share the same crawler |
| `Applebot` | Apple | **The whole Siri / Spotlight / Apple Maps / Apple Intelligence leg breaks**. ⚠️ Until this group is allowed, don't claim your Apple Business Connect profile |
| `Amazonbot` | Amazon | Alexa / Amazon-side retrieval breaks |
| `PerplexityBot` | Perplexity | The Perplexity leg breaks |
| `ClaudeBot` | Anthropic | Claude-side access breaks |
| `Claude-SearchBot` | Anthropic | Claude retrieval breaks |
| `Claude-User` | Anthropic | Users who open your page from inside Claude can't read it |
| `DuckAssistBot` | DuckDuckGo | DuckAssist breaks |

The 6 training legs: `GPTBot` · `Google-Extended` · `Applebot-Extended` · `Meta-ExternalAgent` · `CCBot` · `Bytespider`. Whether you allow them is a content-licensing question, not a visibility question.

```split Figure: the four most common misconceptions — all of them mix up training legs with retrieval legs, or mix one company's rules with another's
Misconception || Fact
Blocking GPTBot protects your content, or gets you shut out of AI || GPTBot is a training leg; 88.2% of sites that block it still get cited
Allowing Google-Extended turns on AI Overviews || AI Overviews and AI Mode go by Googlebot + nosnippet
Block Applebot-Extended and the Apple leg breaks || It's only a training opt-out switch; the Apple leg goes by Applebot
No AI crawler respects robots, so editing it is pointless || All three Anthropic bots respect it; it's ChatGPT-User that doesn't
```

The real danger is the flip side of the first one: blocking `OAI-SearchBot` and `Bingbot` themselves. The second is the most expensive to get wrong, and it rests on two of Google's own sentences. The crawler documentation: "**Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search.**" *AI features and your website* nails down the entry requirement: "**a page must be indexed and eligible to be shown in Google Search with a snippet**", and "**robots.txt directives for Googlebot is the control**". So the gate for AI Overviews is `Googlebot` plus the three nosnippet controls. The point of the fourth is that it saves you work: for the three Anthropic bots (`ClaudeBot` / `Claude-SearchBot` / `Claude-User`), editing robots genuinely takes effect — you don't need to touch the WAF. 88.2% comes from BuzzStream; for the rest of the sources, see （→ 通用版 A.5 出处清单）.

```mermaid Figure: judging a leg means checking the right UA; check the wrong one and you get one of two opposite wrong verdicts — "allowed" or "can't get in"
flowchart LR
  q{"Which leg are you judging"} -->|AI Overviews and AI Mode| g["Googlebot + nosnippet"]:::hl
  q -->|citations inside the Gemini app| ge["Google-Extended"]
  q -->|the Apple leg| ap["Applebot"]
  q -->|ChatGPT retrieval| o1["OAI-SearchBot"]
  q -->|ChatGPT live fetch| o2["ChatGPT-User"]
  q -->|Claude retrieval| c1["Claude-SearchBot"]
  q -->|Claude live fetch| c2["Claude-User"]
  q -->|Perplexity retrieval| p1["PerplexityBot"]
  q -->|Perplexity live fetch| p2["Perplexity-User"]
  q -->|Copilot| bb["Bingbot"]
```

**The zero-JS-execution corollary**: mainstream AI crawlers do not execute JS; the only exceptions are Gemini (which runs on Google's infrastructure and renders in full) and Applebot (evidence in （→ 通用版 1.1 买家怎么问，AI 读什么）). So the price, the core facts and the body copy must be in HTML that's already been rendered server-side — fetch the source with a browser UA and you should be able to string-search it; write them in JS instead and the next 90 days are wasted. The flip side of the same corollary: live fetching doesn't read JSON-LD, so schema is a baseline, not a lever (2.7). This exception also gives the whole book its single most useful troubleshooting criterion: **the Google leg has seats while the ChatGPT leg is zero — that difference by itself is the fingerprint of CSR.**

### Judging the door leg by leg

```mermaid Figure: judging one leg's door — not blocked in robots does not mean it can actually fetch; only a raw fetch returning 200 with the body text found counts as "reachable"
flowchart LR
  f{"Could you fetch robots.txt"} -->|can't fetch| nf["Log not retrieved, retry from another network"]:::warn
  f -->|fetched| s{"Does this UA have a named group"}
  s -->|yes| seg["Judge by the named group"]
  s -->|no| wild["Falls to wildcard group; not a block in itself"]
  seg --> r{"Blocked?"}
  wild --> r
  r -->|blocked| shut["Log this leg as blocked"]
  r -->|not blocked| raw{"Raw-fetch a page with this UA"}
  raw -->|200 and body text found| ok["Log reachable"]:::hl
  raw -->|blocked| blk["Log fetch blocked, with status code"]
  raw -->|not done| unk["Write only: robots unblocked, fetch unverified"]
```

Recording rules outside the figure — skip none of them:

1. **`curl` the whole of robots.txt yourself, download it, parse it group by group, and log it by date.** Grepping for one keyword alone will miss the single most important thing: that this UA was never given its own named group at all.
2. **Check robots on the host that the page actually lives on, not the main domain.** Detail pages, doc sites and help centres are often on a different host or subdomain, and robots is **not shared** with the main domain.
3. **One row per host, one column per leg — never merge into one cell.** The very same platform can easily have some legs blocked and others open; merge it into one cell and the whole chain gets thrown away.
4. **`Google-Extended`, `Applebot-Extended` and `Meta-ExternalAgent` are training / licensing legs — they don't go on this table.**
5. **The table must carry a "verification level" column**: `robots-verified` / `fetch-verified`. A raw-fetch receipt may only say "fetched 200 + body text found" or "fetch blocked (`<status>`)", never "allowed". robots only lifts a protocol-level ban; CSR rendering and WAF / geo-blocking (403s, captchas, redirects) are two other, separate gates, and they are not written in robots.
6. **A platform's robots can change**, so log it in the *Off-site asset decay table* and **rerun it once a quarter**. A refused connection, a timeout or a failed TLS handshake all get logged as "not retrieved"; if a retry still fails, leave it blank — **never guess and fill it in.**
7. **A platform closing its door ≠ that platform has no sales value.** All that's ruled dead is "citations through that one leg", not the business; other legs, and the shopping shelf that runs through data feeds, may still show up as normal. Never mix these two things up in anything you tell clients.

On a host where a given leg is blocked, schedule no content work on that leg, and don't let acceptance count it towards seats. What's being judged here is only whether the door is open, which is a different thing from （→ 通用版 4.6 四态、两层分母与弃权线） judging whether "a given URL can be reached at all". For reproducible, per-platform robots commands, see （→ 通用版 B.1 开门模板）.

> **Example** (training) Measured on 2026-09-21, the robots.txt of TPGateway, Singapore's official training portal: `User-agent: *` → `Disallow: /`; the `Allow: /` groups that follow list only `Googlebot` (including Image / Video / mobile), `Slurp`, `bingbot`, `Applebot`, `Twitterbot` and `Pingdom`. Not one of `GPTBot` / `OAI-SearchBot` / `ChatGPT-User` / `ClaudeBot` / `PerplexityBot` appears — all of them fall under `Disallow: /`. → The Google, Apple and Bing legs are open; the OpenAI, Anthropic and Perplexity legs are blocked: build for the first three legs only, accept work against the first three legs only, and don't count ChatGPT seats.

### Five kinds of site, one extra check each

| Kind of site | Extra check | Criterion and how to handle it |
|---|---|---|
| **Single-page app / client-side rendering** | Fetch the HTML with an ordinary browser UA, grep for one actual sentence of body text | Not found = the body text isn't in the HTML. Move key body text to server-side rendering or prerendering — no "wait and bundle it with the next SEO redesign" |
| **Multiple subdomains** | `curl` a robots.txt for each subdomain separately, one cell per subdomain × retrieval leg | For any subdomain with no receipt, its pages don't count as "door open" this quarter and can't be counted towards deliverables. Doc sites, help centres and trust pages need their door opened and to be server-side rendered too |
| **Login wall / quote wall** | Whether the key figures sit outside the wall | **Anything behind the wall is invisible to AI**: build a public, checkable summary page or a mirror on the main domain |
| **Multilingual sites** | Run the four-identity check once for each language | Never check only one language |
| **Price or stock that changes very fast** | Whether these fields are injected by a client-side script | Yes = as good as not there |

**Judge the two causes separately, and don't reverse the order**: check the robots receipts first (one `curl` per subdomain × leg — the cheapest check). When one subdomain was missed, the blocked leg reads zero and the unblocked legs read normal — a one-line fix; only check CSR once robots is all clear. Reverse the order and you'll schedule a whole quarter of engineering rework for a problem that a single line in robots would have fixed. Only once both are fixed does that leg go from 0 to 1.

> **Example** (B2B software) A site naturally has multiple subdomains: `www` / `docs` / `help` / `blog` / `status` / `app` / `trust`. Check robots for each subdomain separately — a clean main domain doesn't mean `docs` is clean too.

## 2.5 Gate 3, read-only: connect the two free AI reports

**What you'll do in this section**: on D1, connect GSC's generative AI report and Bing Webmaster Tools' AI Performance, actually read each one once, and check that the site field they return really is your own site. When you're done, you'll have two sets of data that are the only free, first-party, unfakeable data in the whole engagement; only open the probe budget after you've read them.

```steps Figure: five steps to connect the two dashboards; reading the wrong site is worse than not connecting at all, because the tool will silently read another site and still return ok
1 | Request access | Domain verification, in the same email as the log access request
2 | Open the GSC report | Generative AI performance, export the CSV and screenshot it
3 | Open the Bing report | Webmaster Tools AI Performance, same as above
4 | Actually read it once | Check the site field — get it wrong and you'll silently read another site
5 | Save it | Two connection screenshots; how to read them is in 4.2
```

Three rules:

- **Get the site field wrong = this box isn't done — don't mark it complete.** When the site parameter isn't configured correctly, the tool will read a different site and still return ok, which is worse than not connecting at all: you'll be treating someone else's numbers as your own baseline. The same goes for other read-only dashboards such as GA4.
- **Can't get domain-verification access → list a red blocker.** It blocks more than just this box: Bing's indexing submission (2.7), and Bing Live Test and GSC's crawled-page view (identities b and c in 2.3) all need domain verification first.
- **This section is connect-and-read only.** How to read each of the two reports, what each one gives you, and what to watch out for are written up once, in （→ 通用版 4.2 题池：三十句从哪来、怎么配、怎么签）; checking Bing for indexing gaps and submitting via IndexNow are changes, and belong in 2.7.

Another free thing to start the same day, in passing, is copying out the actual words buyers used, from your front-desk records or CRM, for the last 15 deals before they closed — it's the question pool's only class A source, and how to do it is also in 4.2.

## 2.6 Fixing the door: nosnippet, the four-step robots merge, the WAF allowlist

**What you'll do in this section**: first confirm the baseline is saved, then on the same day make all the fixes in the order `nosnippet` → robots → WAF, and log the day you fix the door as the split day. The hard gate still applies: **until the baseline is saved, do not change the door, any profile or the facts page.** When you're done, you'll have a merged robots.txt, a receipt with the N × 12 line written on it, and one WAF allowlist rule.

```steps Figure: on the day you fix the door, go in this order — the split day counts from this day; if the precondition isn't met, take none of the steps after it
Precondition | Baseline saved | Baseline Top 20, noise band, the 36 brand-six runs
Gate 1a | nosnippet | Revert all three places, grep comes up empty
Gate 1b | robots | Four-step merge, receipt has the N × 12 sentence, word for word
Gate 1c | WAF | Add the chatgpt-user.json IP range
Log split day | The day you fix the door | Written into the page register
D+7 | Re-pull and recheck | Check the logs for ChatGPT-User, re-pull the four identities
```

The day you fix the door is the split day for the whole engagement; every "before vs after fixing the door" comparison is drawn against it. Each page also has its own split day: the day it goes live (2.8). The book's W1 also counts from this day.

### Gate 1a · nosnippet

```split Figure: each of the three hit locations has its own fix; only when a rerun of the grep comes up empty on all three does it count as passed
Where it hits || How to fix it
nosnippet or a low number in meta robots / googlebot || Change it to max-snippet:-1
The response header X-Robots-Tag || Remove it at the CDN / server layer, or change it to max-snippet:-1
data-nosnippet on a body tag || Delete this attribute
```

**Any hit left unfixed for four weeks → the Google leg counts as zero for the quarter** — write it on the first line, and mark the Google cell "switch not lifted" in the monthly record.

### Gate 1b · The four-step robots merge

```mermaid Figure: simply appending a named group to the end of robots pulls these crawlers out of the wildcard group, and every protection it gave them stops working
flowchart LR
  add["Append a named group"] -->|the most specific group wins| only["That crawler now reads only its own group"]
  only -->|every other group is ignored| lost["The wildcard group's Disallow stops applying"]
  lost --> leak["Admin paths, parameter pages enter AI retrieval"]:::warn
```

What gets let out is usually paths like `/wp-admin/`, `/cart/` and `/?s=`. The consequence isn't a stalled ranking — it's admin paths and parameter pages ending up in AI retrieval.

```steps Figure: four steps to merge robots, never reverse the order; skip step 3 and you've lifted the protection that was already there
1 | Copy the wildcard group | Copy the Disallow lines word for word, record the line count N
2 | Write the retrieval groups | 12 named groups, each with Allow: /
3 | Copy the N lines | Into every named group, until the count reaches N × 12
4 | Finish up | List training groups separately, put the Sitemap line at the end
```

The `Sitemap:` line goes at the end because it doesn't belong to any group. Whether to allow the training crawlers is your own call, but any training group that says `Allow: /` needs those same N lines too. **The receipt must record this sentence word for word:**

> Original wildcard `Disallow` lines = **N**, copied line by line into **N × 12** places (12 named retrieval groups); training groups with `Allow: /` also copied.

**If this sentence isn't on the receipt, gate 1 doesn't count as passed.** The only full text of the robots master template — 12 retrieval groups plus 6 training groups — lives in （→ 通用版 B.1 开门模板）; hand it to whoever edits the site and have them copy straight from there. The `<N lines of Disallow>` in the master template is a placeholder and must be replaced with the original text copied down in step 1 — not one line short, and in every group.

**Won't edit robots → all site-layer work is downgraded to read-only monitoring.**

> `Perplexity-User` does not go in the master template. Perplexity's own crawler documentation says: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules." (https://docs.perplexity.ai/guides/bots). Like `ChatGPT-User`, robots can't control it; the only way to allow it is to add its official IP range in the WAF (published on the same page as https://www.perplexity.com/perplexity-user.json).

### Gate 1c · WAF: the real switch for ChatGPT-User

```split Figure: for ChatGPT-User, Allow in robots is only a declaration — the real switch is in the WAF
Write Allow: / in robots || Add the official IP range to the WAF allowlist
Harmless, but has no effect on ChatGPT-User || The real switch, verifiable in the logs after 7 days
OpenAI's own words: robots.txt rules may not apply || Name the rule allow-chatgpt-user
Fix only this side and you'll think the door is open || The live-fetch path is actually still closed, ChatGPT-User hits stay at 0
```

OpenAI's own complete sentence is: this kind of fetching is triggered by a user action, "Because these actions are initiated by a user, robots.txt rules may not apply". On the WAF side, do three things the same day, and screenshot each one:

1. A screenshot of the "Block AI bots" / "Bot Fight Mode" switch's state.
2. Rate rules: move verified crawlers out of the rate rules, screenshot it, and attach the 429 share from the gate-0 table.
3. Allowlist: add the IP range for `openai.com/chatgpt-user.json`, name the rule `allow-chatgpt-user`, screenshot it.

**D+7 recheck**: re-pull the logs and check whether `ChatGPT-User` has any hits; on the same day, re-pull the four receipts from 2.3 as well and keep them side by side with the D1 set.

## 2.7 Gates 3 and 4: indexing paths, and JSON-LD sealed once

**What you'll do in this section**: handle indexing as two separate legs — use IndexNow for the Bing leg, and profile claiming plus pasting the URL for a quote-back on the ChatGPT leg; then spend half a day sealing the site's JSON-LD once, after which it never goes into a per-page check again. When you're done, you'll have two indexing receipts and one sealed schema template.

### Gate 3 · Indexing in two separate cells

```split Figure: indexing is two separate cells, never merge them — being indexed on Bing doesn't prove the ChatGPT leg has read you
Bing leg || ChatGPT leg
Check indexing page by page with site:, submit any gaps via IndexNow || Foursquare claiming: pick categories word for word from the official category tree
Check IndexNow Insights in Bing Webmaster Tools || Get the Foursquare business status right; claim Yelp too
A submission-to-indexing report exists and can serve as a receipt || Paste the URL and have ChatGPT quote the price line back word for word
Copilot uses this leg directly || Local questions mainly go through Foursquare, Yelp and OpenAI's local index
IndexNow only feeds Bing and Yandex || The criterion is identity d from the four identities; Bing indexing is not a criterion here
```

"ChatGPT runs through Bing's index" no longer holds as of 2026-09: ChatGPT has multiple retrieval sources, and Bing is only one of them (evidence in （→ 通用版 A.1 机制、门与人这一层的证据）). So still do IndexNow, but it isn't the switch for the ChatGPT leg. Claiming a profile is itself a change, and is only done once the baseline is saved; for the full method for the five business profiles, see （→ 通用版 3.4 第三方资质分级、Wikidata 与五处商家档案）. Having `Applebot` allowed, and verified, is a precondition for claiming your Apple Business Connect profile.

### Gate 4 · JSON-LD sealed in half a day

```mermaid Figure: text only gets read by live fetching once it's in visible HTML; schema only takes the indirect route through the Google Knowledge Graph
flowchart LR
  f["The same set of fields"] -->|write first| vis["Visible HTML"]
  vis -->|live fetching can read it| ai["Enters AI answers"]:::hl
  f -->|write second, in passing| sc["schema and sameAs"]
  sc -->|indirect, Google leg only| kg["Google Knowledge Graph"]
  sc -->|live fetching| no["None of them get read"]:::warn
```

Reverse the order and you end up with a site whose fields are all complete but that AI still can't read. Even schema's small indirect effect only exists on the condition that the same fields are already written into visible HTML; do it on its own and the return is close to zero. After adding JSON-LD, AI Overviews (AIO) citations moved −4.6%, the only significant result, and it's negative; the other two platforms weren't significant (evidence in （→ 通用版 A.1 机制、门与人这一层的证据）), so this is a one-off task:

| Rule | What it means |
|---|---|
| **One-off template** | Organization (or your industry's schema subtype) + a Person for every practitioner + sameAs. Set it up once when the site is built, **sealed in half a day** |
| **Not written per page, not checked per page** | Doesn't go on the per-page launch checklist, and doesn't go on the review checklist |
| **Only four things need to be right** | ① the legal name matches `/facts` word for word ② the address and phone number match the five business profiles word for word ③ each practitioner's registration number ④ the sameAs string |
| **A confidence note** | Write one line in the template file — "**indirect, Google leg only, low confidence**" — so the next person doesn't mistake it for a main lever |

The sameAs string links these: regulator register or registration-number lookup pages, association profile pages, Google and Bing business pages, ORCID / Scholar, conference speaker pages, LinkedIn, the Wikidata QID.

Two rules in passing: ❌ don't stack `FAQPage` schema (it's built for rich results and doesn't feed AI citations); ❌ don't add a self-reported `aggregateRating` (on the strictly regulated side this also runs into the ban on testimonials and ratings: Statute text in the example industry). **Do build question-style H2s and a short Q&A at the foot of the page; don't stack schema.**

**The one exception**: when structured data is used as a **data channel** (for example, a product feed feeding a shopping graph), it is useful — but that's a different route, one that doesn't go through live fetching. "Structured data is useful" is only true in this sense, and it is no reason to change anything on your landing pages. For the approved wording to give a client on this point, see （→ 通用版 B.2 口径句与话术）.

## 2.8 The launch gate for every page

**What you'll do in this section**: on the day each page goes live, run it through three checks, then do a quote-back test on day 7; only once every check passes do you mark it "live", and log that page's split day. When you're done, you'll have a timeline for every page from launch to D90, and a page register logging the launch date.

```steps Figure: a page from launch to D90; miss any one of the first four items and you may not mark it "live" — this page's split day is logged as the launch day, not the day work started
Same day | New title visible | Incognito window + hard refresh, confirm with your own eyes
Same day | The four copies match | The same body text and price are present in all four HTML copies
Same day | Submission receipt | IndexNow ping, covers the Bing leg only
D7 | Quote-back test | ChatGPT quotes back one exclusive fact word for word
Mark "live" | Log the split day | The split day is the launch day, not the day work started
D14 | Long-form video | The video gate; miss it and the next page can't start
D30/60/90 | Compared like for like | Same question pool, same engine, same number of passes
```

What happens if each step is missed:

- Miss **new title visible**: whoever opens the page sees the old version.
- **The four HTML copies match**: fetch the source once with each of four UAs — `OAI-SearchBot`, `Bingbot`, `Googlebot`, an ordinary browser — and the same body text and the same price string must be present in all four. Missing this has two possible outcomes: showing machines a different version gets judged as manipulation, or the price never makes it into the answer at all. What this step checks is "are the four copies the same version", which is a different thing from the four identities used to judge the door in 2.3.
- The **D7 quote-back test** is the real acceptance check for the ChatGPT leg: ask ChatGPT once, using the exact wording, about one exclusive checkable fact on that page, or paste the URL and have it quote that line of figures back word for word; screenshot it for the record. n = 1 is valid here too — it tests "was it read or not", not "what position it ranks".
- **D14 long-form video** doesn't hold up "live" — it holds up starting the next page. It belongs to writing pages' three mechanical gates; specs in （→ 通用版 5.4 三十天写页顺序与三道机械闸） and （→ 通用版 6.7 本人口述长视频）.
- **D30 / D60 / D90** must be measured to the same spec at all three points: same question pool, same engine, **the same number of passes**. Seat counts scale linearly with passes — baseline at 5 rounds, D90 run at 10 rounds, and the count doubles even if you did nothing at all; for why, see （→ 通用版 7.5 量具纪律：换尺子就不可比）.

When each page goes live, also log 2–3 checkable fact strings that are unique to it, and look for them again during each monthly retest — this is the only signal in the whole process that can give you page-level cause and effect; for how, see （→ 通用版 7.2 噪声带、页级信号与每月十步）.

## 2.9 Door-layer troubleshooting and this chapter's acceptance checks

**What you'll do in this section**: whenever a monthly retest shows a seat increase that doesn't exceed the noise band, start checking from the door layer, and use the two diagnostic figures to map every symptom to one fix; check it even when seats went up. Only once every one of the ten items on the checklist is ticked does the door count as done.

**When to check**: two points in time. ① Whenever the seat increase doesn't exceed the noise band, start checking here — **never skip a layer** (for the rest of the triage layers, see （→ 通用版 7.3 没动分诊与下月三个点）). ② **Check the door layer once a month, unconditionally, even when seats went up**: its failures are silent — while a page is not being read at all, seats can perfectly well be rising for some other reason, and you won't discover a whole wasted month until the next one.

```mermaid id=door-triage-edge Figure: at the logs / robots / WAF layer, every symptom maps to exactly one fix
flowchart LR
  t["Door-layer symptom"] --> s1["OAI-SearchBot hits = 0"]
  s1 -->|fix| f1["Gate 1's three small steps, 2.6"]
  t --> s2["403 + 429 over 5%"]
  s2 -->|fix| f2["Allowlist it and move it out of the rate rules"]
  t --> s3["All 200 but only the homepage hit"]
  s3 -->|fix| f3["Sitemap, homepage internal links, IndexNow"]
  t --> s4["Admin paths showing up in search"]
  s4 -->|fix| f4["Redo the merge, get the count to N × 12"]
  t --> s5["No movement on the Apple profile"]
  s5 -->|fix| f5["Allow and verify Applebot first"]
  t --> s6["Blaming GPTBot being blocked for not being cited"]
  s6 -->|verdict| f6["Wrong call — don't use it as a criterion"]:::warn
```

```mermaid id=door-triage-render Figure: at the rendering / indexing / nosnippet layer — Google has seats while ChatGPT is zero: check CSR first, not the choice of questions
flowchart LR
  s1["Google has seats, ChatGPT is zero"] -->|check first| csr["CSR dependency"]:::hl
  s2["Four identities: c open, b not"] -->|verdict| csr
  csr -->|fix| ssr["Open an SSR or prerendering task"]
  s3["Four identities: b open, d not"] -->|verdict| waf["WAF, go back to the edge-layer figure"]
  s4["Seats drop after a redesign or plugin change"] -->|verdict| ns["nosnippet has come back"]
  ns -->|fix| m1["Revert all three to max-snippet:-1"]
  s5["Seats haven't moved, cause unclear"] -->|in order| five["Check five things, see below"]
  five -->|any one fails| door["It's a door problem, not the questions"]:::warn
```

When the "cause is unclear", check these five things in order: ① whether the access logs show a real `OAI-SearchBot` hit, and what status code it returned ② whether the three nosnippet controls have been switched back on by a CMS or plugin ③ whether the four HTML copies match word for word ④ Bing indexing checked with `site:` ⑤ paste the URL into ChatGPT and have it quote the price line back word for word.

Rules outside the two figures:

- **Anything that fails: fix it the same day, recheck after 7 days, and don't change topics or add new content pages this month.**
- **Never edit content for the 403 / 429 cell.**
- **Never skip past the door layer and go straight to writing "demand doesn't exist". Stuck on the same layer, unfixed, two months running is an execution problem, not a problem with the questions.**
- **All five gates are mandatory in every industry.** Two more get added depending on the kind of site: a **CSR check** (for content rendered by front-end components such as price lists, course schedules, doc sites, filters and pricing widgets — fetch the HTML with a browser UA and check whether the body text is there); a **visible-text check** (raw-fetch with a retrieval-leg UA, strip the tags, and count the visible body characters and the number of stand-alone statement sentences — a page with numbers but no sentences can't get onto the answer shelf at all).

> **Example** (e-commerce) The visible-text check is mandatory: a page with fewer than 12 statement sentences may not be marked "live". For a measured case where a category page's HTML was large but, once the tags were stripped, only a small scrap of visible body text remained, see A.1.

### Acceptance checklist for this chapter

The door is a multiplier, not a variable: it only has two values, 0 and 1, and with no evidence it counts as 0. Tick these off, and only once every box is ticked does the door count as open:

- [ ] All four columns of the *Crawler hit table* filled in, saved as `crawler_hits_30d.csv`; where logs couldn't be obtained, a red blocker is listed with the catch-up date noted
- [ ] All three `nosnippet` greps come up empty, one screenshot for each of the 8 pages
- [ ] `robots.txt` has all 12 retrieval-class named groups, and the receipt has this written on it: "**Original wildcard `Disallow` lines = N, copied line by line into N × 12 places**"
- [ ] The `Sitemap:` line is at the end of the file
- [ ] The WAF allowlist rule `allow-chatgpt-user` has been added, with a screenshot on file; the 7-day recheck shows `ChatGPT-User` hits
- [ ] Each of gate 2's four criteria has **its own separate receipt**, none merged into one; the two `curl` reads are each screenshotted and marked "for troubleshooting only, not a pass criterion"
- [ ] Both dashboards — Bing Webmaster Tools AI Performance and GSC's generative AI report — are connected, each with a screenshot
- [ ] The JSON-LD template is sealed in half a day, and is never written per page after that
- [ ] Every page that has gone live has passed the launch gate (three checks the same day + the quote-back test on day 7), with the launch date logged
- [ ] The door-layer troubleshooting checklist has been added to the standing monthly actions

Every part of the checklist that's about "recording the current state" is done on D1; every item that actually changes the door is only done once the baseline is saved, and the day you fix the door is already logged as the split day.