The door is a multiplier: if crawlers cannot get in, however many pages you write afterwards get multiplied by zero. With nosnippet switched on, the Google leg is counted as zero for the whole quarter; put the price in JS, and mainstream AI crawlers get not a single word of it. This whole chapter is free and takes half a day to a day, but you do it in two passes. 2.1–2.5 is the D1 read-only audit: get access, record the current state, and change nothing. Only after the baseline has been saved, following → General Edition 4.4 The web control leg and the frozen baseline (the book's only full spec), do you do 2.6–2.9 and fix the door. For how the three gates — door, identity, shelf — relate to each other, see → General Edition 1.3 Three gates and three paths.
Hard gate: until the baseline is saved, do not change the door, any profile or the facts page.
2.1 The five gates at a glance, and gate 0: access logs (read-only)
What you'll do in this section: first see clearly why the door is done in two passes; on D1, the first move is to send the access request, pull the Crawler hit table, and use the four ways of reading it to log problems as to-fix items. When you're done, you'll have a hit table and a to-fix list, and you won't have touched the door.
Why it has to be this way: what the baseline records — page-type mix, the four states per URL, question shape — is all measured on the shelf "while the door is still as it was". Fix the door first and freeze second, and your own changes are already mixed into the before sample; after that, whenever something "went up", you can't tell whether the door did it or it would have risen anyway. So the build order (door first) and the measurement order (freeze first) are two different things — never mix them. On D1 there are only three kinds of thing you can do that do not count as a change: request access, turn on logging, record the current state. The book's week numbers all count from the day you fix the door; see → General Edition 0.1 The book in one sentence, and the 90-day reading order.
- Gate 0Access logsD1 pulls the table; if you can't get it, open a blocker
- Gate 1aSnippet switchesThe three nosnippet controls, D1 greps, 2.6 fixes
- Gate 1brobotsD1 records the six lines, 2.6 does the four-step merge
- Gate 1cWAFD1 checks 403/429, 2.6 adds the allowlist
- Gate 2Four identities, fetched liveD1 keeps one receipt for each
- Gate 3Indexing and dataD1 connects the dashboards, 2.7 submits
- Gate 4JSON-LD2.7 sealed in half a day
Two ordering rules:
- Gate 0 comes before every other gate. It is the only direct evidence, and the only piece of work where "the doing takes half an hour, but getting access can take days". In your first email, ask for all three kinds of access at once: domain verification, permission to edit robots, and permission to export access logs.
- Gate 3 connects the free first-party data before you spend money on probes (see 2.5). Do it the other way round and you'll pay to "discover" something that was already free.
How to pull gate 0
Faking a UA, a site: query, eyeballing it in a browser — all of these are indirect evidence, and each can be wrong in either direction. Without access logs, the whole technical door layer rests on inference.
- Pull 30 daysRaw logsCloudflare Logpush or the host's access.log
- Get IP rangesOfficially published rangesOpenAI, Anthropic, Perplexity
- CIDR filterIdentify by IPNot by UA string
- Produce the hit tableFour columnsSave as crawler_hits_30d.csv
OpenAI's three IP ranges are openai.com/searchbot.json, openai.com/gptbot.json and openai.com/chatgpt-user.json; Anthropic and Perplexity each publish their own official ranges. A hit count based on UA strings has scanners and spoofed traffic mixed into it, which invalidates the whole thing as evidence. For the template of the four-column hit table (crawler name, hits in 30 days, top 20 hit URLs, status-code distribution), see → General Edition B.1 Door templates.
Two fallbacks when you can't get it: no export access → open a task the same day and list a red blocker, and let gate 2 stand in with criteria b–d for now (see 2.3); the site never kept logs at all → turn on Logpush / access logs the same day, and the first table only exists after 30 days.
How to read the hit table
Three rules outside the figure:
- At the layer where
OAI-SearchBothits are 0, do not write "wrong topic", and do not skip ahead to the "change topic" layer. If the door isn't open, or the site simply hasn't been discovered, that's a door problem. - 403 / 429 mean the WAF or a rate rule is blocking, not robots. Editing robots will not fix it.
- This table is three things at once: gate 2's criterion a, the first check in the monthly troubleshooting pass, and the only direct evidence of whether a given leg is actually open. When you can't get the logs, don't say "the door is open"; note the catch-up date on the first line. Saying you have no evidence when you have none is a book-wide rule of how you write things up (see → General Edition D.1 Number discipline: how to label numbers, and what stays internal).
For how to fix each reading, see the two diagnostic figures in 2.9.
2.2 Gate 1 read-only audit: what to check in nosnippet, robots and the WAF
What you'll do in this section: spend 10 minutes grepping the three nosnippet controls, then spend half an hour recording the current state of robots and the WAF in six lines. When you're done, you'll have a Six-line door read and 8 screenshots — record only, change nothing; every item you record gets fixed in 2.6.
Check nosnippet first, then robots
Google's own documentation states it plainly: to restrict how a page's content shows up in Search (including AI Overviews and AI Mode), you use nosnippet, data-nosnippet, max-snippet and noindex (developers.google.com/search/docs/crawling-indexing/robots-meta-tag). This is the one switch on the Google leg that can zero you out with a single click, and CMS templates and SEO plugins often turn it on by default. It takes 10 minutes to check — check it first.
Do this on 8 pages: the homepage, the price page, the /facts page, plus 5 template pages (a service page, a practitioner's or author's person page, an old blog post, a list page, a landing page). curl each page for the HTML plus the response headers.
D1's receipt: the results of the three greps (all empty, or which one hit and on which page), one screenshot per page for all 8. This switch gets turned back on after a redesign, a template change or a plugin update, so it is also a standing item in the monthly door-layer check (see 2.9).
⚠️ Many people treat
Google-Extendedas the AI Overviews switch and spend effort arguing over whether to allow it, while missing this real switch. For why it isn'tGoogle-Extended, see the four misconceptions in 2.4.
robots and WAF: record six lines
- 1Retrieval UAsRecord each of the 12, allowed or blocked
- 2Wildcard group NN Disallow lines, copy them out line by line, word for word
- 3Training UAsAll 6, mark every one "unrelated to whether you get cited"
- 4Snippet switchescurl -I the homepage and the price page, check all three places
- 5CDN/WAFBlock AI bots, 429s, allowlist
- 6Access logsWhether you have them, the date you requested export access
The three failure paths are: skip line 2, and merging robots later strips away the protection that was already there (2.6); miss line 4, and nosnippet stays on with nobody noticing, so the Google leg counts as zero for the quarter; skip asking about line 6, and gate 0 can't start work the same day, leaving the whole door layer as nothing but inference — which is why this line should be asked as early as possible.
Line 5 needs three things nailed down: ① whether the "Block AI bots" / "Bot Fight Mode" switch is on, off, or doesn't exist; ② whether rate rules hit crawlers with 429s (check against the 429 share in the gate-0 table); ③ whether the IP range for openai.com/chatgpt-user.json is on the allowlist.
Where to check each line: for lines 1 and 3, open https://<domain>/robots.txt and read it group by group; for line 5, check the CDN / WAF dashboard; for line 6, ask ops and record "has Logpush" / "has access.log" / "kept no logs". The six lines take roughly 10, 5, 5, 10, 10 and 5 minutes in turn. For the list and purpose of the 12 retrieval UAs and 6 training UAs, see 2.4; for a blank record sheet, see → General Edition B.1 Door templates.
⚠️ Allowed in robots does not mean actually let through — the CDN / WAF layer can still return a straight 403 regardless of robots (inferred, 85% confidence). "Allowed" from the audit is only the paper-level first layer; the only evidence of actually being let through is gate 0's Crawler hit table.
2.3 Gate 2: four identities, fetched live (read-only)
What you'll do in this section: stop judging the door by looking at content through curl -A, and instead look at the same URL through four paths, then read the differences between the four to tell CSR, WAF and a closed door apart. When you're done, you'll have four separate receipts — never merge them into one.
curl -A "OAI-SearchBot" only proves how your infrastructure reacts to that string. The kind of pass-through on the left isn't hypothetical: Cloudflare Verified Bots really does allow by IP + rDNS.
Four identities: one URL as four paths see it
Ranked by how trustworthy each one is; all of them are free.
| Identity | Where you get it | Pass criterion | What it can prove |
|---|---|---|---|
| a · What the real crawler sees | Gate 0's access logs, CIDR-filtered against the official IP ranges such as openai.com/searchbot.json | This URL was hit by OAI-SearchBot within the last 30 days and returned 200 | The only direct evidence, applies across every leg |
| b · The HTML Bing fetched | Bing Webmaster Tools → URL Inspection → the source returned by Live Test | You can string-search the source and find the price figure and the body text | The Bing / Copilot leg; also the free sample closest to "what a non-rendering crawler sees" |
| c · The HTML Google fetched | Google Search Console (GSC) → URL Inspection → the HTML under "Crawled page" | Same as above, string-searchable | Only proves Google's side (Google renders), never use it alone as your conclusion |
| d · What ChatGPT's live-fetch path sees | Paste this URL into ChatGPT, verbatim: "Open this link and copy out, word for word, the line on the page that gives the price." | It can quote it back word for word (not paraphrase it, not estimate it) | The ChatGPT-User live-fetch leg is open |
curl is not a fifth identity — it's a troubleshooting tool, and it has to be split into two separate reads: ① fetch the HTML with an ordinary browser UA and check whether the body text and price are there — this is the right way to judge CSR, and CSR has nothing to do with which UA you use; ② run it again with the OAI-SearchBot UA, look only at the status code, not the content, to judge the WAF. Screenshot each of the two separately, and write on the receipt "for troubleshooting only, not a pass criterion". The half-day quick-start version (→ General Edition 4.8 Two shortcuts and this chapter's checklist) also judges JS following step ①.
How to read the differences between the four
Outside the figure, also note:
- "Pass, degraded" must have "no direct evidence, fill in a after 30 days" noted on the first line.
- CSR dependency = Google's rendering can get it, but a non-rendering crawler cannot. Log it as a to-fix item and open an SSR / prerendering task the same day; push all page work back as a whole, and add no new content pages this month.
- Never edit content for the WAF cell — that belongs to the WAF allowlist in 2.6.
- For the door-not-open cell, never write "demand doesn't exist", and never skip ahead to "change topic".
D1 is the audit. On day 7 after fixing the door (2.6), pull all four receipts again and lay them side by side against the D1 set.
2.4 What each crawler is for, and judging the door leg by leg (read-only)
What you'll do in this section: tell the 12 retrieval crawlers and the 6 training crawlers apart, and know which leg breaks when you block which one; when you hit a platform, multiple subdomains, a login wall or a multilingual site, judge the door leg by leg. When you're done, you'll have a Per-leg door table: one row per host, one column per leg — never merge into one cell.
This is the single easiest thing in the whole chapter to get wrong: block the wrong one and nothing happens; allow the wrong one and the whole quarter is wasted.
12 retrieval legs: what blocking each one costs
| Crawler | Belongs to | What blocking it costs |
|---|---|---|
OAI-SearchBot | OpenAI | The ChatGPT leg is zero for the quarter |
ChatGPT-User | OpenAI | Users who open your page from inside ChatGPT can't read it. ⚠️ The real switch is in the WAF, not robots (2.6) |
OAI-AdsBot | OpenAI | Only affects ad scenarios; if you don't run ads, you don't have to allow it |
Bingbot | Microsoft | Copilot breaks outright; one of ChatGPT's eight retrieval sources goes missing |
Googlebot | You lose AI Overviews / AI Mode and organic rankings together: all three share the same crawler | |
Applebot | Apple | The whole Siri / Spotlight / Apple Maps / Apple Intelligence leg breaks. ⚠️ Until this group is allowed, don't claim your Apple Business Connect profile |
Amazonbot | Amazon | Alexa / Amazon-side retrieval breaks |
PerplexityBot | Perplexity | The Perplexity leg breaks |
ClaudeBot | Anthropic | Claude-side access breaks |
Claude-SearchBot | Anthropic | Claude retrieval breaks |
Claude-User | Anthropic | Users who open your page from inside Claude can't read it |
DuckAssistBot | DuckDuckGo | DuckAssist breaks |
The 6 training legs: GPTBot · Google-Extended · Applebot-Extended · Meta-ExternalAgent · CCBot · Bytespider. Whether you allow them is a content-licensing question, not a visibility question.
The real danger is the flip side of the first one: blocking OAI-SearchBot and Bingbot themselves. The second is the most expensive to get wrong, and it rests on two of Google's own sentences. The crawler documentation: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." AI features and your website nails down the entry requirement: "a page must be indexed and eligible to be shown in Google Search with a snippet", and "robots.txt directives for Googlebot is the control". So the gate for AI Overviews is Googlebot plus the three nosnippet controls. The point of the fourth is that it saves you work: for the three Anthropic bots (ClaudeBot / Claude-SearchBot / Claude-User), editing robots genuinely takes effect — you don't need to touch the WAF. 88.2% comes from BuzzStream; for the rest of the sources, see → General Edition A.5 List of sources.
The zero-JS-execution corollary: mainstream AI crawlers do not execute JS; the only exceptions are Gemini (which runs on Google's infrastructure and renders in full) and Applebot (evidence in → General Edition 1.1 How buyers ask, and what AI reads). So the price, the core facts and the body copy must be in HTML that's already been rendered server-side — fetch the source with a browser UA and you should be able to string-search it; write them in JS instead and the next 90 days are wasted. The flip side of the same corollary: live fetching doesn't read JSON-LD, so schema is a baseline, not a lever (2.7). This exception also gives the whole book its single most useful troubleshooting criterion: the Google leg has seats while the ChatGPT leg is zero — that difference by itself is the fingerprint of CSR.
Judging the door leg by leg
Recording rules outside the figure — skip none of them:
curlthe whole of robots.txt yourself, download it, parse it group by group, and log it by date. Grepping for one keyword alone will miss the single most important thing: that this UA was never given its own named group at all.- Check robots on the host that the page actually lives on, not the main domain. Detail pages, doc sites and help centres are often on a different host or subdomain, and robots is not shared with the main domain.
- One row per host, one column per leg — never merge into one cell. The very same platform can easily have some legs blocked and others open; merge it into one cell and the whole chain gets thrown away.
Google-Extended,Applebot-ExtendedandMeta-ExternalAgentare training / licensing legs — they don't go on this table.- The table must carry a "verification level" column:
robots-verified/fetch-verified. A raw-fetch receipt may only say "fetched 200 + body text found" or "fetch blocked (<status>)", never "allowed". robots only lifts a protocol-level ban; CSR rendering and WAF / geo-blocking (403s, captchas, redirects) are two other, separate gates, and they are not written in robots. - A platform's robots can change, so log it in the Off-site asset decay table and rerun it once a quarter. A refused connection, a timeout or a failed TLS handshake all get logged as "not retrieved"; if a retry still fails, leave it blank — never guess and fill it in.
- A platform closing its door ≠ that platform has no sales value. All that's ruled dead is "citations through that one leg", not the business; other legs, and the shopping shelf that runs through data feeds, may still show up as normal. Never mix these two things up in anything you tell clients.
On a host where a given leg is blocked, schedule no content work on that leg, and don't let acceptance count it towards seats. What's being judged here is only whether the door is open, which is a different thing from → General Edition 4.6 Four states, two denominators and the abstain line judging whether "a given URL can be reached at all". For reproducible, per-platform robots commands, see → General Edition B.1 Door templates.
Example (training) Measured on 2026-09-21, the robots.txt of TPGateway, Singapore's official training portal:
User-agent: *→Disallow: /; theAllow: /groups that follow list onlyGooglebot(including Image / Video / mobile),Slurp,bingbot,Applebot,TwitterbotandPingdom. Not one ofGPTBot/OAI-SearchBot/ChatGPT-User/ClaudeBot/PerplexityBotappears — all of them fall underDisallow: /. → The Google, Apple and Bing legs are open; the OpenAI, Anthropic and Perplexity legs are blocked: build for the first three legs only, accept work against the first three legs only, and don't count ChatGPT seats.
Five kinds of site, one extra check each
| Kind of site | Extra check | Criterion and how to handle it |
|---|---|---|
| Single-page app / client-side rendering | Fetch the HTML with an ordinary browser UA, grep for one actual sentence of body text | Not found = the body text isn't in the HTML. Move key body text to server-side rendering or prerendering — no "wait and bundle it with the next SEO redesign" |
| Multiple subdomains | curl a robots.txt for each subdomain separately, one cell per subdomain × retrieval leg | For any subdomain with no receipt, its pages don't count as "door open" this quarter and can't be counted towards deliverables. Doc sites, help centres and trust pages need their door opened and to be server-side rendered too |
| Login wall / quote wall | Whether the key figures sit outside the wall | Anything behind the wall is invisible to AI: build a public, checkable summary page or a mirror on the main domain |
| Multilingual sites | Run the four-identity check once for each language | Never check only one language |
| Price or stock that changes very fast | Whether these fields are injected by a client-side script | Yes = as good as not there |
Judge the two causes separately, and don't reverse the order: check the robots receipts first (one curl per subdomain × leg — the cheapest check). When one subdomain was missed, the blocked leg reads zero and the unblocked legs read normal — a one-line fix; only check CSR once robots is all clear. Reverse the order and you'll schedule a whole quarter of engineering rework for a problem that a single line in robots would have fixed. Only once both are fixed does that leg go from 0 to 1.
Example (B2B software) A site naturally has multiple subdomains:
www/docs/help/blog/status/app/trust. Check robots for each subdomain separately — a clean main domain doesn't meandocsis clean too.
2.5 Gate 3, read-only: connect the two free AI reports
What you'll do in this section: on D1, connect GSC's generative AI report and Bing Webmaster Tools' AI Performance, actually read each one once, and check that the site field they return really is your own site. When you're done, you'll have two sets of data that are the only free, first-party, unfakeable data in the whole engagement; only open the probe budget after you've read them.
- 1Request accessDomain verification, in the same email as the log access request
- 2Open the GSC reportGenerative AI performance, export the CSV and screenshot it
- 3Open the Bing reportWebmaster Tools AI Performance, same as above
- 4Actually read it onceCheck the site field — get it wrong and you'll silently read another site
- 5Save itTwo connection screenshots; how to read them is in 4.2
Three rules:
- Get the site field wrong = this box isn't done — don't mark it complete. When the site parameter isn't configured correctly, the tool will read a different site and still return ok, which is worse than not connecting at all: you'll be treating someone else's numbers as your own baseline. The same goes for other read-only dashboards such as GA4.
- Can't get domain-verification access → list a red blocker. It blocks more than just this box: Bing's indexing submission (2.7), and Bing Live Test and GSC's crawled-page view (identities b and c in 2.3) all need domain verification first.
- This section is connect-and-read only. How to read each of the two reports, what each one gives you, and what to watch out for are written up once, in → General Edition 4.2 The question pool: where the 30 questions come from, how they are balanced, how they are signed; checking Bing for indexing gaps and submitting via IndexNow are changes, and belong in 2.7.
Another free thing to start the same day, in passing, is copying out the actual words buyers used, from your front-desk records or CRM, for the last 15 deals before they closed — it's the question pool's only class A source, and how to do it is also in 4.2.
2.6 Fixing the door: nosnippet, the four-step robots merge, the WAF allowlist
What you'll do in this section: first confirm the baseline is saved, then on the same day make all the fixes in the order nosnippet → robots → WAF, and log the day you fix the door as the split day. The hard gate still applies: until the baseline is saved, do not change the door, any profile or the facts page. When you're done, you'll have a merged robots.txt, a receipt with the N × 12 line written on it, and one WAF allowlist rule.
- PreconditionBaseline savedBaseline Top 20, noise band, the 36 brand-six runs
- Gate 1anosnippetRevert all three places, grep comes up empty
- Gate 1brobotsFour-step merge, receipt has the N × 12 sentence, word for word
- Gate 1cWAFAdd the chatgpt-user.json IP range
- Log split dayThe day you fix the doorWritten into the page register
- D+7Re-pull and recheckCheck the logs for ChatGPT-User, re-pull the four identities
The day you fix the door is the split day for the whole engagement; every "before vs after fixing the door" comparison is drawn against it. Each page also has its own split day: the day it goes live (2.8). The book's W1 also counts from this day.
Gate 1a · nosnippet
Any hit left unfixed for four weeks → the Google leg counts as zero for the quarter — write it on the first line, and mark the Google cell "switch not lifted" in the monthly record.
Gate 1b · The four-step robots merge
What gets let out is usually paths like /wp-admin/, /cart/ and /?s=. The consequence isn't a stalled ranking — it's admin paths and parameter pages ending up in AI retrieval.
- 1Copy the wildcard groupCopy the Disallow lines word for word, record the line count N
- 2Write the retrieval groups12 named groups, each with Allow: /
- 3Copy the N linesInto every named group, until the count reaches N × 12
- 4Finish upList training groups separately, put the Sitemap line at the end
The Sitemap: line goes at the end because it doesn't belong to any group. Whether to allow the training crawlers is your own call, but any training group that says Allow: / needs those same N lines too. The receipt must record this sentence word for word:
Original wildcard
Disallowlines = N, copied line by line into N × 12 places (12 named retrieval groups); training groups withAllow: /also copied.
If this sentence isn't on the receipt, gate 1 doesn't count as passed. The only full text of the robots master template — 12 retrieval groups plus 6 training groups — lives in → General Edition B.1 Door templates; hand it to whoever edits the site and have them copy straight from there. The <N lines of Disallow> in the master template is a placeholder and must be replaced with the original text copied down in step 1 — not one line short, and in every group.
Won't edit robots → all site-layer work is downgraded to read-only monitoring.
Perplexity-Userdoes not go in the master template. Perplexity's own crawler documentation says: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules." (https://docs.perplexity.ai/guides/bots). LikeChatGPT-User, robots can't control it; the only way to allow it is to add its official IP range in the WAF (published on the same page as https://www.perplexity.com/perplexity-user.json).
Gate 1c · WAF: the real switch for ChatGPT-User
OpenAI's own complete sentence is: this kind of fetching is triggered by a user action, "Because these actions are initiated by a user, robots.txt rules may not apply". On the WAF side, do three things the same day, and screenshot each one:
- A screenshot of the "Block AI bots" / "Bot Fight Mode" switch's state.
- Rate rules: move verified crawlers out of the rate rules, screenshot it, and attach the 429 share from the gate-0 table.
- Allowlist: add the IP range for
openai.com/chatgpt-user.json, name the ruleallow-chatgpt-user, screenshot it.
D+7 recheck: re-pull the logs and check whether ChatGPT-User has any hits; on the same day, re-pull the four receipts from 2.3 as well and keep them side by side with the D1 set.
2.7 Gates 3 and 4: indexing paths, and JSON-LD sealed once
What you'll do in this section: handle indexing as two separate legs — use IndexNow for the Bing leg, and profile claiming plus pasting the URL for a quote-back on the ChatGPT leg; then spend half a day sealing the site's JSON-LD once, after which it never goes into a per-page check again. When you're done, you'll have two indexing receipts and one sealed schema template.
Gate 3 · Indexing in two separate cells
"ChatGPT runs through Bing's index" no longer holds as of 2026-09: ChatGPT has multiple retrieval sources, and Bing is only one of them (evidence in → General Edition A.1 Evidence for mechanics, the door and identity). So still do IndexNow, but it isn't the switch for the ChatGPT leg. Claiming a profile is itself a change, and is only done once the baseline is saved; for the full method for the five business profiles, see → General Edition 3.4 Third-party credential tiers, Wikidata and the five business profiles. Having Applebot allowed, and verified, is a precondition for claiming your Apple Business Connect profile.
Gate 4 · JSON-LD sealed in half a day
Reverse the order and you end up with a site whose fields are all complete but that AI still can't read. Even schema's small indirect effect only exists on the condition that the same fields are already written into visible HTML; do it on its own and the return is close to zero. After adding JSON-LD, AI Overviews (AIO) citations moved −4.6%, the only significant result, and it's negative; the other two platforms weren't significant (evidence in → General Edition A.1 Evidence for mechanics, the door and identity), so this is a one-off task:
| Rule | What it means |
|---|---|
| One-off template | Organization (or your industry's schema subtype) + a Person for every practitioner + sameAs. Set it up once when the site is built, sealed in half a day |
| Not written per page, not checked per page | Doesn't go on the per-page launch checklist, and doesn't go on the review checklist |
| Only four things need to be right | ① the legal name matches /facts word for word ② the address and phone number match the five business profiles word for word ③ each practitioner's registration number ④ the sameAs string |
| A confidence note | Write one line in the template file — "indirect, Google leg only, low confidence" — so the next person doesn't mistake it for a main lever |
The sameAs string links these: regulator register or registration-number lookup pages, association profile pages, Google and Bing business pages, ORCID / Scholar, conference speaker pages, LinkedIn, the Wikidata QID.
Two rules in passing: ❌ don't stack FAQPage schema (it's built for rich results and doesn't feed AI citations); ❌ don't add a self-reported aggregateRating (on the strictly regulated side this also runs into the ban on testimonials and ratings: Statute text in the example industry). Do build question-style H2s and a short Q&A at the foot of the page; don't stack schema.
The one exception: when structured data is used as a data channel (for example, a product feed feeding a shopping graph), it is useful — but that's a different route, one that doesn't go through live fetching. "Structured data is useful" is only true in this sense, and it is no reason to change anything on your landing pages. For the approved wording to give a client on this point, see → General Edition B.2 Approved wording and scripts.
2.8 The launch gate for every page
What you'll do in this section: on the day each page goes live, run it through three checks, then do a quote-back test on day 7; only once every check passes do you mark it "live", and log that page's split day. When you're done, you'll have a timeline for every page from launch to D90, and a page register logging the launch date.
- Same dayNew title visibleIncognito window + hard refresh, confirm with your own eyes
- Same dayThe four copies matchThe same body text and price are present in all four HTML copies
- Same daySubmission receiptIndexNow ping, covers the Bing leg only
- D7Quote-back testChatGPT quotes back one exclusive fact word for word
- Mark "live"Log the split dayThe split day is the launch day, not the day work started
- D14Long-form videoThe video gate; miss it and the next page can't start
- D30/60/90Compared like for likeSame question pool, same engine, same number of passes
What happens if each step is missed:
- Miss new title visible: whoever opens the page sees the old version.
- The four HTML copies match: fetch the source once with each of four UAs —
OAI-SearchBot,Bingbot,Googlebot, an ordinary browser — and the same body text and the same price string must be present in all four. Missing this has two possible outcomes: showing machines a different version gets judged as manipulation, or the price never makes it into the answer at all. What this step checks is "are the four copies the same version", which is a different thing from the four identities used to judge the door in 2.3. - The D7 quote-back test is the real acceptance check for the ChatGPT leg: ask ChatGPT once, using the exact wording, about one exclusive checkable fact on that page, or paste the URL and have it quote that line of figures back word for word; screenshot it for the record. n = 1 is valid here too — it tests "was it read or not", not "what position it ranks".
- D14 long-form video doesn't hold up "live" — it holds up starting the next page. It belongs to writing pages' three mechanical gates; specs in → General Edition 5.4 The 30-day writing order and the three mechanical gates and → General Edition 6.7 Long videos narrated by the named expert.
- D30 / D60 / D90 must be measured to the same spec at all three points: same question pool, same engine, the same number of passes. Seat counts scale linearly with passes — baseline at 5 rounds, D90 run at 10 rounds, and the count doubles even if you did nothing at all; for why, see → General Edition 7.5 Ruler discipline: change the ruler and nothing is comparable.
When each page goes live, also log 2–3 checkable fact strings that are unique to it, and look for them again during each monthly retest — this is the only signal in the whole process that can give you page-level cause and effect; for how, see → General Edition 7.2 The noise band, the page-level signal and the ten monthly steps.
2.9 Door-layer troubleshooting and this chapter's acceptance checks
What you'll do in this section: whenever a monthly retest shows a seat increase that doesn't exceed the noise band, start checking from the door layer, and use the two diagnostic figures to map every symptom to one fix; check it even when seats went up. Only once every one of the ten items on the checklist is ticked does the door count as done.
When to check: two points in time. ① Whenever the seat increase doesn't exceed the noise band, start checking here — never skip a layer (for the rest of the triage layers, see → General Edition 7.3 Not-moved triage and next month's three points). ② Check the door layer once a month, unconditionally, even when seats went up: its failures are silent — while a page is not being read at all, seats can perfectly well be rising for some other reason, and you won't discover a whole wasted month until the next one.
When the "cause is unclear", check these five things in order: ① whether the access logs show a real OAI-SearchBot hit, and what status code it returned ② whether the three nosnippet controls have been switched back on by a CMS or plugin ③ whether the four HTML copies match word for word ④ Bing indexing checked with site: ⑤ paste the URL into ChatGPT and have it quote the price line back word for word.
Rules outside the two figures:
- Anything that fails: fix it the same day, recheck after 7 days, and don't change topics or add new content pages this month.
- Never edit content for the 403 / 429 cell.
- Never skip past the door layer and go straight to writing "demand doesn't exist". Stuck on the same layer, unfixed, two months running is an execution problem, not a problem with the questions.
- All five gates are mandatory in every industry. Two more get added depending on the kind of site: a CSR check (for content rendered by front-end components such as price lists, course schedules, doc sites, filters and pricing widgets — fetch the HTML with a browser UA and check whether the body text is there); a visible-text check (raw-fetch with a retrieval-leg UA, strip the tags, and count the visible body characters and the number of stand-alone statement sentences — a page with numbers but no sentences can't get onto the answer shelf at all).
Example (e-commerce) The visible-text check is mandatory: a page with fewer than 12 statement sentences may not be marked "live". For a measured case where a category page's HTML was large but, once the tags were stripped, only a small scrap of visible body text remained, see A.1.
Acceptance checklist for this chapter
The door is a multiplier, not a variable: it only has two values, 0 and 1, and with no evidence it counts as 0. Tick these off, and only once every box is ticked does the door count as open:
- [ ] All four columns of the Crawler hit table filled in, saved as
crawler_hits_30d.csv; where logs couldn't be obtained, a red blocker is listed with the catch-up date noted - [ ] All three
nosnippetgreps come up empty, one screenshot for each of the 8 pages - [ ]
robots.txthas all 12 retrieval-class named groups, and the receipt has this written on it: "Original wildcardDisallowlines = N, copied line by line into N × 12 places" - [ ] The
Sitemap:line is at the end of the file - [ ] The WAF allowlist rule
allow-chatgpt-userhas been added, with a screenshot on file; the 7-day recheck showsChatGPT-Userhits - [ ] Each of gate 2's four criteria has its own separate receipt, none merged into one; the two
curlreads are each screenshotted and marked "for troubleshooting only, not a pass criterion" - [ ] Both dashboards — Bing Webmaster Tools AI Performance and GSC's generative AI report — are connected, each with a screenshot
- [ ] The JSON-LD template is sealed in half a day, and is never written per page after that
- [ ] Every page that has gone live has passed the launch gate (three checks the same day + the quote-back test on day 7), with the launch date logged
- [ ] The door-layer troubleshooting checklist has been added to the standing monthly actions
Every part of the checklist that's about "recording the current state" is done on D1; every item that actually changes the door is only done once the baseline is saved, and the day you fix the door is already logged as the split day.