Skip to main content

GEO Playbook · General Edition · Chapter 2 (3 of 17)

Let AI read your website: check first, then fix the settings

Access logs, robots.txt, firewalls and nosnippet; save how things stand before you change anything

The door is a multiplier: if crawlers cannot get in, however many pages you write afterwards get multiplied by zero. With nosnippet switched on, the Google leg is counted as zero for the whole quarter; put the price in JS, and mainstream AI crawlers get not a single word of it. This whole chapter is free and takes half a day to a day, but you do it in two passes. 2.1–2.5 is the D1 read-only audit: get access, record the current state, and change nothing. Only after the baseline has been saved, following → General Edition 4.4 The web control leg and the frozen baseline (the book's only full spec), do you do 2.6–2.9 and fix the door. For how the three gates — door, identity, shelf — relate to each other, see → General Edition 1.3 Three gates and three paths.

Hard gate: until the baseline is saved, do not change the door, any profile or the facts page.

2.1 The five gates at a glance, and gate 0: access logs (read-only)

What you'll do in this section: first see clearly why the door is done in two passes; on D1, the first move is to send the access request, pull the Crawler hit table, and use the four ways of reading it to log problems as to-fix items. When you're done, you'll have a hit table and a to-fix list, and you won't have touched the door.

change nothing

baseline Top 20, noise band

counted against the same denominator

D1 read-only audit

Freeze baseline 4.1–4.4

Baseline saved

Fix the door, log split day

Before and after comparable

Fix the door before freezing

Freezes a shelf you've already changed

Before and after not comparable

Figure: audit read-only first, freeze the baseline, then fix the door — only then are before and after comparable; reverse the order and what you freeze is a shelf you've already changed yourself

Why it has to be this way: what the baseline records — page-type mix, the four states per URL, question shape — is all measured on the shelf "while the door is still as it was". Fix the door first and freeze second, and your own changes are already mixed into the before sample; after that, whenever something "went up", you can't tell whether the door did it or it would have risen anyway. So the build order (door first) and the measurement order (freeze first) are two different things — never mix them. On D1 there are only three kinds of thing you can do that do not count as a change: request access, turn on logging, record the current state. The book's week numbers all count from the day you fix the door; see → General Edition 0.1 The book in one sentence, and the 90-day reading order.

  1. Gate 0Access logsD1 pulls the table; if you can't get it, open a blocker
  2. Gate 1aSnippet switchesThe three nosnippet controls, D1 greps, 2.6 fixes
  3. Gate 1brobotsD1 records the six lines, 2.6 does the four-step merge
  4. Gate 1cWAFD1 checks 403/429, 2.6 adds the allowlist
  5. Gate 2Four identities, fetched liveD1 keeps one receipt for each
  6. Gate 3Indexing and dataD1 connects the dashboards, 2.7 submits
  7. Gate 4JSON-LD2.7 sealed in half a day
Figure: go through the five gates in the same order twice — D1 only looks, changes nothing; once the baseline is saved, fix them in this same order

Two ordering rules:

  1. Gate 0 comes before every other gate. It is the only direct evidence, and the only piece of work where "the doing takes half an hour, but getting access can take days". In your first email, ask for all three kinds of access at once: domain verification, permission to edit robots, and permission to export access logs.
  2. Gate 3 connects the free first-party data before you spend money on probes (see 2.5). Do it the other way round and you'll pay to "discover" something that was already free.

How to pull gate 0

Faking a UA, a site: query, eyeballing it in a browser — all of these are indirect evidence, and each can be wrong in either direction. Without access logs, the whole technical door layer rests on inference.

  1. Pull 30 daysRaw logsCloudflare Logpush or the host's access.log
  2. Get IP rangesOfficially published rangesOpenAI, Anthropic, Perplexity
  3. CIDR filterIdentify by IPNot by UA string
  4. Produce the hit tableFour columnsSave as crawler_hits_30d.csv
Figure: four steps to pull the logs; identify a crawler by its IP range only, because anyone can fake a UA string

OpenAI's three IP ranges are openai.com/searchbot.json, openai.com/gptbot.json and openai.com/chatgpt-user.json; Anthropic and Perplexity each publish their own official ranges. A hit count based on UA strings has scanners and spoofed traffic mixed into it, which invalidates the whole thing as evidence. For the template of the four-column hit table (crawler name, hits in 30 days, top 20 hit URLs, status-code distribution), see → General Edition B.1 Door templates.

Two fallbacks when you can't get it: no export access → open a task the same day and list a red blocker, and let gate 2 stand in with criteria b–d for now (see 2.3); the site never kept logs at all → turn on Logpush / access logs the same day, and the first table only exists after 30 days.

How to read the hit table

verdict

verdict

verdict

verdict

write only

Crawler hit table

OAI-SearchBot hits = 0

Door not open, log to-fix 2.6

403 + 429 over 5% combined

WAF is blocking, log to-fix 2.6

All 200 but only the homepage hit

Sitemap and internal links, log to-fix 2.7

ChatGPT-User has hits

Live-fetch leg is open, log monthly

Can't get the logs

No direct evidence, fill in after 30 days

Figure: each reading on the hit table maps to exactly one to-fix item; before the baseline is saved, record only, change nothing

Three rules outside the figure:

  • At the layer where OAI-SearchBot hits are 0, do not write "wrong topic", and do not skip ahead to the "change topic" layer. If the door isn't open, or the site simply hasn't been discovered, that's a door problem.
  • 403 / 429 mean the WAF or a rate rule is blocking, not robots. Editing robots will not fix it.
  • This table is three things at once: gate 2's criterion a, the first check in the monthly troubleshooting pass, and the only direct evidence of whether a given leg is actually open. When you can't get the logs, don't say "the door is open"; note the catch-up date on the first line. Saying you have no evidence when you have none is a book-wide rule of how you write things up (see → General Edition D.1 Number discipline: how to label numbers, and what stays internal).

For how to fix each reading, see the two diagnostic figures in 2.9.

2.2 Gate 1 read-only audit: what to check in nosnippet, robots and the WAF

What you'll do in this section: spend 10 minutes grepping the three nosnippet controls, then spend half an hour recording the current state of robots and the WAF in six lines. When you're done, you'll have a Six-line door read and 8 screenshots — record only, change nothing; every item you record gets fixed in 2.6.

Check nosnippet first, then robots

Google's own documentation states it plainly: to restrict how a page's content shows up in Search (including AI Overviews and AI Mode), you use nosnippet, data-nosnippet, max-snippet and noindex (developers.google.com/search/docs/crawling-indexing/robots-meta-tag). This is the one switch on the Google leg that can zero you out with a single click, and CMS templates and SEO plugins often turn it on by default. It takes 10 minutes to check — check it first.

Do this on 8 pages: the homepage, the price page, the /facts page, plus 5 template pages (a service page, a practitioner's or author's person page, an old blog post, a list page, a landing page). curl each page for the HTML plus the response headers.

Where to checkWhat a hit looks like
meta robots / googlebot in the HTMLnosnippet, or max-snippet:0 / :20
The HTTP response header X-Robots-TagSame as above; you need curl -I to see it
Tag attributes on key parts of the bodydata-nosnippet; most often found on price tables, FAQ answers and fact blocks
Figure: the three nosnippet controls hide in three places; looking only at the HTML misses the response-header one

D1's receipt: the results of the three greps (all empty, or which one hit and on which page), one screenshot per page for all 8. This switch gets turned back on after a redesign, a template change or a plugin update, so it is also a standing item in the monthly door-layer check (see 2.9).

⚠️ Many people treat Google-Extended as the AI Overviews switch and spend effort arguing over whether to allow it, while missing this real switch. For why it isn't Google-Extended, see the four misconceptions in 2.4.

robots and WAF: record six lines

  1. 1Retrieval UAsRecord each of the 12, allowed or blocked
  2. 2Wildcard group NN Disallow lines, copy them out line by line, word for word
  3. 3Training UAsAll 6, mark every one "unrelated to whether you get cited"
  4. 4Snippet switchescurl -I the homepage and the price page, check all three places
  5. 5CDN/WAFBlock AI bots, 429s, allowlist
  6. 6Access logsWhether you have them, the date you requested export access
Figure: six lines for the read-only audit — miss even one and it doesn't count as checked; lines 2, 4 and 6 each sit on a failure path that can waste the whole quarter

The three failure paths are: skip line 2, and merging robots later strips away the protection that was already there (2.6); miss line 4, and nosnippet stays on with nobody noticing, so the Google leg counts as zero for the quarter; skip asking about line 6, and gate 0 can't start work the same day, leaving the whole door layer as nothing but inference — which is why this line should be asked as early as possible.

Line 5 needs three things nailed down: ① whether the "Block AI bots" / "Bot Fight Mode" switch is on, off, or doesn't exist; ② whether rate rules hit crawlers with 429s (check against the 429 share in the gate-0 table); ③ whether the IP range for openai.com/chatgpt-user.json is on the allowlist.

Where to check each line: for lines 1 and 3, open https://<domain>/robots.txt and read it group by group; for line 5, check the CDN / WAF dashboard; for line 6, ask ops and record "has Logpush" / "has access.log" / "kept no logs". The six lines take roughly 10, 5, 5, 10, 10 and 5 minutes in turn. For the list and purpose of the 12 retrieval UAs and 6 training UAs, see 2.4; for a blank record sheet, see → General Edition B.1 Door templates.

⚠️ Allowed in robots does not mean actually let through — the CDN / WAF layer can still return a straight 403 regardless of robots (inferred, 85% confidence). "Allowed" from the audit is only the paper-level first layer; the only evidence of actually being let through is gate 0's Crawler hit table.

2.3 Gate 2: four identities, fetched live (read-only)

What you'll do in this section: stop judging the door by looking at content through curl -A, and instead look at the same URL through four paths, then read the differences between the four to tell CSR, WAF and a closed door apart. When you're done, you'll have four separate receipts — never merge them into one.

False negative: you conclude "blocked"False positive: you conclude "open"
The WAF allows verified crawlers through by IP + rDNSThe site serves a degraded SSR version to non-browser UAs
You fake the UA from your office and get blocked as a spoofed crawlerThe real retrieval path actually gets the CSR version
So you go fix something that wasn't brokenSo you spend the whole quarter writing pages AI can't read
Figure: use curl -A to impersonate a crawler and read the content back, and you can get the verdict wrong in either direction

curl -A "OAI-SearchBot" only proves how your infrastructure reacts to that string. The kind of pass-through on the left isn't hypothetical: Cloudflare Verified Bots really does allow by IP + rDNS.

Four identities: one URL as four paths see it

Ranked by how trustworthy each one is; all of them are free.

IdentityWhere you get itPass criterionWhat it can prove
a · What the real crawler seesGate 0's access logs, CIDR-filtered against the official IP ranges such as openai.com/searchbot.jsonThis URL was hit by OAI-SearchBot within the last 30 days and returned 200The only direct evidence, applies across every leg
b · The HTML Bing fetchedBing Webmaster Tools → URL Inspection → the source returned by Live TestYou can string-search the source and find the price figure and the body textThe Bing / Copilot leg; also the free sample closest to "what a non-rendering crawler sees"
c · The HTML Google fetchedGoogle Search Console (GSC) → URL Inspection → the HTML under "Crawled page"Same as above, string-searchableOnly proves Google's side (Google renders), never use it alone as your conclusion
d · What ChatGPT's live-fetch path seesPaste this URL into ChatGPT, verbatim: "Open this link and copy out, word for word, the line on the page that gives the price."It can quote it back word for word (not paraphrase it, not estimate it)The ChatGPT-User live-fetch leg is open

curl is not a fifth identity — it's a troubleshooting tool, and it has to be split into two separate reads: ① fetch the HTML with an ordinary browser UA and check whether the body text and price are there — this is the right way to judge CSR, and CSR has nothing to do with which UA you use; ② run it again with the OAI-SearchBot UA, look only at the status code, not the content, to judge the WAF. Screenshot each of the two separately, and write on the receipt "for troubleshooting only, not a pass criterion". The half-day quick-start version (→ General Edition 4.8 Two shortcuts and this chapter's checklist) also judges JS following step ①.

How to read the differences between the four

hit and 200

all 0

403/429 over 5%

no logs

yes

no

both open

b open, d not

any other combination

How does a look in the logs

Pass

Door not open, or not discovered

WAF or a rate rule is blocking

Is c open and b not?

CSR dependency

Is b open? Is d open?

Pass, degraded

No verdict yet, fill in a

Figure: the four-identity decision tree — the conclusion comes from the differences between the four, not from ticking each one off

Outside the figure, also note:

  • "Pass, degraded" must have "no direct evidence, fill in a after 30 days" noted on the first line.
  • CSR dependency = Google's rendering can get it, but a non-rendering crawler cannot. Log it as a to-fix item and open an SSR / prerendering task the same day; push all page work back as a whole, and add no new content pages this month.
  • Never edit content for the WAF cell — that belongs to the WAF allowlist in 2.6.
  • For the door-not-open cell, never write "demand doesn't exist", and never skip ahead to "change topic".

D1 is the audit. On day 7 after fixing the door (2.6), pull all four receipts again and lay them side by side against the D1 set.

2.4 What each crawler is for, and judging the door leg by leg (read-only)

What you'll do in this section: tell the 12 retrieval crawlers and the 6 training crawlers apart, and know which leg breaks when you block which one; when you hit a platform, multiple subdomains, a login wall or a multilingual site, judge the door leg by leg. When you're done, you'll have a Per-leg door table: one row per host, one column per leg — never merge into one cell.

This is the single easiest thing in the whole chapter to get wrong: block the wrong one and nothing happens; allow the wrong one and the whole quarter is wasted.

Retrieval legs (12)Training legs (6)
Decide whether you can appear in AI answers at allOnly about content licensing, unrelated to whether you get cited
Block OAI-SearchBot and the ChatGPT leg is zero for the quarterBlock GPTBot and you can still be cited
Block Googlebot and you lose AI Overviews and organic rankings togetherGoogle-Extended has nothing to do with AI Overviews
Allow or not: allowAllow or not: your own call
Figure: block a retrieval leg and that leg breaks; Disallow every training leg and not one retrieval leg is affected

12 retrieval legs: what blocking each one costs

CrawlerBelongs toWhat blocking it costs
OAI-SearchBotOpenAIThe ChatGPT leg is zero for the quarter
ChatGPT-UserOpenAIUsers who open your page from inside ChatGPT can't read it. ⚠️ The real switch is in the WAF, not robots (2.6)
OAI-AdsBotOpenAIOnly affects ad scenarios; if you don't run ads, you don't have to allow it
BingbotMicrosoftCopilot breaks outright; one of ChatGPT's eight retrieval sources goes missing
GooglebotGoogleYou lose AI Overviews / AI Mode and organic rankings together: all three share the same crawler
ApplebotAppleThe whole Siri / Spotlight / Apple Maps / Apple Intelligence leg breaks. ⚠️ Until this group is allowed, don't claim your Apple Business Connect profile
AmazonbotAmazonAlexa / Amazon-side retrieval breaks
PerplexityBotPerplexityThe Perplexity leg breaks
ClaudeBotAnthropicClaude-side access breaks
Claude-SearchBotAnthropicClaude retrieval breaks
Claude-UserAnthropicUsers who open your page from inside Claude can't read it
DuckAssistBotDuckDuckGoDuckAssist breaks

The 6 training legs: GPTBot · Google-Extended · Applebot-Extended · Meta-ExternalAgent · CCBot · Bytespider. Whether you allow them is a content-licensing question, not a visibility question.

MisconceptionFact
Blocking GPTBot protects your content, or gets you shut out of AIGPTBot is a training leg; 88.2% of sites that block it still get cited
Allowing Google-Extended turns on AI OverviewsAI Overviews and AI Mode go by Googlebot + nosnippet
Block Applebot-Extended and the Apple leg breaksIt's only a training opt-out switch; the Apple leg goes by Applebot
No AI crawler respects robots, so editing it is pointlessAll three Anthropic bots respect it; it's ChatGPT-User that doesn't
Figure: the four most common misconceptions — all of them mix up training legs with retrieval legs, or mix one company's rules with another's

The real danger is the flip side of the first one: blocking OAI-SearchBot and Bingbot themselves. The second is the most expensive to get wrong, and it rests on two of Google's own sentences. The crawler documentation: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." AI features and your website nails down the entry requirement: "a page must be indexed and eligible to be shown in Google Search with a snippet", and "robots.txt directives for Googlebot is the control". So the gate for AI Overviews is Googlebot plus the three nosnippet controls. The point of the fourth is that it saves you work: for the three Anthropic bots (ClaudeBot / Claude-SearchBot / Claude-User), editing robots genuinely takes effect — you don't need to touch the WAF. 88.2% comes from BuzzStream; for the rest of the sources, see → General Edition A.5 List of sources.

AI Overviews and AI Mode

citations inside the Gemini app

the Apple leg

ChatGPT retrieval

ChatGPT live fetch

Claude retrieval

Claude live fetch

Perplexity retrieval

Perplexity live fetch

Copilot

Which leg are you judging

Googlebot + nosnippet

Google-Extended

Applebot

OAI-SearchBot

ChatGPT-User

Claude-SearchBot

Claude-User

PerplexityBot

Perplexity-User

Bingbot

Figure: judging a leg means checking the right UA; check the wrong one and you get one of two opposite wrong verdicts — "allowed" or "can't get in"

The zero-JS-execution corollary: mainstream AI crawlers do not execute JS; the only exceptions are Gemini (which runs on Google's infrastructure and renders in full) and Applebot (evidence in → General Edition 1.1 How buyers ask, and what AI reads). So the price, the core facts and the body copy must be in HTML that's already been rendered server-side — fetch the source with a browser UA and you should be able to string-search it; write them in JS instead and the next 90 days are wasted. The flip side of the same corollary: live fetching doesn't read JSON-LD, so schema is a baseline, not a lever (2.7). This exception also gives the whole book its single most useful troubleshooting criterion: the Google leg has seats while the ChatGPT leg is zero — that difference by itself is the fingerprint of CSR.

Judging the door leg by leg

can't fetch

fetched

yes

no

blocked

not blocked

200 and body text found

blocked

not done

Could you fetch robots.txt

Log not retrieved, retry from another network

Does this UA have a named group

Judge by the named group

Falls to wildcard group; not a block in itself

Blocked?

Log this leg as blocked

Raw-fetch a page with this UA

Log reachable

Log fetch blocked, with status code

Write only: robots unblocked, fetch unverified

Figure: judging one leg's door — not blocked in robots does not mean it can actually fetch; only a raw fetch returning 200 with the body text found counts as "reachable"

Recording rules outside the figure — skip none of them:

  1. curl the whole of robots.txt yourself, download it, parse it group by group, and log it by date. Grepping for one keyword alone will miss the single most important thing: that this UA was never given its own named group at all.
  2. Check robots on the host that the page actually lives on, not the main domain. Detail pages, doc sites and help centres are often on a different host or subdomain, and robots is not shared with the main domain.
  3. One row per host, one column per leg — never merge into one cell. The very same platform can easily have some legs blocked and others open; merge it into one cell and the whole chain gets thrown away.
  4. Google-Extended, Applebot-Extended and Meta-ExternalAgent are training / licensing legs — they don't go on this table.
  5. The table must carry a "verification level" column: robots-verified / fetch-verified. A raw-fetch receipt may only say "fetched 200 + body text found" or "fetch blocked (<status>)", never "allowed". robots only lifts a protocol-level ban; CSR rendering and WAF / geo-blocking (403s, captchas, redirects) are two other, separate gates, and they are not written in robots.
  6. A platform's robots can change, so log it in the Off-site asset decay table and rerun it once a quarter. A refused connection, a timeout or a failed TLS handshake all get logged as "not retrieved"; if a retry still fails, leave it blank — never guess and fill it in.
  7. A platform closing its door ≠ that platform has no sales value. All that's ruled dead is "citations through that one leg", not the business; other legs, and the shopping shelf that runs through data feeds, may still show up as normal. Never mix these two things up in anything you tell clients.

On a host where a given leg is blocked, schedule no content work on that leg, and don't let acceptance count it towards seats. What's being judged here is only whether the door is open, which is a different thing from → General Edition 4.6 Four states, two denominators and the abstain line judging whether "a given URL can be reached at all". For reproducible, per-platform robots commands, see → General Edition B.1 Door templates.

Example (training) Measured on 2026-09-21, the robots.txt of TPGateway, Singapore's official training portal: User-agent: * → Disallow: /; the Allow: / groups that follow list only Googlebot (including Image / Video / mobile), Slurp, bingbot, Applebot, Twitterbot and Pingdom. Not one of GPTBot / OAI-SearchBot / ChatGPT-User / ClaudeBot / PerplexityBot appears — all of them fall under Disallow: /. → The Google, Apple and Bing legs are open; the OpenAI, Anthropic and Perplexity legs are blocked: build for the first three legs only, accept work against the first three legs only, and don't count ChatGPT seats.

Five kinds of site, one extra check each

Kind of siteExtra checkCriterion and how to handle it
Single-page app / client-side renderingFetch the HTML with an ordinary browser UA, grep for one actual sentence of body textNot found = the body text isn't in the HTML. Move key body text to server-side rendering or prerendering — no "wait and bundle it with the next SEO redesign"
Multiple subdomainscurl a robots.txt for each subdomain separately, one cell per subdomain × retrieval legFor any subdomain with no receipt, its pages don't count as "door open" this quarter and can't be counted towards deliverables. Doc sites, help centres and trust pages need their door opened and to be server-side rendered too
Login wall / quote wallWhether the key figures sit outside the wallAnything behind the wall is invisible to AI: build a public, checkable summary page or a mirror on the main domain
Multilingual sitesRun the four-identity check once for each languageNever check only one language
Price or stock that changes very fastWhether these fields are injected by a client-side scriptYes = as good as not there

Judge the two causes separately, and don't reverse the order: check the robots receipts first (one curl per subdomain × leg — the cheapest check). When one subdomain was missed, the blocked leg reads zero and the unblocked legs read normal — a one-line fix; only check CSR once robots is all clear. Reverse the order and you'll schedule a whole quarter of engineering rework for a problem that a single line in robots would have fixed. Only once both are fixed does that leg go from 0 to 1.

Example (B2B software) A site naturally has multiple subdomains: www / docs / help / blog / status / app / trust. Check robots for each subdomain separately — a clean main domain doesn't mean docs is clean too.

2.5 Gate 3, read-only: connect the two free AI reports

What you'll do in this section: on D1, connect GSC's generative AI report and Bing Webmaster Tools' AI Performance, actually read each one once, and check that the site field they return really is your own site. When you're done, you'll have two sets of data that are the only free, first-party, unfakeable data in the whole engagement; only open the probe budget after you've read them.

  1. 1Request accessDomain verification, in the same email as the log access request
  2. 2Open the GSC reportGenerative AI performance, export the CSV and screenshot it
  3. 3Open the Bing reportWebmaster Tools AI Performance, same as above
  4. 4Actually read it onceCheck the site field — get it wrong and you'll silently read another site
  5. 5Save itTwo connection screenshots; how to read them is in 4.2
Figure: five steps to connect the two dashboards; reading the wrong site is worse than not connecting at all, because the tool will silently read another site and still return ok

Three rules:

  • Get the site field wrong = this box isn't done — don't mark it complete. When the site parameter isn't configured correctly, the tool will read a different site and still return ok, which is worse than not connecting at all: you'll be treating someone else's numbers as your own baseline. The same goes for other read-only dashboards such as GA4.
  • Can't get domain-verification access → list a red blocker. It blocks more than just this box: Bing's indexing submission (2.7), and Bing Live Test and GSC's crawled-page view (identities b and c in 2.3) all need domain verification first.
  • This section is connect-and-read only. How to read each of the two reports, what each one gives you, and what to watch out for are written up once, in → General Edition 4.2 The question pool: where the 30 questions come from, how they are balanced, how they are signed; checking Bing for indexing gaps and submitting via IndexNow are changes, and belong in 2.7.

Another free thing to start the same day, in passing, is copying out the actual words buyers used, from your front-desk records or CRM, for the last 15 deals before they closed — it's the question pool's only class A source, and how to do it is also in 4.2.

2.6 Fixing the door: nosnippet, the four-step robots merge, the WAF allowlist

What you'll do in this section: first confirm the baseline is saved, then on the same day make all the fixes in the order nosnippet → robots → WAF, and log the day you fix the door as the split day. The hard gate still applies: until the baseline is saved, do not change the door, any profile or the facts page. When you're done, you'll have a merged robots.txt, a receipt with the N × 12 line written on it, and one WAF allowlist rule.

  1. PreconditionBaseline savedBaseline Top 20, noise band, the 36 brand-six runs
  2. Gate 1anosnippetRevert all three places, grep comes up empty
  3. Gate 1brobotsFour-step merge, receipt has the N × 12 sentence, word for word
  4. Gate 1cWAFAdd the chatgpt-user.json IP range
  5. Log split dayThe day you fix the doorWritten into the page register
  6. D+7Re-pull and recheckCheck the logs for ChatGPT-User, re-pull the four identities
Figure: on the day you fix the door, go in this order — the split day counts from this day; if the precondition isn't met, take none of the steps after it

The day you fix the door is the split day for the whole engagement; every "before vs after fixing the door" comparison is drawn against it. Each page also has its own split day: the day it goes live (2.8). The book's W1 also counts from this day.

Gate 1a · nosnippet

Where it hitsHow to fix it
nosnippet or a low number in meta robots / googlebotChange it to max-snippet:-1
The response header X-Robots-TagRemove it at the CDN / server layer, or change it to max-snippet:-1
data-nosnippet on a body tagDelete this attribute
Figure: each of the three hit locations has its own fix; only when a rerun of the grep comes up empty on all three does it count as passed

Any hit left unfixed for four weeks → the Google leg counts as zero for the quarter — write it on the first line, and mark the Google cell "switch not lifted" in the monthly record.

Gate 1b · The four-step robots merge

the most specific group wins

every other group is ignored

Append a named group

That crawler now reads only its own group

The wildcard group's Disallow stops applying

Admin paths, parameter pages enter AI retrieval

Figure: simply appending a named group to the end of robots pulls these crawlers out of the wildcard group, and every protection it gave them stops working

What gets let out is usually paths like /wp-admin/, /cart/ and /?s=. The consequence isn't a stalled ranking — it's admin paths and parameter pages ending up in AI retrieval.

  1. 1Copy the wildcard groupCopy the Disallow lines word for word, record the line count N
  2. 2Write the retrieval groups12 named groups, each with Allow: /
  3. 3Copy the N linesInto every named group, until the count reaches N × 12
  4. 4Finish upList training groups separately, put the Sitemap line at the end
Figure: four steps to merge robots, never reverse the order; skip step 3 and you've lifted the protection that was already there

The Sitemap: line goes at the end because it doesn't belong to any group. Whether to allow the training crawlers is your own call, but any training group that says Allow: / needs those same N lines too. The receipt must record this sentence word for word:

Original wildcard Disallow lines = N, copied line by line into N × 12 places (12 named retrieval groups); training groups with Allow: / also copied.

If this sentence isn't on the receipt, gate 1 doesn't count as passed. The only full text of the robots master template — 12 retrieval groups plus 6 training groups — lives in → General Edition B.1 Door templates; hand it to whoever edits the site and have them copy straight from there. The <N lines of Disallow> in the master template is a placeholder and must be replaced with the original text copied down in step 1 — not one line short, and in every group.

Won't edit robots → all site-layer work is downgraded to read-only monitoring.

Perplexity-User does not go in the master template. Perplexity's own crawler documentation says: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules." (https://docs.perplexity.ai/guides/bots). Like ChatGPT-User, robots can't control it; the only way to allow it is to add its official IP range in the WAF (published on the same page as https://www.perplexity.com/perplexity-user.json).

Gate 1c · WAF: the real switch for ChatGPT-User

Write Allow: / in robotsAdd the official IP range to the WAF allowlist
Harmless, but has no effect on ChatGPT-UserThe real switch, verifiable in the logs after 7 days
OpenAI's own words: robots.txt rules may not applyName the rule allow-chatgpt-user
Fix only this side and you'll think the door is openThe live-fetch path is actually still closed, ChatGPT-User hits stay at 0
Figure: for ChatGPT-User, Allow in robots is only a declaration — the real switch is in the WAF

OpenAI's own complete sentence is: this kind of fetching is triggered by a user action, "Because these actions are initiated by a user, robots.txt rules may not apply". On the WAF side, do three things the same day, and screenshot each one:

  1. A screenshot of the "Block AI bots" / "Bot Fight Mode" switch's state.
  2. Rate rules: move verified crawlers out of the rate rules, screenshot it, and attach the 429 share from the gate-0 table.
  3. Allowlist: add the IP range for openai.com/chatgpt-user.json, name the rule allow-chatgpt-user, screenshot it.

D+7 recheck: re-pull the logs and check whether ChatGPT-User has any hits; on the same day, re-pull the four receipts from 2.3 as well and keep them side by side with the D1 set.

2.7 Gates 3 and 4: indexing paths, and JSON-LD sealed once

What you'll do in this section: handle indexing as two separate legs — use IndexNow for the Bing leg, and profile claiming plus pasting the URL for a quote-back on the ChatGPT leg; then spend half a day sealing the site's JSON-LD once, after which it never goes into a per-page check again. When you're done, you'll have two indexing receipts and one sealed schema template.

Gate 3 · Indexing in two separate cells

Bing legChatGPT leg
Check indexing page by page with site:, submit any gaps via IndexNowFoursquare claiming: pick categories word for word from the official category tree
Check IndexNow Insights in Bing Webmaster ToolsGet the Foursquare business status right; claim Yelp too
A submission-to-indexing report exists and can serve as a receiptPaste the URL and have ChatGPT quote the price line back word for word
Copilot uses this leg directlyLocal questions mainly go through Foursquare, Yelp and OpenAI's local index
IndexNow only feeds Bing and YandexThe criterion is identity d from the four identities; Bing indexing is not a criterion here
Figure: indexing is two separate cells, never merge them — being indexed on Bing doesn't prove the ChatGPT leg has read you

"ChatGPT runs through Bing's index" no longer holds as of 2026-09: ChatGPT has multiple retrieval sources, and Bing is only one of them (evidence in → General Edition A.1 Evidence for mechanics, the door and identity). So still do IndexNow, but it isn't the switch for the ChatGPT leg. Claiming a profile is itself a change, and is only done once the baseline is saved; for the full method for the five business profiles, see → General Edition 3.4 Third-party credential tiers, Wikidata and the five business profiles. Having Applebot allowed, and verified, is a precondition for claiming your Apple Business Connect profile.

Gate 4 · JSON-LD sealed in half a day

write first

live fetching can read it

write second, in passing

indirect, Google leg only

live fetching

The same set of fields

Visible HTML

Enters AI answers

schema and sameAs

Google Knowledge Graph

None of them get read

Figure: text only gets read by live fetching once it's in visible HTML; schema only takes the indirect route through the Google Knowledge Graph

Reverse the order and you end up with a site whose fields are all complete but that AI still can't read. Even schema's small indirect effect only exists on the condition that the same fields are already written into visible HTML; do it on its own and the return is close to zero. After adding JSON-LD, AI Overviews (AIO) citations moved −4.6%, the only significant result, and it's negative; the other two platforms weren't significant (evidence in → General Edition A.1 Evidence for mechanics, the door and identity), so this is a one-off task:

RuleWhat it means
One-off templateOrganization (or your industry's schema subtype) + a Person for every practitioner + sameAs. Set it up once when the site is built, sealed in half a day
Not written per page, not checked per pageDoesn't go on the per-page launch checklist, and doesn't go on the review checklist
Only four things need to be right① the legal name matches /facts word for word ② the address and phone number match the five business profiles word for word ③ each practitioner's registration number ④ the sameAs string
A confidence noteWrite one line in the template file — "indirect, Google leg only, low confidence" — so the next person doesn't mistake it for a main lever

The sameAs string links these: regulator register or registration-number lookup pages, association profile pages, Google and Bing business pages, ORCID / Scholar, conference speaker pages, LinkedIn, the Wikidata QID.

Two rules in passing: ❌ don't stack FAQPage schema (it's built for rich results and doesn't feed AI citations); ❌ don't add a self-reported aggregateRating (on the strictly regulated side this also runs into the ban on testimonials and ratings: Statute text in the example industry). Do build question-style H2s and a short Q&A at the foot of the page; don't stack schema.

The one exception: when structured data is used as a data channel (for example, a product feed feeding a shopping graph), it is useful — but that's a different route, one that doesn't go through live fetching. "Structured data is useful" is only true in this sense, and it is no reason to change anything on your landing pages. For the approved wording to give a client on this point, see → General Edition B.2 Approved wording and scripts.

2.8 The launch gate for every page

What you'll do in this section: on the day each page goes live, run it through three checks, then do a quote-back test on day 7; only once every check passes do you mark it "live", and log that page's split day. When you're done, you'll have a timeline for every page from launch to D90, and a page register logging the launch date.

  1. Same dayNew title visibleIncognito window + hard refresh, confirm with your own eyes
  2. Same dayThe four copies matchThe same body text and price are present in all four HTML copies
  3. Same daySubmission receiptIndexNow ping, covers the Bing leg only
  4. D7Quote-back testChatGPT quotes back one exclusive fact word for word
  5. Mark "live"Log the split dayThe split day is the launch day, not the day work started
  6. D14Long-form videoThe video gate; miss it and the next page can't start
  7. D30/60/90Compared like for likeSame question pool, same engine, same number of passes
Figure: a page from launch to D90; miss any one of the first four items and you may not mark it "live" — this page's split day is logged as the launch day, not the day work started

What happens if each step is missed:

  • Miss new title visible: whoever opens the page sees the old version.
  • The four HTML copies match: fetch the source once with each of four UAs — OAI-SearchBot, Bingbot, Googlebot, an ordinary browser — and the same body text and the same price string must be present in all four. Missing this has two possible outcomes: showing machines a different version gets judged as manipulation, or the price never makes it into the answer at all. What this step checks is "are the four copies the same version", which is a different thing from the four identities used to judge the door in 2.3.
  • The D7 quote-back test is the real acceptance check for the ChatGPT leg: ask ChatGPT once, using the exact wording, about one exclusive checkable fact on that page, or paste the URL and have it quote that line of figures back word for word; screenshot it for the record. n = 1 is valid here too — it tests "was it read or not", not "what position it ranks".
  • D14 long-form video doesn't hold up "live" — it holds up starting the next page. It belongs to writing pages' three mechanical gates; specs in → General Edition 5.4 The 30-day writing order and the three mechanical gates and → General Edition 6.7 Long videos narrated by the named expert.
  • D30 / D60 / D90 must be measured to the same spec at all three points: same question pool, same engine, the same number of passes. Seat counts scale linearly with passes — baseline at 5 rounds, D90 run at 10 rounds, and the count doubles even if you did nothing at all; for why, see → General Edition 7.5 Ruler discipline: change the ruler and nothing is comparable.

When each page goes live, also log 2–3 checkable fact strings that are unique to it, and look for them again during each monthly retest — this is the only signal in the whole process that can give you page-level cause and effect; for how, see → General Edition 7.2 The noise band, the page-level signal and the ten monthly steps.

2.9 Door-layer troubleshooting and this chapter's acceptance checks

What you'll do in this section: whenever a monthly retest shows a seat increase that doesn't exceed the noise band, start checking from the door layer, and use the two diagnostic figures to map every symptom to one fix; check it even when seats went up. Only once every one of the ten items on the checklist is ticked does the door count as done.

When to check: two points in time. ① Whenever the seat increase doesn't exceed the noise band, start checking here — never skip a layer (for the rest of the triage layers, see → General Edition 7.3 Not-moved triage and next month's three points). ② Check the door layer once a month, unconditionally, even when seats went up: its failures are silent — while a page is not being read at all, seats can perfectly well be rising for some other reason, and you won't discover a whole wasted month until the next one.

fix

fix

fix

fix

fix

verdict

Door-layer symptom

OAI-SearchBot hits = 0

Gate 1's three small steps, 2.6

403 + 429 over 5%

Allowlist it and move it out of the rate rules

All 200 but only the homepage hit

Sitemap, homepage internal links, IndexNow

Admin paths showing up in search

Redo the merge, get the count to N × 12

No movement on the Apple profile

Allow and verify Applebot first

Blaming GPTBot being blocked for not being cited

Wrong call — don't use it as a criterion

Figure: at the logs / robots / WAF layer, every symptom maps to exactly one fix

check first

verdict

fix

verdict

verdict

fix

in order

any one fails

Google has seats, ChatGPT is zero

CSR dependency

Four identities: c open, b not

Open an SSR or prerendering task

Four identities: b open, d not

WAF, go back to the edge-layer figure

Seats drop after a redesign or plugin change

nosnippet has come back

Revert all three to max-snippet:-1

Seats haven't moved, cause unclear

Check five things, see below

It's a door problem, not the questions

Figure: at the rendering / indexing / nosnippet layer — Google has seats while ChatGPT is zero: check CSR first, not the choice of questions

When the "cause is unclear", check these five things in order: ① whether the access logs show a real OAI-SearchBot hit, and what status code it returned ② whether the three nosnippet controls have been switched back on by a CMS or plugin ③ whether the four HTML copies match word for word ④ Bing indexing checked with site: ⑤ paste the URL into ChatGPT and have it quote the price line back word for word.

Rules outside the two figures:

  • Anything that fails: fix it the same day, recheck after 7 days, and don't change topics or add new content pages this month.
  • Never edit content for the 403 / 429 cell.
  • Never skip past the door layer and go straight to writing "demand doesn't exist". Stuck on the same layer, unfixed, two months running is an execution problem, not a problem with the questions.
  • All five gates are mandatory in every industry. Two more get added depending on the kind of site: a CSR check (for content rendered by front-end components such as price lists, course schedules, doc sites, filters and pricing widgets — fetch the HTML with a browser UA and check whether the body text is there); a visible-text check (raw-fetch with a retrieval-leg UA, strip the tags, and count the visible body characters and the number of stand-alone statement sentences — a page with numbers but no sentences can't get onto the answer shelf at all).

Example (e-commerce) The visible-text check is mandatory: a page with fewer than 12 statement sentences may not be marked "live". For a measured case where a category page's HTML was large but, once the tags were stripped, only a small scrap of visible body text remained, see A.1.

Acceptance checklist for this chapter

The door is a multiplier, not a variable: it only has two values, 0 and 1, and with no evidence it counts as 0. Tick these off, and only once every box is ticked does the door count as open:

  • [ ] All four columns of the Crawler hit table filled in, saved as crawler_hits_30d.csv; where logs couldn't be obtained, a red blocker is listed with the catch-up date noted
  • [ ] All three nosnippet greps come up empty, one screenshot for each of the 8 pages
  • [ ] robots.txt has all 12 retrieval-class named groups, and the receipt has this written on it: "Original wildcard Disallow lines = N, copied line by line into N × 12 places"
  • [ ] The Sitemap: line is at the end of the file
  • [ ] The WAF allowlist rule allow-chatgpt-user has been added, with a screenshot on file; the 7-day recheck shows ChatGPT-User hits
  • [ ] Each of gate 2's four criteria has its own separate receipt, none merged into one; the two curl reads are each screenshotted and marked "for troubleshooting only, not a pass criterion"
  • [ ] Both dashboards — Bing Webmaster Tools AI Performance and GSC's generative AI report — are connected, each with a screenshot
  • [ ] The JSON-LD template is sealed in half a day, and is never written per page after that
  • [ ] Every page that has gone live has passed the launch gate (three checks the same day + the quote-back test on day 7), with the launch date logged
  • [ ] The door-layer troubleshooting checklist has been added to the standing monthly actions

Every part of the checklist that's about "recording the current state" is done on D1; every item that actually changes the door is only done once the baseline is saved, and the day you fix the door is already logged as the split day.

Back to contents · GEO Playbook: General Edition

This chapter is published under a CC BY 4.0 licence · © Canlah AI. To republish or adapt it, credit “Canlah AI · GEO Playbook” and link to this page.

A condensed version for AI assistants is on GitHub, and the Markdown version of this chapter can go straight to an AI assistant. The quick guide and full-book downloads are in the downloads section. The measurements behind the numbers in this book are on the dataset page (CC BY 4.0).