# Chapter 7 · Monthly retest: how to tell whether it's working and what to do next month

> GEO Playbook · General Edition v1.0 · Canlah AI · CC BY 4.0 · Web page: https://canlah.ai/playbook/measure/
> Markdown edition for AI assistants, same content as the web page. Figures are code blocks (wireframe / mermaid / bars / steps / split); "→" links open the matching section on the web, and the same URL with .md is its Markdown edition.

Ask again every month with the ruler you froze for good. The verdict is never "it worked" or "it didn't" — write **which three points to hit next month**. Alongside it, keep three firm numbers that don't depend on AI's own randomness, to prove that this month's work really landed.

## 7.1 What a retest produces, and the rulers

**What you'll do in this section**: turn the monthly retest's verdict from "did it work" into "which three points to hit next month". When you're done, you'll have five rulers, each answering one question, the limits each one must print exactly as they are, and four sentences you'll never say again.

The order is fixed: **read the free first-party data first, then look at the probes that cost money.**

### Why we do not judge with statistical tests

```split Figure: On a small sample, statistical judgement can only output "not demonstrated"; raw counts + a noise band are what actually produce next month's work orders
Statistical judgement (internal reference only) || Raw counts + noise band (used for the verdict)
Wilson three-state: moved / seems to have moved / not moved || Seat count, pages on the list, factual errors (count)
The threshold for "moved" is out of reach on a small sample || Retesting 3 times in the same week is enough to measure the noise band
Default output: "not demonstrated" || Output: next month's three work orders
Only used to trigger not-moved triage || Used for the verdict, and to schedule the work orders
Permutation test: at n ≤ 10 it changes no decision || Don't run it — that's just ritual
```

Run it on a statistical basis, and week 13's default output is necessarily "not demonstrated" — **you walk in already carrying the conclusion "this didn't work"**, whether or not anything actually moved (for how the threshold is worked out, see the evidence in （→ 通用版 A.2 选点、写页、站外与复测的证据）). The fix isn't to loosen the statistics — it's **to change what you accept as proof**. Rankings are moved by work orders, not by p-values.

If you're doing this yourself, **the minimum viable ruler set = three numbers + one page-level signal + one web control leg**. You don't need a Wilson interval, a control page, or a permutation test — but you do need to honestly separate "actually moved" from "noise". The three numbers are the seat count, on the list and cited, and factual errors (count) below; the page-level signal and the noise band are in 7.2; the full specification of the web control leg appears only in （→ 通用版 4.4 网页对照腿与冻结基线（全书唯一完整规格））.

### Name the legs correctly first

The two API legs' official names are **① OpenAI API leg** and **② Gemini model leg**. The diagram mapping the three products to their three rulers is in （→ 通用版 1.2 两条腿：ChatGPT 找源头，AI Mode 找二手）; here we only add the three mistakes people most often make at retest:

- Treating the Gemini model leg's number as "visibility in Google's AI Overviews". **Gemini API ≠ AI Overviews ≠ AI Mode** — the fetch path, the way the answer is generated, and the source pool all differ; passing one leg's number off as another's is measuring A to sign off on B. Visibility for AI Overviews / AI Mode **uses only the impressions from the Search Console generative AI report**.
- Checking `Google-Extended` to judge whether AI Overviews can crawl you. **You should check `Googlebot`.** `Google-Extended` governs training and the Gemini app side; blocking it does not affect whether AI Overviews indexes you — it's `Googlebot` being blocked that means the door is shut.
- Describing the API legs' results as "what users will see in ChatGPT". Where the only evidence is from an API leg, the wording must always be qualified as "on the model API".

### Five rulers, one question each

```mermaid Figure: Five rulers, one question each — print them side by side only, each with its limits attached; add them into one score and any one going bad gets covered up by another
flowchart LR
  m1["On the list X/20 (two lines)"] -->|are you on the page| out["Side by side, each with its limits"]:::hl
  m2["Seat count = times named"] -->|how many times named| out
  m3["Factual error count"] -->|are the facts right| out
  m4["Same-week retest noise band"] -->|the verdict threshold| out
  m5["GSC AI Overviews impressions"] -->|impressions on Google's side| out
  m1 -.->|add up| bad["Combined into one visibility score"]:::warn
  m2 -.->|add up| bad
```

How each ruler is counted, and the limits that must be printed exactly as they are on the report:

| Ruler | How you count it | The limits that must be printed exactly as they are |
|---|---|---|
| **Pages on the list** (report as two lines: **on the list X/20** and **on the list and cited this month Y/20**) | Of the frozen baseline Top 20 sources, how many pages now actually feature you (open each one by hand + screenshot). **The denominator, 20, does not change all quarter** | **On the list ≠ cited**: AI may cite this page without mentioning your line — being on the list is a necessary, not a sufficient, condition. The reading is fixed: on-the-list rises but on-the-list-and-cited doesn't, **for two months running** → you're investing in a container that isn't being cited; next month, invest only in sources that were actually cited this month |
| **Seat count** | How many times your name was named across this month's N question-and-answer runs; also log **this month's total names** (the sum of every business's name count) and each competitor's seat count | It's a sampled count under your own question pool, engine and number of passes — it does not represent the whole market; total names shifts with how long the answer is, so **you must print total names and seat share together**; it **scales linearly with the number of passes**, so if the number of passes changed, the two numbers must not be shown side by side or subtracted from each other |
| **Factual errors (count)** | A person reads the brand six (6 questions × 3 rounds × 2 engines = 36 runs / month, frozen for the quarter, defined in （→ 通用版 3.5 品牌六问与发现错误之后）), copies down every fact AI got wrong, and attaches a screenshot of the original | The hardest cell in the whole ruler set; it measures only whether the facts are right, not ranking. **If you can only watch one number, watch this one**: the error is copied by AI from the source page it retrieved, so fixing the source moves both legs together — of all the rulers, it is the least sensitive to skew between the API and web paths |
| **Same-week retest noise band** | See 7.2 | The noise band itself is also a sampled value |
| **AI Overviews / AI Mode impressions** | GSC generative AI report impressions, exported by page and by country (free, first-party) | Only impressions — no names, no position; it's already included in total impressions, so **it must not be added to total impressions**; flag it when a blank value is exported as 0 |

**Seats are simply times named, and the measure has only one name.** Do not write "the raw seat count" and "the answer-level named count" as two separate lines — they are the same number, and the only effect of splitting them into two is that the noise figure you're banned from reporting on its own sneaks back into the acceptance test under a different name.

### How you split changes by industry; the definitions do not

The definition of a ruler doesn't change by a single word; the only thing you can change is how you split it:

- **Seats can be split into two**: when a category has both an answer shelf and a shopping shelf, count brand seats (answer shelf) separately from SKU card placement and top-spot share (shopping shelf), each with its own noise band. **This answer's total names is fixed by measuring the category itself** — never reuse another category's number.
- **The denominator can be split by work unit**: when two fronts' named sets barely overlap (by product line, or by school stage × subject), report the denominators separately; merge them into one denominator and the ups and downs of two independent fronts will cancel each other out.
- **Factual errors can be split into two columns**: log "errors in your own source" separately from "third-party errors" — the two have completely different repair paths and time-to-effect (7.3, layer 2).
- **How many rounds go into the noise band is an industry variable** — measure it yourself in the first quarter. The fourth reference number (things like enquiry source) is chosen per industry; **it's the only cell that varies by industry, and it never enters acceptance**.

> **Example** (e-commerce) Seats split into two: brand seats (answer shelf) + SKU card placement and top-spot share (shopping shelf); the noise band is also measured as two, one for brand seats and one for SKU card placement.

The source skew between the web version and the API is systematic, not noise, and adding more rounds doesn't fix it — the number appears exactly once, in （→ 通用版 4.4 网页对照腿与冻结基线（全书唯一完整规格））.

### Four things never to say

```split Figure: Four things never to say, each paired with a reproducible way to say it instead
Never say || Say instead
"We're on the list, so AI will recommend us" || On the list X/20, of which cited this month Y/20
"Seat share 1.29% = market share 1.29%" || Asked K times, named X times, noise band ±N
"Up 2 seats, heading the right way" (within the band) || Within normal variation, cannot be determined
Adding two rulers into a combined "visibility score" || Each ruler on its own line, side by side, with its limitations
```

Writing a report, briefing your boss or a client, or even summing up in your own head — the four sentences on the left must never appear. Why: on the list ≠ cited; seat share is a sampled count under your own question pool, not the market; a rise that hasn't cleared the noise band is treating noise as a result; and two rulers with different denominators and different failure modes, once added together, let either one go bad while the other one covers for it. The noise-band statement that must be printed word for word in the monthly report is in （→ 通用版 B.2 口径句与话术）.

Retesting also has three more criteria; the full statement is in （→ 通用版 1.4 不要做的三类事） — here they're just applied:

- An observation that is single-run, single-engine and n ≤ 3 can only serve as directional evidence — **you cannot declare "zero visibility"**;
- The wording of the question pool can produce a false zero: when bare category words dominate, suspect the ruler first (7.3, layer 4); evidence in （→ 通用版 A.2 选点、写页、站外与复测的证据）;
- **Even if pages-on-the-list has risen, still run triage** — the trigger looks only at seats (7.3).

## 7.2 The noise band, the page-level signal and the ten monthly steps

**What you'll do in this section**: get three more things every month — a measured noise-band value, a list of page-level-signal hits, and a monthly schedule with people and hours assigned. With the noise band you know which rises and falls you're allowed to write into your conclusion; with the page-level signal you know which page was actually read and taken.

### The noise band: the cell people most often get wrong

**What to do**: run the same batch of questions, same engine, same number of passes, again **3 times within the same week**, giving X′ / X″ / X‴. **How long**: about 1.5 hours/month (the machine runs it; a person only checks it lands on disk). **How to check**: three independently saved answer packs, all timestamped within the same week; the noise-band value is written on page 1 of the monthly report.

```mermaid Figure: A rise in seats doesn't mean you can write "increased" — it has to pass four checks in a row, and each one you fail has a fixed way to write it
flowchart LR
  g0{"Month 1 or 2?"} -->|yes| w0["List side by side only, no verdict"]
  g0 -->|no| s{"Each leg's sample ≥ 80% of plan"}
  s -->|no| inc["Mark 'instrument incomplete', no comparing"]:::warn
  s -->|yes| q1{"Increase exceeds the noise band?"}
  q1 -->|no| nd["Within normal variation, cannot be determined"]
  q1 -->|yes| q2{"Same direction two months running?"}
  q2 -->|no| pend["Above the band this month, to be confirmed"]
  q2 -->|yes| up["Write 'increased'"]:::hl
```

The wording at each check in the figure is fixed: for months 1–2, mark "noise band baseline not yet stable (K/3 months recorded)"; in a month where the sample falls short, print each leg's figures separately, **do not print a total, and do not compare with last month**. What the figure can't show is how the noise band itself is defined:

| Rule | Why |
|---|---|
| **Run it 3 times, not once** | If you use one sampling difference as the threshold for another, then under the null hypothesis that "nothing happened" **the two have the same distribution, so the false-positive rate is about 50%**. Running it 3 times brings the false-positive rate down from 100% to 50%, not down to 0 |
| **Noise band = the range (max − min); the floor is fixed at ±1 seat** — even if all three runs come out identical, record it as ±1 | When the range is 0, any +1 would be judged "increased" — the gate fails in reverse |
| **The noise band baseline is the median of three months, not the maximum** | Taking the maximum is a **monotonically non-decreasing ratchet** — the longer you run it, the harder it becomes to declare any improvement, and you end up in a closed loop where "the worse the ruler, the more it excuses you" |
| **If the noise band widens for two months running → add more passes and retest**; you may not invoke "instrument defect, not counted this month" again | An excuse clause used more than once becomes an excuse machine. Adding more passes is changing the ruler — per 7.5, start a new baseline |
| **If one leg's valid sample that month < 80% of the planned volume → mark that month "instrument incomplete"** | When one leg is missing for the whole month, the absolute cross-engine count gets cut in half. This is the same problem as "adding passes doubles the seat count" — just the other side of it |

### The page-level signal: the only page-level evidence of cause

```mermaid Figure: The page-level signal measures "was it read or not", not "which position" — so n = 1 still counts
flowchart LR
  a["Log 2–3 unique facts at launch"] -->|retest monthly| b["String-search the saved answer text"]
  b -->|quoted back| c["This page was read and taken as a source"]:::hl
```

What you log must be a **checkable fact string unique to that page** — for example, the `S$6,800` fixed into that page's price list. Adjectives, or sentences that also appear on other pages, don't count; finding them proves nothing about which page was actually read.

It has nothing to do with seat count or the noise band, and reuses the answer-text packs you're already storing — **the marginal cost is close to zero**. Seat count tells you whether things are rising overall; the page-level signal tells you which page is doing the work — record the two numbers separately, and never infer one from the other.

### Ten monthly steps, grouped into eight stages

```steps Figure: Ten monthly steps grouped into eight stages; the first six stages are just bookkeeping — what moves rankings is stage 7, triage
1 | First-party data | Step 1: Bing AI Performance + GSC, 0.5 h
2 | Retest | Step 2: same question pool, same engine, same number of passes, 2 h
3 | Noise band | Step 3: run again 3 times in the same week, 1.5 h
4 | Count seats | Step 4: names, position, total names + page-level signal, 1.5 h
5 | Recheck on-the-list | Step 5: X/20 + cited this month or not, 1.5 h
6 | Profile + brand-six audit | Steps 6–8: profile consistency, 36 brand-six runs, conventional search audit, 2.5 h
7 | Triage | Step 9: the five layers in 7.3, 1 h
8 | Monthly report | Step 10: sent on a fixed date, 2 h; total about 12.5 h
```

Steps 1–4 are only responsible for accurately measuring what the shelf looks like right now; what actually moves rankings is step 9 — **without step 9, the first eight steps are just bookkeeping**.

**Jump-the-queue rule**: if AI produces a false statement about you or shows a negative tendency → **jump the queue and fix it first; don't wait for the monthly report** (the five situations that jump the queue are in （→ 通用版 3.5 品牌六问与发现错误之后）).

The web control leg's 40 minutes **is not in this table** — schedule it as a separate day at the start of each month; the full specification is in （→ 通用版 4.4 网页对照腿与冻结基线（全书唯一完整规格））.

Who does each step, what they hand over, and who checks it:

| Step | Who | Output | Who checks | Points the figure doesn't show |
|---|---|---|---|---|
| 1 | The report owner | `first_party_<month>.json` | Self-check | Export this month's landing-page impressions and Citation Share |
| 2 | The person running probes | `probe_<month>/`, the raw text of every answer | Self-check | Print "not measured" for any engine not measured because the quota ran out |
| 3 | The person running probes | `probe_<month>_retest/` + the noise-band value | The report owner | — |
| 4 | The person running probes | `seats_<month>.csv` + the page-level-signal hit list | The report owner | Which businesses were named in each answer, at what position, and **this answer's total names** |
| 5 | The off-site lead | `offsite_ledger_<month>.csv` | The project lead | Open every item in the baseline Top 20 by hand: are you on it, at what position, is the information out of date; while there, tag each one with the five binary page-type features |
| 6 | The profile owner | Same ledger as above | The project lead | The three review numbers are a **passive observation**; five-profile consistency |
| 7 | The person running probes | `brand-6.jsonl` + `entity_errors.md` | Read by a person, no skipping | 6 questions × 3 rounds × 2 engines = 36 runs, producing the factual-errors count |
| 8 | The person who edits the site | Audit page (with the four identities' four receipts + screenshots of the three `nosnippet` greps) | The project lead | Non-brand clicks, the canonical Google picked, the index status of key URLs, that the four identities' fetches still match (defined in （→ 通用版 2.3 闸二：实抓的四身份（只读））), whether the three `nosnippet` controls have recurred this month |
| 9 | The project lead | Next month's work orders | Look it over again the next day | Run the 7.3 triage, produce next month's three points |
| 10 | The report owner | The monthly report + the raw answer pack | Read page by page by a person | Sent on a fixed date |

## 7.3 Not-moved triage and next month's three points

**What you'll do in this section**: for a month when seats don't clear the noise band, come away with a triage verdict of "stopped at which layer", plus next month's three work orders, laid out mechanically by priority order; add to that a fixed-format page 1 for the monthly report, and a competitor board for your eyes only.

### There is only one trigger

> **This month's seat increase did not clear the noise band → always start at layer 1; no skipping layers. Whether pages-on-the-list rose or not has nothing to do with this rule.**

Why you can't add "and pages-on-the-list did not increase": pages-on-the-list is **the output of your own work**, and a monthly +1 is almost routine. Fold it into the trigger as an AND condition and you've fitted the whole triage table with a **switch you can flip off whenever you like**, and no month all quarter will ever really go through triage. **No variable you control yourself may appear in the trigger condition** — any new trigger condition proposed later must be checked against this rule.

### Five layers, cheapest first

```mermaid Figure: Not-moved triage checks layers cheapest-to-rule-out first, layer by layer, no skipping; the door and identity layers are checked every month unconditionally, even when seats rose
flowchart LR
  t["Seat increase did not clear the noise band"] --> L1{"1 · door: can it be read?"}
  L1 -->|fails| f1["fix same day, recheck in 7 days"]:::warn
  L1 -->|passes| L2{"2 · identity: got the right business?"}
  L2 -->|fails| f2["page work resets to zero, fix identity first"]:::warn
  L2 -->|passes| L3{"3 · citation slots line up?"}
  L3 -->|no| f3["handled by the three verdicts in the table below"]
  L3 -->|yes| L4{"4 · split day and ruler"}
  L4 -->|hits| f4["doesn't count as not-moved this month"]
  L4 -->|clears| L5["5 · change topic: switch battlefield"]:::hl
  m["Every month, unconditional"] -.-> L1
  m -.-> L2
```

The door and identity layers are checked unconditionally because failure at either one is **silent**: when a page was never read at all, or AI has the wrong business, seat count can perfectly well be rising for some other reason, and you won't find out until next month — by which point a whole month has been wasted. Which fix each door-layer symptom maps to is looked up in the two diagnostic figures in （→ 通用版 2.9 门层排查与本章验收）; they aren't redrawn here. Three of these are the easiest to misjudge: the Google leg has seats while the ChatGPT leg is zero — **check CSR first, not your question selection**; don't use `GPTBot` as your evidence, it's a training crawler — for the retrieval legs look at `OAI-SearchBot` and `ChatGPT-User`; and don't draw a conclusion from what `curl -A "OAI-SearchBot"` returns — both directions of that test lead to the wrong verdict.

| Layer | What to check that day | Test and action | How to say it in the monthly report |
|---|---|---|---|
| **1 · Door** | ① Access logs: in the past 30 days, has a real `OAI-SearchBot` hit these pages, and what status code came back (**the only direct evidence**) ② have the three `nosnippet` controls been switched back on by the CMS or a plugin ③ do the four identities match word for word ④ `site:` search to check Bing indexing ⑤ paste the URL into ChatGPT and have it **quote back the price line word for word** ⑥ is the Foursquare category and business status right, is the Yelp entry there | Failing any one = **a door problem, not a targeting problem**: fix it the same day, recheck in 7 days, **don't change topics or add new content pages this month** | "Last month ChatGPT simply wasn't reading those pages at all; the reason was X, it's fixed today, we'll recheck in 7 days" |
| **2 · Identity** | Read the brand six by hand: is it talking about this business? Are the name, address, registration number and price right? Any namesake confusion? Any complaint or negative tendency? | Wrong business, or facts wrong → **all page work resets to zero and gets recalculated**; fix identity first: tick off and screenshot each of the six controllable sources (own-site `/facts` → Google Business Profile → Bing Places → Foursquare → Yelp → Apple Business Connect) + send a correction letter to the cited source. **Mark every error as either "third-party page source" or "business profile source"** — mixing the two makes profile-type errors run late systematically | "AI mistook you for another business (or got your price wrong); this month we deal with that first" |
| **3 · Citation slots** | The intent-level top 10 URLs by times cited this month, checked one by one against the four columns of the citation-slot table (`four states / judged on / basis / judged by`) | **Judged reachable at the time, entry point gone now** = the shelf really changed → switch to "hold", go after that new source, put it first on this month's off-site work orders<br>**Judged "to ask" at the time, or the basis column is blank** = the judgement drifted → fill in the basis this month and recompute the abstain line; this month it doesn't count as a wrong target<br>**Fewer than 2 reachable-enough URLs in the intent-level top 10** = wrong target → move this intent out of the main-attack pool and swap in another; pages already invested in become long-tail landing pages and get no more investment | "A got into some directory last month and pushed you from 3rd to 5th; go fight for this this month" / "Last month's basis for this one wasn't kept on record; fill it in this month, then judge" / "We can't get into this slot for this question because we picked the wrong battlefield at the start; swap it" |
| **4 · Split day and ruler** | ① the page has been live less than two weeks ② the URL changed after launch ③ are the frozen questions bare words ④ has the engine / model / number of passes been touched ⑤ is the noise band bigger than last month ⑥ is the web-control-leg overlap < 50% ⑦ was one leg's valid sample < 80% of the planned volume | Any hit → recompute the split day, or mark the question pool "instrument defect", and **it doesn't count as "not moved" this month**. If ⑥ hits, this month's seat conclusion always gets the non-extrapolation sentence printed alongside it (（→ 通用版 4.4 网页对照腿与冻结基线（全书唯一完整规格））), **and it must not be used as a reason to change topics**. The "instrument defect" excuse **may not be used in two consecutive months**: the second time, always add passes and retest instead | "This batch of pages has been live less than two weeks; the first comparable point comes next month" |
| **5 · Change topic** | Layers 1–4 all pass, and it still hasn't moved | Seats are saturated, or the demand isn't real → **this is the only case that's actually "change topic"**: go back to target-picking and pick again, and it must be **a different battlefield, not just different wording** | "The shortlist for this question is locked; swap in a battlefield we can actually get into" |

Two disciplines:

- **You may not skip layers 1–4 and go straight to writing "demand doesn't exist".**
- **Stopping at the same layer, unfixed, two months running = an execution problem, not a shelf problem** — handle it ahead of the queue next month, and the same excuse may not be used again.

The definitions of the four states, the abstain line and the citation-slot table are in （→ 通用版 4.6 四态、两层分母与弃权线） and （→ 通用版 6.2 发信闸门、付费收录与夺取表）.

### Turn the triage result mechanically into next month's three points

```mermaid Figure: The triage result is mechanically converted into next month's three points, by priority order; stop once you've filled three, and if you can't, fall back by content form — no guessing
flowchart LR
  r["Triage result"] -->|layers 1–2 failed| p1["Point 1: fix the door or identity"]:::warn
  r -->|a new reachable source appeared| p2["Point 2: capture one off-site slot"]
  r -->|on the list but past position 4| p3["Point 3: improve one position"]:::hl
  p1 --> full{"three points filled?"}
  p2 --> full
  p3 --> full
  full -->|filled| stop["Stop, assign"]
  full -->|short · professional| A["Fallback A: educational thick page"]
  full -->|short · consumer| B["Fallback B: profiles and facts"]
  A -->|still short| C["Fallback C: expand the price page"]
  B -->|still short| C
  C -->|still short| D["Fallback D: Chinese page or long video"]
```

How each cell in the figure picks that one item, and what work order it assigns:

| Priority | Which one to take | Next month's work order |
|---|---|---|
| **Point 1** | Triage failed at layer 1 or layer 2 | Fix the door / fix identity. **Once triggered, it automatically takes point 1, and no new content pages are added this month** |
| **Point 2** | A source that has **newly appeared** among this month's cited sources, is judged "reachable" under the four states, is allowed on this side, and does not have you on it | Take the 1 source **actually cited this month**, with the highest citation count, where **the competitor's way in ≠ platform-edited entry**, and assign an off-site capture work order |
| **Point 3** | A source where you're **already on the list** but ranked past position 4 | Take the 1 with the highest citation count, and assign an update-letter / supply-materials work order. **Improving position is cheaper than winning a new slot** (cross-industry pooled basis, not verified in this local industry — carry the qualifying sentence exactly when you print it; sample in （→ 通用版 A.2 选点、写页、站外与复测的证据）) |
| **Fallback A** | Three points can't be filled, and the content form is **professional** | Educational-thick-page work order: fold the landing page for the billable intent with the lowest seat count into its intent cluster and thicken it. How thick follows the page-type spec — word count is not a criterion (（→ 通用版 5.7 所有页型都成立的七条）) |
| **Fallback B** | Three points can't be filled, and the content form is **consumer** | Profiles-and-facts work order: fill in the fact fields on all five business profiles, expand the `/facts` page, complete itemised prices on the price page (wording follows the side, （→ 通用版 5.2 三侧速查（一）：先判身份；价格、促销与赠送、结果数字）), and complete outbound links to official and regulatory bodies. **No review work order is assigned on the strictly regulated side**: that side may not actively ask for reviews |
| **Fallback C** | Still short | Expand a price page: price and fee queries trigger AI Overviews at a high rate; evidence in （→ 通用版 A.2 选点、写页、站外与复测的证据） |
| **Fallback D** | Still short | Complete the Chinese pages, or pair an existing thick page with a long video narrated by the named expert (（→ 通用版 6.7 本人口述长视频）) |

Content form (professional / consumer) is the second judgement, made after the side has been decided; how to judge it is in （→ 通用版 5.9 内容形态、FAQ、中文页与外语页）. The basis for splitting the fallback by form is a cross-industry query sample — **it does not include Singapore, at about 75% confidence** — the sample's basis is in （→ 通用版 A.2 选点、写页、站外与复测的证据）.

### Reddit: read the number monthly; do not fix a conclusion

Read once every month at retest: **Reddit's share of this month's cited sources** = the number of this month's cited URLs that are reddit.com ÷ the total number of this month's cited URLs (two decimal places, logged on page 2 of the monthly report).

| This month's share | Action |
|---|---|
| **> 2%** | **Schedule one Reddit work order this month** (target whichever leg it's showing up in) |
| **≤ 2%** | No investment this month — just log the number; read it again next month |
| **> 2% for two months running** | Only then may Reddit be written into the regular work rhythm |

Why we don't fix a rule of "never invest in Reddit for ChatGPT": ChatGPT's Reddit citations once fell sharply in a single month, but the same kind of collapse had happened earlier and fully recovered, and ChatGPT never stopped reading Reddit — what stopped was putting threads **into the answer**. That's "the answer surface temporarily narrowing", not "this source is dead" (numbers and dates in （→ 通用版 A.2 选点、写页、站外与复测的证据）). **A single month's swing is never written into a long-term rule.**

### Page 1 of the monthly report: three firm numbers and two measures printed beside them

```split Figure: On page 1 of the monthly report, the three numbers on the left prove the work landed; the two measures on the right decide whether those numbers can be read as a rise or fall
Three firm numbers (don't depend on AI's swings) || Two measures that must be printed alongside them
Pages on the list X/20 (including on the list and cited) || This month's noise band ±N seats
Reviews and star rating || Web control leg · named-set overlap
Citation slots gained / lost || Web control leg · source overlap
```

The reviews cell is explained by side: **on a side that may not actively ask for reviews, reviews are only an observation item**; on a side that may ask for them, reviews are printed in the report but are never the main effort (what each side may do is in （→ 通用版 5.3 三侧速查（二）：证言与评价、比较、榜单、头衔、外链、FAQ 与图注）).

Page 2 of the monthly report is the competitor board: who AI is recommending this month, who's up and who's down, who pushed you out (down to which question, which source page), your seat share, and **the competitor's way in** (bylined submission / labelled sponsored or paid / platform-edited entry / user-submitted profile / cannot tell — pick one of five, all observable from the page). Fill in this column each week while you're at it, 2 minutes per entry — your competitors have already run the experiment of "what works on this shelf" for you, and **without this column you only know which page to get into, not how to get in**. The page-1 sample and the competitor board's column headers are in （→ 通用版 B.4 工单、台账与月报表头）.

> ⚠️ **The competitor board names third-party organisations and is internal material**: it must not be used in any advertising, posted, forwarded, or made public. Material that can go external is marked item by item.

## 7.4 Holding position and the 90-day settlement

**What you'll do in this section**: for a client who already holds positions, come away with a monthly five-item holding checklist and one sentence you can say externally; at week 13 (D90), come away with an honest before-and-after comparison and a keep-or-cut verdict reached by the criteria, not by hope.

### Holding position: add nothing new, keep what you have

```steps Figure: Holding position is five items a month; add nothing new, only keep what you have — skip one and you lose ground silently
Question ② | Does it recognise you | How AI answers when asked "is X reliable"
Question ④ | Are the facts right | Including whether the price was misquoted
Question ⑤ | Negative or excluded | Whether there's a negative or exclusion tendency
M questions | Retest the M questions | The M questions from the battlefield pile and the hold pile, same-week retest for the noise band
Off-site | Competitors and off-site | Competitor board + the off-site three numbers
```

The first three items are three of the questions in the brand six (the six questions are defined in （→ 通用版 3.5 品牌六问与发现错误之后）). Four things the figure can't show:

- **What to run every month**: the M questions + the brand six (6 questions × 3 rounds × 2 engines = **36 runs/month**) × the same set of engines, archiving the original text of every answer; same-week retest for the noise band; recheck the five profiles' consistency. **Once the question composition, rounds and engines are set, they're frozen — no changes for the rest of the quarter** — change them and the factual-errors counts before and after stop being comparable.
- **How to check**: ① factual errors (count): baseline X → this month Y, each one with a screenshot of the AI's original text, **anyone can reproduce it just by asking the same question** (X is taken only from that one freeze-week run of 36, （→ 通用版 4.3 粗筛分堆、满轮与品牌六问基线）) ② the number of times a negative or exclusion tendency appeared ③ profile consistency 5/5.
- **Correction**: fix everything on sources you control **within 3 working days** (the own-site facts page, the five business profiles) + send a correction letter under your own name to every third-party page that got it wrong, and keep a record. Whether and when a third party adopts it is up to them — **we make no promise there, only the two follow-ups at +7 and +21, plus a monthly recheck**. Notify the relevant parties within 1 working day of discovery.
- **Claiming profiles (one-off, about 2 hours)**: Foursquare + Yelp + Apple Business, all free. **70% confidence**: the approved line is that it's free and done in passing, **not that "filling it in gets you recommended by ChatGPT"**. The evidence (which data sources ChatGPT's local answers draw on, and which year the partnership was signed) is in （→ 通用版 A.1 机制、门与人这一层的证据）; how this confidence figure is printed is in （→ 通用版 A.2 选点、写页、站外与复测的证据）.

> **Say only the verifiable cell externally**: "At baseline, AI got X facts about us wrong; this month it's down to Y, each with a screenshot of the original, and you can reproduce it yourself just by asking."
> ❌ Do not say "most of the 90-day increase will land on brand questions" — that is a **tabletop projection, not a measurement** (70% confidence), for internal reference only.

### Why ninety days

The public line: visibility usually only starts moving after **two to eight weeks**, and a meaningful lift in citation share takes **three to four months**. Seats don't stay won — ask the same question three times, and only a tiny few of the sources cited last time are still there; **an answer has a limited number of names, and the moment a competitor follows suit, the gain is cancelled out**. So this is ongoing occupation, not a one-off build; off-site assets decay too (（→ 通用版 6.6 年度榜单与站外资产腐烂）) — this is the only answer to "why keep doing this" that you can give without lying.

### Three things, each in its place

| Settlement measure | What it is | Strength |
|---|---|---|
| **Deliverables list** | Thick pages launched and accepted, claimed profiles and statutory-register materials completed and checkable, paid slots live, retrieval-crawler door fixed with live-fetch evidence attached, month-by-month original-text comparison reports | **Can be treated as a hard commitment**: everything is in your own hands, self-verifiable, and involves no prediction of a third-party engine's or third-party editor's behaviour |
| **Seat count** (before-and-after side by side + noise band) | Three points — day 0 / day 60 / day 90 — **the same batch of questions, the same set of engines, the same number of passes**, counted item by item, printing total names, seat share and this quarter's measured noise band alongside them | **List side by side only, never promise a figure** |
| **Pages on the list X/20** (split into two lines, "on the list" and "on the list and cited") | The denominator is the frozen baseline Top 20 | **Can serve as hard evidence** (it's either on the page or it isn't, unaffected by sampling noise) |

The Wilson three-state appears only on the internal review page; externally, at most, it's used to describe a trend.

> ⚠️ **Seat count scales linearly with the number of passes**: baseline 5 rounds, D90 10 rounds — **it would double even if you did nothing**, and "seat increase > noise band" would necessarily come true. So **the number of passes is frozen for the whole quarter; D90 adds not a single extra pass**. If you must add passes, the only option is to **start a new baseline**, writing "**not comparable with the previous baseline**" plainly on the same page; the old and new **must not be added together or drawn on the same trend chart**.

### D90 settlement: the order of steps

```steps Figure: D90 is settled in this order; the main verdict comes only from stage 1, and the wider sample never enters the before-and-after comparison
1 | Three-point comparison | Day 0, D60, D90; same questions, same engines, same number of passes, 0.5 d
2 | Wider sample, stored separately | 10 passes × 2 engines per question, produces only the shelf map and competitor board, 1 d
3 | Same-week retest | 3 times, produces this quarter's noise band, 0.5 d
4 | Count seats | Shelf map at day 0 vs D90, including total names, 0.5 d
5 | Final on-the-list check | Screenshot the baseline Top 20 item by item, marked as two lines, 0.5 d
6 | Stop list | Stop investing in pages asked 30+ times cumulatively and never cited
7 | Notarised sample | Logged out, real browser, no more than ten questions, 0.5 d
8 | Write the report | Before / after + a separate brand settlement, 1.5 d
```

Points the figure can't show:

- The wider sample is stored in `probe_d90_wide/`, **kept separate from `probe_d90/`; the two must never be added together or shown side by side on the same trend chart**, and the sampling spec is labelled separately in the figure. The shelf map is internal material.
- The stop list goes into the report with reasons given. The notarised sample **is never the headline number** — it only backs up reproducibility.
- Stage 8's "before / after" writes four things: the shelf three months ago / now; which questions went from absent to present, which didn't move and why (citing which layer triage stopped at); how competitor positions changed; the off-site three-number curve. A separate page for the brand: factual-errors count, baseline X → now Y, each with a screenshot attached, plus a *Try it yourself* card.
- There is also a separate internal three-state review (the difference between your pages and the control pages, before and after), **done once a quarter, not part of the monthly process, and never external**.

### The settlement verdict

```mermaid Figure: Settlement checks admission and deliverables first, then whether there's a determinable increase, and finally which layer triage stopped at; you may not conclude "no effect" without completing triage
flowchart LR
  a{"Will publish a checkable price, edit site?"} -->|never willing| g0["Hold position only"]
  a -->|yes| d{"Are all deliverables accepted?"}
  d -->|something unfinished| fix["Finish deliverables, push deadline"]:::warn
  d -->|all done| q{"Seats exceed the band, or cited pages +3"}
  q -->|yes| ext["Keep expanding: next batch of questions"]:::hl
  q -->|no| t{"Which layer did triage stop at?"}
  t -->|layers 1–4 all pass| hold["Switch to hold, no gloss on the report"]
  t -->|layer 3, not reachable| swap["Wrong target, swap the whole topic"]
  t -->|stopped at layer 1, 2 or 4| no["Don't conclude no effect"]
```

The full statement of the criteria:

- **Keep expanding**: the seat increase > this quarter's measured noise band (**must use the same number of passes, same engine, same question pool**; if the number of passes was changed for any reason this quarter, this cell is judged "not comparable" across the board and may not be used — look at pages-on-the-list instead), or **on-the-list-and-cited pages increase by ≥3**. Keep expanding = move to the next batch of questions and redo target-picking and the shelf snapshot.
- **Switch to hold**: deliverables all passed, seat increase did not clear the noise band, on-the-list didn't increase either, and triage passed layers 1–4 all the way through. Stop adding new pages, and only do the five holding items above; write plainly in the report, "this quarter did not achieve a determinable seat increase on the main-attack questions" — **no glossing over it**.
- **Finish deliverables first**: **never offset a deliverables shortfall with a visibility number**.
- **Wrong target**: triage stopped at layer 3, "citation slots are all at unreachable sources" — move this whole intent out of the main-attack pool and actively change topic. **Don't count a targeting mistake as "this industry can't be done"**.
- **Admission condition**: the website must state a **definite, checkable price**, in the form set by the side (（→ 通用版 5.2 三侧速查（一）：先判身份；价格、促销与赠送、结果数字）'s price figure); a client who's never willing to write one, or whose site can never be changed, gets hold position only. The price page is one of the few cells on an owned site that can actually get into an answer — **not giving a price is the same as giving up that cell**.
- ⚠️ **On the strictly regulated side, admission is judged by itemised fixed pricing**: on this side, writing a "price range" is itself non-compliant. Judging a client as "not giving a price" because it has no price range would wrongly disqualify a project that is actually compliant and simply can't write its price using the general template.

**The decision rests with the data, not with hope.** But "no effect" must be a conclusion reached only **after triage is complete and the seat change has cleared the noise band** — **not a side effect of the ruler's resolution being too coarse**. At settlement, these three sentences may not be said to yourself:

> ❌ "Give it another three months and it will definitely rise"
> ❌ "The industry average is six months"
> ❌ "The data is trending up" (when the seat increase has not cleared the noise band and on-the-list hasn't increased either)

## 7.5 Ruler discipline: change the ruler and nothing is comparable

**What you'll do in this section**: pin down "what counts as changing the ruler" into a reference table, and pin down "whether the baseline should be re-frozen this month" into one mechanical test. When you're done, you'll have a way to handle a changed ruler, the rule for pinning the probe model, two ways to freeze the baseline when the shelf gets rewritten in bulk, and a list of which numbers can be printed and which stay internal only.

### What counts as changing the ruler

| Action | Consequence | What's allowed |
|---|---|---|
| Change engine | **The ruler has changed** | Start a new baseline, and write plainly "not comparable with the previous one" |
| Change the AI model (including a minor version) | **The ruler has changed** | Same as above |
| Change the phrasing / question pool / ratios | **The ruler has changed** | Same as above |
| Change the number of passes | Seat count **scales linearly** | Same as above, and the old and new must not be added together or shown on the same chart |
| Change the web control leg's 5 questions | **The calibration ruler has changed** | Do not change it for the whole quarter. If it must change, change it next quarter |
| Change the brand six's number of questions or rounds | Factual-errors count **before and after is not comparable** | Frozen for the whole quarter |

**Across versions, list the numbers only — do not compute a rise or fall.** When numbers from two different baselines sit on the same page, they must be printed in two separate blocks, with a line "Below: a different ruler" between them. This table and the noise-band rule in 7.2 are two scales of the same test: the noise band governs sampling variation inside the same ruler; this table governs the moments when "you've actually already switched to a different ruler".

### Probe models: pinned, never drifting with your main model

The model used to run probes **has a fixed version — it never drifts along with your day-to-day main model**. Name the legs as in 7.1; pin the model version per the table below:

| Leg | Rule |
|---|---|
| **① OpenAI API leg** | **Must connect directly to the official API**. It measures ChatGPT itself — routing through any relay or proxy pool is measuring A to sign off on B |
| **② Gemini model leg** | May go through a proxy pool (when the free tier's direct quota isn't enough, the whole leg comes back empty — a "high-fidelity" path that can't get through is no more honest than a proxy path that can). **If you go through a pool, pin the exact model version**: a floating alias returns 503, and that whole month's sample is gone |
| **Both legs** | The version only ever moves on an explicit decision, and **every leg must move together**. Changing the version = changing the ruler, and it makes the new report incomparable with old reports already sent out |

The cost has to be written into how the report is worded: when sampling through a proxy pool, region and the model's minor version can't be independently verified — print this sentence in the methodology note.

### Whether to re-freeze the baseline this month: one mechanical test

```mermaid Figure: Whether to re-freeze the baseline this month rests on one mechanical test only — never a gut feeling
flowchart LR
  m["Monthly recheck of baseline Top 20"] -->|check the update date item by item| c{"≥6 items updated after freezing"}
  c -->|no| keep["Don't re-freeze, keep comparing in this segment"]
  c -->|yes| new["Re-freeze this month, open a new segment"]:::hl
  new --> old["Old segment stays historical only"]
  new --> nc["Write old vs new X/20 as not comparable"]:::warn
```

Why this test is needed: put "the denominator doesn't change all quarter" (7.1) together with "the baseline isn't frozen during a window of concentrated shelf rewrites", and a gap opens up — the baseline freezes in some month, and afterwards most of the Top 20 pages get rewritten in bulk, so the denominator has already come apart from the shelf, yet discipline says you still can't change it. The test in the figure is what plugs that gap (not measured; the mechanism is clear).

Three points the figure can't show: "updated" only counts the **visible update date** on the URL, compared against the last freeze date; the old and new segments' X/20 **must not be drawn as one single line on a chart**; the monthly report for the re-freeze month **must state exactly which URLs triggered it**.

### Two ways to freeze the baseline when the shelf has predictable rewrite windows

Some shelves have bulk-rewrite windows you can see coming: the same batch of sources gets rewritten around a particular point in time, and a rewrite = the candidate set gets swapped out (not measured; the mechanism is clear). Freeze the baseline inside the window, or let the baseline and the retest fall into two different phases, and the ruler is effectively broken for that month — any rise or fall is entirely produced by the window, not by anything you did.

```split Figure: When you hit a bulk-rewrite window, pick one of two options and write it into the kickoff document — never choose on the day of the retest
Option 1 · Avoid it || Option 2 · Record one per phase
Applies when: the window can be avoided || Applies when: it can't be avoided, e.g. the window covers half the year
Schedule both the baseline freeze and the retest outside the window || Measure once before the window and once after
If a measurement point lands on the window, move it earlier or later — never move the work schedule || Record two noise bands separately: normal-phase band / window band
Every comparison lands in the normal phase || After that, compare same phase to same phase only, never across phases
```

Whichever option you choose, write the phase you're in at the time of freezing into the **first line of the baseline-freeze file**, and label the phase on every line of the monthly comparison; when numbers from different phases are shown side by side, state plainly "different phases, not comparable". Why this must be fixed in advance: getting this wrong doesn't throw an error — it just makes the whole month's report point the wrong way. Just like a misjudged door, the costs are asymmetric.

> **Example** (e-commerce) During a major sale, prices, stock status, platform ranking and buyer search terms all change at once, and both seat rulers' noise bands blow out. Normal phase = outside the two weeks before a major sale and outside the two weeks after one; for any month where the retest window and the baseline window fall in different promotional phases, always write "cannot be determined (cross-promotional-phase)" — you may write neither a rise nor a fall.

### Different confidence, different printing

Low-confidence work still gets done — it just has to be **flagged when printed**, carrying the qualifying sentence exactly as given. Which of this chapter's statements get printed, and which stay internal only:

| Statement | Printed or not | How to print it |
|---|---|---|
| This month's measured noise band | Printed | The value goes on page 1 of the monthly report; before the first measured value is obtained, seat counts are only listed side by side, with no verdict |
| Completing the five profiles benefits ChatGPT's local answers (from a public partnership) | Printed | State it on the basis of "free, done in passing" — never promise an effect |
| Improving position is cheaper than winning a new slot | Printed | Carry the qualifying sentence exactly: cross-industry pooled basis, not verified in this local industry |
| Fallbacks split by content form | Printed | Carry "does not include Singapore, about 75% confidence" |
| The rate at which a review is quoted word for word into an answer (external live search, a very small sample) | Not printed externally | Used only to decide direction of work; the monthly report prints your own measured values for the M questions, split into recommendation-type / non-recommendation-type lines, and the two lines must be shown together |
| The share of the 90-day increase that lands on brand questions | Not printed | A tabletop projection, not measured, internal reference only (the "do not say" line from holding position in 7.4) |
| An intent where "0 businesses are named, and someone is already fighting for the citation slot" (external live search, one-off) | The specific figure is not printed | Used only for internal ranking; the report prints the named-business count and citation-slot ownership from your own baseline numbers, with the retest date and engine stated |

The reviews row requires the two lines side by side because printing only the recommendation-type line would lead people towards asking for reviews, which the strictly regulated side cannot do at all. The last row carries one more red line: ❌ no material may ever say "no one is doing this" or "competition is near zero".

The percentage, sample and source behind each row are in （→ 通用版 A.2 选点、写页、站外与复测的证据）; the general rule for how to flag a number and what stays for your eyes only is in （→ 通用版 D.1 数字纪律：数字怎么标、什么只能自己看）.