Skip to main content

GEO Playbook · General Edition · Chapter 7 (12 of 17)

Monthly retest: how to tell whether it's working and what to do next month

Ask the same questions every month, and tell real change from random noise

Ask again every month with the ruler you froze for good. The verdict is never "it worked" or "it didn't" — write which three points to hit next month. Alongside it, keep three firm numbers that don't depend on AI's own randomness, to prove that this month's work really landed.

7.1 What a retest produces, and the rulers

What you'll do in this section: turn the monthly retest's verdict from "did it work" into "which three points to hit next month". When you're done, you'll have five rulers, each answering one question, the limits each one must print exactly as they are, and four sentences you'll never say again.

The order is fixed: read the free first-party data first, then look at the probes that cost money.

Why we do not judge with statistical tests

Statistical judgement (internal reference only)Raw counts + noise band (used for the verdict)
Wilson three-state: moved / seems to have moved / not movedSeat count, pages on the list, factual errors (count)
The threshold for "moved" is out of reach on a small sampleRetesting 3 times in the same week is enough to measure the noise band
Default output: "not demonstrated"Output: next month's three work orders
Only used to trigger not-moved triageUsed for the verdict, and to schedule the work orders
Permutation test: at n ≤ 10 it changes no decisionDon't run it — that's just ritual
Figure: On a small sample, statistical judgement can only output "not demonstrated"; raw counts + a noise band are what actually produce next month's work orders

Run it on a statistical basis, and week 13's default output is necessarily "not demonstrated" — you walk in already carrying the conclusion "this didn't work", whether or not anything actually moved (for how the threshold is worked out, see the evidence in → General Edition A.2 Evidence for picking targets, writing pages, off-site and retests). The fix isn't to loosen the statistics — it's to change what you accept as proof. Rankings are moved by work orders, not by p-values.

If you're doing this yourself, the minimum viable ruler set = three numbers + one page-level signal + one web control leg. You don't need a Wilson interval, a control page, or a permutation test — but you do need to honestly separate "actually moved" from "noise". The three numbers are the seat count, on the list and cited, and factual errors (count) below; the page-level signal and the noise band are in 7.2; the full specification of the web control leg appears only in → General Edition 4.4 The web control leg and the frozen baseline (the book's only full spec).

Name the legs correctly first

The two API legs' official names are ① OpenAI API leg and ② Gemini model leg. The diagram mapping the three products to their three rulers is in → General Edition 1.2 Two legs: ChatGPT looks for the source, AI Mode for second-hand summaries; here we only add the three mistakes people most often make at retest:

  • Treating the Gemini model leg's number as "visibility in Google's AI Overviews". Gemini API ≠ AI Overviews ≠ AI Mode — the fetch path, the way the answer is generated, and the source pool all differ; passing one leg's number off as another's is measuring A to sign off on B. Visibility for AI Overviews / AI Mode uses only the impressions from the Search Console generative AI report.
  • Checking Google-Extended to judge whether AI Overviews can crawl you. You should check Googlebot. Google-Extended governs training and the Gemini app side; blocking it does not affect whether AI Overviews indexes you — it's Googlebot being blocked that means the door is shut.
  • Describing the API legs' results as "what users will see in ChatGPT". Where the only evidence is from an API leg, the wording must always be qualified as "on the model API".

Five rulers, one question each

are you on the page

how many times named

are the facts right

the verdict threshold

impressions on Google's side

add up

add up

On the list X/20 (two lines)

Side by side, each with its limits

Seat count = times named

Factual error count

Same-week retest noise band

GSC AI Overviews impressions

Combined into one visibility score

Figure: Five rulers, one question each — print them side by side only, each with its limits attached; add them into one score and any one going bad gets covered up by another

How each ruler is counted, and the limits that must be printed exactly as they are on the report:

RulerHow you count itThe limits that must be printed exactly as they are
Pages on the list (report as two lines: on the list X/20 and on the list and cited this month Y/20)Of the frozen baseline Top 20 sources, how many pages now actually feature you (open each one by hand + screenshot). The denominator, 20, does not change all quarterOn the list ≠ cited: AI may cite this page without mentioning your line — being on the list is a necessary, not a sufficient, condition. The reading is fixed: on-the-list rises but on-the-list-and-cited doesn't, for two months running → you're investing in a container that isn't being cited; next month, invest only in sources that were actually cited this month
Seat countHow many times your name was named across this month's N question-and-answer runs; also log this month's total names (the sum of every business's name count) and each competitor's seat countIt's a sampled count under your own question pool, engine and number of passes — it does not represent the whole market; total names shifts with how long the answer is, so you must print total names and seat share together; it scales linearly with the number of passes, so if the number of passes changed, the two numbers must not be shown side by side or subtracted from each other
Factual errors (count)A person reads the brand six (6 questions × 3 rounds × 2 engines = 36 runs / month, frozen for the quarter, defined in → General Edition 3.5 The brand six questions, and what to do when you find an error), copies down every fact AI got wrong, and attaches a screenshot of the originalThe hardest cell in the whole ruler set; it measures only whether the facts are right, not ranking. If you can only watch one number, watch this one: the error is copied by AI from the source page it retrieved, so fixing the source moves both legs together — of all the rulers, it is the least sensitive to skew between the API and web paths
Same-week retest noise bandSee 7.2The noise band itself is also a sampled value
AI Overviews / AI Mode impressionsGSC generative AI report impressions, exported by page and by country (free, first-party)Only impressions — no names, no position; it's already included in total impressions, so it must not be added to total impressions; flag it when a blank value is exported as 0

Seats are simply times named, and the measure has only one name. Do not write "the raw seat count" and "the answer-level named count" as two separate lines — they are the same number, and the only effect of splitting them into two is that the noise figure you're banned from reporting on its own sneaks back into the acceptance test under a different name.

How you split changes by industry; the definitions do not

The definition of a ruler doesn't change by a single word; the only thing you can change is how you split it:

  • Seats can be split into two: when a category has both an answer shelf and a shopping shelf, count brand seats (answer shelf) separately from SKU card placement and top-spot share (shopping shelf), each with its own noise band. This answer's total names is fixed by measuring the category itself — never reuse another category's number.
  • The denominator can be split by work unit: when two fronts' named sets barely overlap (by product line, or by school stage × subject), report the denominators separately; merge them into one denominator and the ups and downs of two independent fronts will cancel each other out.
  • Factual errors can be split into two columns: log "errors in your own source" separately from "third-party errors" — the two have completely different repair paths and time-to-effect (7.3, layer 2).
  • How many rounds go into the noise band is an industry variable — measure it yourself in the first quarter. The fourth reference number (things like enquiry source) is chosen per industry; it's the only cell that varies by industry, and it never enters acceptance.

Example (e-commerce) Seats split into two: brand seats (answer shelf) + SKU card placement and top-spot share (shopping shelf); the noise band is also measured as two, one for brand seats and one for SKU card placement.

The source skew between the web version and the API is systematic, not noise, and adding more rounds doesn't fix it — the number appears exactly once, in → General Edition 4.4 The web control leg and the frozen baseline (the book's only full spec).

Four things never to say

Never saySay instead
"We're on the list, so AI will recommend us"On the list X/20, of which cited this month Y/20
"Seat share 1.29% = market share 1.29%"Asked K times, named X times, noise band ±N
"Up 2 seats, heading the right way" (within the band)Within normal variation, cannot be determined
Adding two rulers into a combined "visibility score"Each ruler on its own line, side by side, with its limitations
Figure: Four things never to say, each paired with a reproducible way to say it instead

Writing a report, briefing your boss or a client, or even summing up in your own head — the four sentences on the left must never appear. Why: on the list ≠ cited; seat share is a sampled count under your own question pool, not the market; a rise that hasn't cleared the noise band is treating noise as a result; and two rulers with different denominators and different failure modes, once added together, let either one go bad while the other one covers for it. The noise-band statement that must be printed word for word in the monthly report is in → General Edition B.2 Approved wording and scripts.

Retesting also has three more criteria; the full statement is in → General Edition 1.4 Three kinds of things not to do — here they're just applied:

  • An observation that is single-run, single-engine and n ≤ 3 can only serve as directional evidence — you cannot declare "zero visibility";
  • The wording of the question pool can produce a false zero: when bare category words dominate, suspect the ruler first (7.3, layer 4); evidence in → General Edition A.2 Evidence for picking targets, writing pages, off-site and retests;
  • Even if pages-on-the-list has risen, still run triage — the trigger looks only at seats (7.3).

7.2 The noise band, the page-level signal and the ten monthly steps

What you'll do in this section: get three more things every month — a measured noise-band value, a list of page-level-signal hits, and a monthly schedule with people and hours assigned. With the noise band you know which rises and falls you're allowed to write into your conclusion; with the page-level signal you know which page was actually read and taken.

The noise band: the cell people most often get wrong

What to do: run the same batch of questions, same engine, same number of passes, again 3 times within the same week, giving X′ / X″ / X‴. How long: about 1.5 hours/month (the machine runs it; a person only checks it lands on disk). How to check: three independently saved answer packs, all timestamped within the same week; the noise-band value is written on page 1 of the monthly report.

yes

no

no

yes

no

yes

no

yes

Month 1 or 2?

List side by side only, no verdict

Each leg's sample ≥ 80% of plan

Mark 'instrument incomplete', no comparing

Increase exceeds the noise band?

Within normal variation, cannot be determined

Same direction two months running?

Above the band this month, to be confirmed

Write 'increased'

Figure: A rise in seats doesn't mean you can write "increased" — it has to pass four checks in a row, and each one you fail has a fixed way to write it

The wording at each check in the figure is fixed: for months 1–2, mark "noise band baseline not yet stable (K/3 months recorded)"; in a month where the sample falls short, print each leg's figures separately, do not print a total, and do not compare with last month. What the figure can't show is how the noise band itself is defined:

RuleWhy
Run it 3 times, not onceIf you use one sampling difference as the threshold for another, then under the null hypothesis that "nothing happened" the two have the same distribution, so the false-positive rate is about 50%. Running it 3 times brings the false-positive rate down from 100% to 50%, not down to 0
Noise band = the range (max − min); the floor is fixed at ±1 seat — even if all three runs come out identical, record it as ±1When the range is 0, any +1 would be judged "increased" — the gate fails in reverse
The noise band baseline is the median of three months, not the maximumTaking the maximum is a monotonically non-decreasing ratchet — the longer you run it, the harder it becomes to declare any improvement, and you end up in a closed loop where "the worse the ruler, the more it excuses you"
If the noise band widens for two months running → add more passes and retest; you may not invoke "instrument defect, not counted this month" againAn excuse clause used more than once becomes an excuse machine. Adding more passes is changing the ruler — per 7.5, start a new baseline
If one leg's valid sample that month < 80% of the planned volume → mark that month "instrument incomplete"When one leg is missing for the whole month, the absolute cross-engine count gets cut in half. This is the same problem as "adding passes doubles the seat count" — just the other side of it

The page-level signal: the only page-level evidence of cause

retest monthly

quoted back

Log 2–3 unique facts at launch

String-search the saved answer text

This page was read and taken as a source

Figure: The page-level signal measures "was it read or not", not "which position" — so n = 1 still counts

What you log must be a checkable fact string unique to that page — for example, the S$6,800 fixed into that page's price list. Adjectives, or sentences that also appear on other pages, don't count; finding them proves nothing about which page was actually read.

It has nothing to do with seat count or the noise band, and reuses the answer-text packs you're already storing — the marginal cost is close to zero. Seat count tells you whether things are rising overall; the page-level signal tells you which page is doing the work — record the two numbers separately, and never infer one from the other.

Ten monthly steps, grouped into eight stages

  1. 1First-party dataStep 1: Bing AI Performance + GSC, 0.5 h
  2. 2RetestStep 2: same question pool, same engine, same number of passes, 2 h
  3. 3Noise bandStep 3: run again 3 times in the same week, 1.5 h
  4. 4Count seatsStep 4: names, position, total names + page-level signal, 1.5 h
  5. 5Recheck on-the-listStep 5: X/20 + cited this month or not, 1.5 h
  6. 6Profile + brand-six auditSteps 6–8: profile consistency, 36 brand-six runs, conventional search audit, 2.5 h
  7. 7TriageStep 9: the five layers in 7.3, 1 h
  8. 8Monthly reportStep 10: sent on a fixed date, 2 h; total about 12.5 h
Figure: Ten monthly steps grouped into eight stages; the first six stages are just bookkeeping — what moves rankings is stage 7, triage

Steps 1–4 are only responsible for accurately measuring what the shelf looks like right now; what actually moves rankings is step 9 — without step 9, the first eight steps are just bookkeeping.

Jump-the-queue rule: if AI produces a false statement about you or shows a negative tendency → jump the queue and fix it first; don't wait for the monthly report (the five situations that jump the queue are in → General Edition 3.5 The brand six questions, and what to do when you find an error).

The web control leg's 40 minutes is not in this table — schedule it as a separate day at the start of each month; the full specification is in → General Edition 4.4 The web control leg and the frozen baseline (the book's only full spec).

Who does each step, what they hand over, and who checks it:

StepWhoOutputWho checksPoints the figure doesn't show
1The report ownerfirst_party_<month>.jsonSelf-checkExport this month's landing-page impressions and Citation Share
2The person running probesprobe_<month>/, the raw text of every answerSelf-checkPrint "not measured" for any engine not measured because the quota ran out
3The person running probesprobe_<month>_retest/ + the noise-band valueThe report owner—
4The person running probesseats_<month>.csv + the page-level-signal hit listThe report ownerWhich businesses were named in each answer, at what position, and this answer's total names
5The off-site leadoffsite_ledger_<month>.csvThe project leadOpen every item in the baseline Top 20 by hand: are you on it, at what position, is the information out of date; while there, tag each one with the five binary page-type features
6The profile ownerSame ledger as aboveThe project leadThe three review numbers are a passive observation; five-profile consistency
7The person running probesbrand-6.jsonl + entity_errors.mdRead by a person, no skipping6 questions × 3 rounds × 2 engines = 36 runs, producing the factual-errors count
8The person who edits the siteAudit page (with the four identities' four receipts + screenshots of the three nosnippet greps)The project leadNon-brand clicks, the canonical Google picked, the index status of key URLs, that the four identities' fetches still match (defined in → General Edition 2.3 Gate 2: four identities, fetched live (read-only)), whether the three nosnippet controls have recurred this month
9The project leadNext month's work ordersLook it over again the next dayRun the 7.3 triage, produce next month's three points
10The report ownerThe monthly report + the raw answer packRead page by page by a personSent on a fixed date

7.3 Not-moved triage and next month's three points

What you'll do in this section: for a month when seats don't clear the noise band, come away with a triage verdict of "stopped at which layer", plus next month's three work orders, laid out mechanically by priority order; add to that a fixed-format page 1 for the monthly report, and a competitor board for your eyes only.

There is only one trigger

This month's seat increase did not clear the noise band → always start at layer 1; no skipping layers. Whether pages-on-the-list rose or not has nothing to do with this rule.

Why you can't add "and pages-on-the-list did not increase": pages-on-the-list is the output of your own work, and a monthly +1 is almost routine. Fold it into the trigger as an AND condition and you've fitted the whole triage table with a switch you can flip off whenever you like, and no month all quarter will ever really go through triage. No variable you control yourself may appear in the trigger condition — any new trigger condition proposed later must be checked against this rule.

Five layers, cheapest first

fails

passes

fails

passes

no

yes

hits

clears

Seat increase did not clear the noise band

1 · door: can it be read?

fix same day, recheck in 7 days

2 · identity: got the right business?

page work resets to zero, fix identity first

3 · citation slots line up?

handled by the three verdicts in the table below

4 · split day and ruler

doesn't count as not-moved this month

5 · change topic: switch battlefield

Every month, unconditional

Figure: Not-moved triage checks layers cheapest-to-rule-out first, layer by layer, no skipping; the door and identity layers are checked every month unconditionally, even when seats rose

The door and identity layers are checked unconditionally because failure at either one is silent: when a page was never read at all, or AI has the wrong business, seat count can perfectly well be rising for some other reason, and you won't find out until next month — by which point a whole month has been wasted. Which fix each door-layer symptom maps to is looked up in the two diagnostic figures in → General Edition 2.9 Door-layer troubleshooting and this chapter's acceptance checks; they aren't redrawn here. Three of these are the easiest to misjudge: the Google leg has seats while the ChatGPT leg is zero — check CSR first, not your question selection; don't use GPTBot as your evidence, it's a training crawler — for the retrieval legs look at OAI-SearchBot and ChatGPT-User; and don't draw a conclusion from what curl -A "OAI-SearchBot" returns — both directions of that test lead to the wrong verdict.

LayerWhat to check that dayTest and actionHow to say it in the monthly report
1 · Door① Access logs: in the past 30 days, has a real OAI-SearchBot hit these pages, and what status code came back (the only direct evidence) ② have the three nosnippet controls been switched back on by the CMS or a plugin ③ do the four identities match word for word ④ site: search to check Bing indexing ⑤ paste the URL into ChatGPT and have it quote back the price line word for word ⑥ is the Foursquare category and business status right, is the Yelp entry thereFailing any one = a door problem, not a targeting problem: fix it the same day, recheck in 7 days, don't change topics or add new content pages this month"Last month ChatGPT simply wasn't reading those pages at all; the reason was X, it's fixed today, we'll recheck in 7 days"
2 · IdentityRead the brand six by hand: is it talking about this business? Are the name, address, registration number and price right? Any namesake confusion? Any complaint or negative tendency?Wrong business, or facts wrong → all page work resets to zero and gets recalculated; fix identity first: tick off and screenshot each of the six controllable sources (own-site /facts → Google Business Profile → Bing Places → Foursquare → Yelp → Apple Business Connect) + send a correction letter to the cited source. Mark every error as either "third-party page source" or "business profile source" — mixing the two makes profile-type errors run late systematically"AI mistook you for another business (or got your price wrong); this month we deal with that first"
3 · Citation slotsThe intent-level top 10 URLs by times cited this month, checked one by one against the four columns of the citation-slot table (four states / judged on / basis / judged by)Judged reachable at the time, entry point gone now = the shelf really changed → switch to "hold", go after that new source, put it first on this month's off-site work orders
Judged "to ask" at the time, or the basis column is blank = the judgement drifted → fill in the basis this month and recompute the abstain line; this month it doesn't count as a wrong target
Fewer than 2 reachable-enough URLs in the intent-level top 10 = wrong target → move this intent out of the main-attack pool and swap in another; pages already invested in become long-tail landing pages and get no more investment
"A got into some directory last month and pushed you from 3rd to 5th; go fight for this this month" / "Last month's basis for this one wasn't kept on record; fill it in this month, then judge" / "We can't get into this slot for this question because we picked the wrong battlefield at the start; swap it"
4 · Split day and ruler① the page has been live less than two weeks ② the URL changed after launch ③ are the frozen questions bare words ④ has the engine / model / number of passes been touched ⑤ is the noise band bigger than last month ⑥ is the web-control-leg overlap < 50% ⑦ was one leg's valid sample < 80% of the planned volumeAny hit → recompute the split day, or mark the question pool "instrument defect", and it doesn't count as "not moved" this month. If ⑥ hits, this month's seat conclusion always gets the non-extrapolation sentence printed alongside it (→ General Edition 4.4 The web control leg and the frozen baseline (the book's only full spec)), and it must not be used as a reason to change topics. The "instrument defect" excuse may not be used in two consecutive months: the second time, always add passes and retest instead"This batch of pages has been live less than two weeks; the first comparable point comes next month"
5 · Change topicLayers 1–4 all pass, and it still hasn't movedSeats are saturated, or the demand isn't real → this is the only case that's actually "change topic": go back to target-picking and pick again, and it must be a different battlefield, not just different wording"The shortlist for this question is locked; swap in a battlefield we can actually get into"

Two disciplines:

  • You may not skip layers 1–4 and go straight to writing "demand doesn't exist".
  • Stopping at the same layer, unfixed, two months running = an execution problem, not a shelf problem — handle it ahead of the queue next month, and the same excuse may not be used again.

The definitions of the four states, the abstain line and the citation-slot table are in → General Edition 4.6 Four states, two denominators and the abstain line and → General Edition 6.2 The send gate, paid listings and the citation-slot table.

Turn the triage result mechanically into next month's three points

layers 1–2 failed

a new reachable source appeared

on the list but past position 4

filled

short · professional

short · consumer

still short

still short

still short

Triage result

Point 1: fix the door or identity

Point 2: capture one off-site slot

Point 3: improve one position

three points filled?

Stop, assign

Fallback A: educational thick page

Fallback B: profiles and facts

Fallback C: expand the price page

Fallback D: Chinese page or long video

Figure: The triage result is mechanically converted into next month's three points, by priority order; stop once you've filled three, and if you can't, fall back by content form — no guessing

How each cell in the figure picks that one item, and what work order it assigns:

PriorityWhich one to takeNext month's work order
Point 1Triage failed at layer 1 or layer 2Fix the door / fix identity. Once triggered, it automatically takes point 1, and no new content pages are added this month
Point 2A source that has newly appeared among this month's cited sources, is judged "reachable" under the four states, is allowed on this side, and does not have you on itTake the 1 source actually cited this month, with the highest citation count, where the competitor's way in ≠ platform-edited entry, and assign an off-site capture work order
Point 3A source where you're already on the list but ranked past position 4Take the 1 with the highest citation count, and assign an update-letter / supply-materials work order. Improving position is cheaper than winning a new slot (cross-industry pooled basis, not verified in this local industry — carry the qualifying sentence exactly when you print it; sample in → General Edition A.2 Evidence for picking targets, writing pages, off-site and retests)
Fallback AThree points can't be filled, and the content form is professionalEducational-thick-page work order: fold the landing page for the billable intent with the lowest seat count into its intent cluster and thicken it. How thick follows the page-type spec — word count is not a criterion (→ General Edition 5.7 Seven rules that hold for every page type)
Fallback BThree points can't be filled, and the content form is consumerProfiles-and-facts work order: fill in the fact fields on all five business profiles, expand the /facts page, complete itemised prices on the price page (wording follows the side, → General Edition 5.2 Quick reference by side (1): identify the advertiser first; prices, promotions and freebies, result numbers), and complete outbound links to official and regulatory bodies. No review work order is assigned on the strictly regulated side: that side may not actively ask for reviews
Fallback CStill shortExpand a price page: price and fee queries trigger AI Overviews at a high rate; evidence in → General Edition A.2 Evidence for picking targets, writing pages, off-site and retests
Fallback DStill shortComplete the Chinese pages, or pair an existing thick page with a long video narrated by the named expert (→ General Edition 6.7 Long videos narrated by the named expert)

Content form (professional / consumer) is the second judgement, made after the side has been decided; how to judge it is in → General Edition 5.9 Content form, FAQ, Chinese pages and other-language pages. The basis for splitting the fallback by form is a cross-industry query sample — it does not include Singapore, at about 75% confidence — the sample's basis is in → General Edition A.2 Evidence for picking targets, writing pages, off-site and retests.

Reddit: read the number monthly; do not fix a conclusion

Read once every month at retest: Reddit's share of this month's cited sources = the number of this month's cited URLs that are reddit.com ÷ the total number of this month's cited URLs (two decimal places, logged on page 2 of the monthly report).

This month's shareAction
> 2%Schedule one Reddit work order this month (target whichever leg it's showing up in)
≤ 2%No investment this month — just log the number; read it again next month
> 2% for two months runningOnly then may Reddit be written into the regular work rhythm

Why we don't fix a rule of "never invest in Reddit for ChatGPT": ChatGPT's Reddit citations once fell sharply in a single month, but the same kind of collapse had happened earlier and fully recovered, and ChatGPT never stopped reading Reddit — what stopped was putting threads into the answer. That's "the answer surface temporarily narrowing", not "this source is dead" (numbers and dates in → General Edition A.2 Evidence for picking targets, writing pages, off-site and retests). A single month's swing is never written into a long-term rule.

Page 1 of the monthly report: three firm numbers and two measures printed beside them

Three firm numbers (don't depend on AI's swings)Two measures that must be printed alongside them
Pages on the list X/20 (including on the list and cited)This month's noise band ±N seats
Reviews and star ratingWeb control leg · named-set overlap
Citation slots gained / lostWeb control leg · source overlap
Figure: On page 1 of the monthly report, the three numbers on the left prove the work landed; the two measures on the right decide whether those numbers can be read as a rise or fall

The reviews cell is explained by side: on a side that may not actively ask for reviews, reviews are only an observation item; on a side that may ask for them, reviews are printed in the report but are never the main effort (what each side may do is in → General Edition 5.3 Quick reference by side (2): testimonials and reviews, comparisons, lists, titles, outbound links, FAQ and captions).

Page 2 of the monthly report is the competitor board: who AI is recommending this month, who's up and who's down, who pushed you out (down to which question, which source page), your seat share, and the competitor's way in (bylined submission / labelled sponsored or paid / platform-edited entry / user-submitted profile / cannot tell — pick one of five, all observable from the page). Fill in this column each week while you're at it, 2 minutes per entry — your competitors have already run the experiment of "what works on this shelf" for you, and without this column you only know which page to get into, not how to get in. The page-1 sample and the competitor board's column headers are in → General Edition B.4 Headers for work orders, ledgers and monthly reports.

⚠️ The competitor board names third-party organisations and is internal material: it must not be used in any advertising, posted, forwarded, or made public. Material that can go external is marked item by item.

7.4 Holding position and the 90-day settlement

What you'll do in this section: for a client who already holds positions, come away with a monthly five-item holding checklist and one sentence you can say externally; at week 13 (D90), come away with an honest before-and-after comparison and a keep-or-cut verdict reached by the criteria, not by hope.

Holding position: add nothing new, keep what you have

  1. Question ②Does it recognise youHow AI answers when asked "is X reliable"
  2. Question ④Are the facts rightIncluding whether the price was misquoted
  3. Question ⑤Negative or excludedWhether there's a negative or exclusion tendency
  4. M questionsRetest the M questionsThe M questions from the battlefield pile and the hold pile, same-week retest for the noise band
  5. Off-siteCompetitors and off-siteCompetitor board + the off-site three numbers
Figure: Holding position is five items a month; add nothing new, only keep what you have — skip one and you lose ground silently

The first three items are three of the questions in the brand six (the six questions are defined in → General Edition 3.5 The brand six questions, and what to do when you find an error). Four things the figure can't show:

  • What to run every month: the M questions + the brand six (6 questions × 3 rounds × 2 engines = 36 runs/month) × the same set of engines, archiving the original text of every answer; same-week retest for the noise band; recheck the five profiles' consistency. Once the question composition, rounds and engines are set, they're frozen — no changes for the rest of the quarter — change them and the factual-errors counts before and after stop being comparable.
  • How to check: ① factual errors (count): baseline X → this month Y, each one with a screenshot of the AI's original text, anyone can reproduce it just by asking the same question (X is taken only from that one freeze-week run of 36, → General Edition 4.3 Coarse screen and piles, full rounds, and the brand-six baseline) ② the number of times a negative or exclusion tendency appeared ③ profile consistency 5/5.
  • Correction: fix everything on sources you control within 3 working days (the own-site facts page, the five business profiles) + send a correction letter under your own name to every third-party page that got it wrong, and keep a record. Whether and when a third party adopts it is up to them — we make no promise there, only the two follow-ups at +7 and +21, plus a monthly recheck. Notify the relevant parties within 1 working day of discovery.
  • Claiming profiles (one-off, about 2 hours): Foursquare + Yelp + Apple Business, all free. 70% confidence: the approved line is that it's free and done in passing, not that "filling it in gets you recommended by ChatGPT". The evidence (which data sources ChatGPT's local answers draw on, and which year the partnership was signed) is in → General Edition A.1 Evidence for mechanics, the door and identity; how this confidence figure is printed is in → General Edition A.2 Evidence for picking targets, writing pages, off-site and retests.

Say only the verifiable cell externally: "At baseline, AI got X facts about us wrong; this month it's down to Y, each with a screenshot of the original, and you can reproduce it yourself just by asking." ❌ Do not say "most of the 90-day increase will land on brand questions" — that is a tabletop projection, not a measurement (70% confidence), for internal reference only.

Why ninety days

The public line: visibility usually only starts moving after two to eight weeks, and a meaningful lift in citation share takes three to four months. Seats don't stay won — ask the same question three times, and only a tiny few of the sources cited last time are still there; an answer has a limited number of names, and the moment a competitor follows suit, the gain is cancelled out. So this is ongoing occupation, not a one-off build; off-site assets decay too (→ General Edition 6.6 Annual rankings and off-site asset decay) — this is the only answer to "why keep doing this" that you can give without lying.

Three things, each in its place

Settlement measureWhat it isStrength
Deliverables listThick pages launched and accepted, claimed profiles and statutory-register materials completed and checkable, paid slots live, retrieval-crawler door fixed with live-fetch evidence attached, month-by-month original-text comparison reportsCan be treated as a hard commitment: everything is in your own hands, self-verifiable, and involves no prediction of a third-party engine's or third-party editor's behaviour
Seat count (before-and-after side by side + noise band)Three points — day 0 / day 60 / day 90 — the same batch of questions, the same set of engines, the same number of passes, counted item by item, printing total names, seat share and this quarter's measured noise band alongside themList side by side only, never promise a figure
Pages on the list X/20 (split into two lines, "on the list" and "on the list and cited")The denominator is the frozen baseline Top 20Can serve as hard evidence (it's either on the page or it isn't, unaffected by sampling noise)

The Wilson three-state appears only on the internal review page; externally, at most, it's used to describe a trend.

⚠️ Seat count scales linearly with the number of passes: baseline 5 rounds, D90 10 rounds — it would double even if you did nothing, and "seat increase > noise band" would necessarily come true. So the number of passes is frozen for the whole quarter; D90 adds not a single extra pass. If you must add passes, the only option is to start a new baseline, writing "not comparable with the previous baseline" plainly on the same page; the old and new must not be added together or drawn on the same trend chart.

D90 settlement: the order of steps

  1. 1Three-point comparisonDay 0, D60, D90; same questions, same engines, same number of passes, 0.5 d
  2. 2Wider sample, stored separately10 passes × 2 engines per question, produces only the shelf map and competitor board, 1 d
  3. 3Same-week retest3 times, produces this quarter's noise band, 0.5 d
  4. 4Count seatsShelf map at day 0 vs D90, including total names, 0.5 d
  5. 5Final on-the-list checkScreenshot the baseline Top 20 item by item, marked as two lines, 0.5 d
  6. 6Stop listStop investing in pages asked 30+ times cumulatively and never cited
  7. 7Notarised sampleLogged out, real browser, no more than ten questions, 0.5 d
  8. 8Write the reportBefore / after + a separate brand settlement, 1.5 d
Figure: D90 is settled in this order; the main verdict comes only from stage 1, and the wider sample never enters the before-and-after comparison

Points the figure can't show:

  • The wider sample is stored in probe_d90_wide/, kept separate from probe_d90/; the two must never be added together or shown side by side on the same trend chart, and the sampling spec is labelled separately in the figure. The shelf map is internal material.
  • The stop list goes into the report with reasons given. The notarised sample is never the headline number — it only backs up reproducibility.
  • Stage 8's "before / after" writes four things: the shelf three months ago / now; which questions went from absent to present, which didn't move and why (citing which layer triage stopped at); how competitor positions changed; the off-site three-number curve. A separate page for the brand: factual-errors count, baseline X → now Y, each with a screenshot attached, plus a Try it yourself card.
  • There is also a separate internal three-state review (the difference between your pages and the control pages, before and after), done once a quarter, not part of the monthly process, and never external.

The settlement verdict

never willing

yes

something unfinished

all done

yes

no

layers 1–4 all pass

layer 3, not reachable

stopped at layer 1, 2 or 4

Will publish a checkable price, edit site?

Hold position only

Are all deliverables accepted?

Finish deliverables, push deadline

Seats exceed the band, or cited pages +3

Keep expanding: next batch of questions

Which layer did triage stop at?

Switch to hold, no gloss on the report

Wrong target, swap the whole topic

Don't conclude no effect

Figure: Settlement checks admission and deliverables first, then whether there's a determinable increase, and finally which layer triage stopped at; you may not conclude "no effect" without completing triage

The full statement of the criteria:

  • Keep expanding: the seat increase > this quarter's measured noise band (must use the same number of passes, same engine, same question pool; if the number of passes was changed for any reason this quarter, this cell is judged "not comparable" across the board and may not be used — look at pages-on-the-list instead), or on-the-list-and-cited pages increase by ≥3. Keep expanding = move to the next batch of questions and redo target-picking and the shelf snapshot.
  • Switch to hold: deliverables all passed, seat increase did not clear the noise band, on-the-list didn't increase either, and triage passed layers 1–4 all the way through. Stop adding new pages, and only do the five holding items above; write plainly in the report, "this quarter did not achieve a determinable seat increase on the main-attack questions" — no glossing over it.
  • Finish deliverables first: never offset a deliverables shortfall with a visibility number.
  • Wrong target: triage stopped at layer 3, "citation slots are all at unreachable sources" — move this whole intent out of the main-attack pool and actively change topic. Don't count a targeting mistake as "this industry can't be done".
  • Admission condition: the website must state a definite, checkable price, in the form set by the side (→ General Edition 5.2 Quick reference by side (1): identify the advertiser first; prices, promotions and freebies, result numbers's price figure); a client who's never willing to write one, or whose site can never be changed, gets hold position only. The price page is one of the few cells on an owned site that can actually get into an answer — not giving a price is the same as giving up that cell.
  • ⚠️ On the strictly regulated side, admission is judged by itemised fixed pricing: on this side, writing a "price range" is itself non-compliant. Judging a client as "not giving a price" because it has no price range would wrongly disqualify a project that is actually compliant and simply can't write its price using the general template.

The decision rests with the data, not with hope. But "no effect" must be a conclusion reached only after triage is complete and the seat change has cleared the noise band — not a side effect of the ruler's resolution being too coarse. At settlement, these three sentences may not be said to yourself:

❌ "Give it another three months and it will definitely rise" ❌ "The industry average is six months" ❌ "The data is trending up" (when the seat increase has not cleared the noise band and on-the-list hasn't increased either)

7.5 Ruler discipline: change the ruler and nothing is comparable

What you'll do in this section: pin down "what counts as changing the ruler" into a reference table, and pin down "whether the baseline should be re-frozen this month" into one mechanical test. When you're done, you'll have a way to handle a changed ruler, the rule for pinning the probe model, two ways to freeze the baseline when the shelf gets rewritten in bulk, and a list of which numbers can be printed and which stay internal only.

What counts as changing the ruler

ActionConsequenceWhat's allowed
Change engineThe ruler has changedStart a new baseline, and write plainly "not comparable with the previous one"
Change the AI model (including a minor version)The ruler has changedSame as above
Change the phrasing / question pool / ratiosThe ruler has changedSame as above
Change the number of passesSeat count scales linearlySame as above, and the old and new must not be added together or shown on the same chart
Change the web control leg's 5 questionsThe calibration ruler has changedDo not change it for the whole quarter. If it must change, change it next quarter
Change the brand six's number of questions or roundsFactual-errors count before and after is not comparableFrozen for the whole quarter

Across versions, list the numbers only — do not compute a rise or fall. When numbers from two different baselines sit on the same page, they must be printed in two separate blocks, with a line "Below: a different ruler" between them. This table and the noise-band rule in 7.2 are two scales of the same test: the noise band governs sampling variation inside the same ruler; this table governs the moments when "you've actually already switched to a different ruler".

Probe models: pinned, never drifting with your main model

The model used to run probes has a fixed version — it never drifts along with your day-to-day main model. Name the legs as in 7.1; pin the model version per the table below:

LegRule
① OpenAI API legMust connect directly to the official API. It measures ChatGPT itself — routing through any relay or proxy pool is measuring A to sign off on B
② Gemini model legMay go through a proxy pool (when the free tier's direct quota isn't enough, the whole leg comes back empty — a "high-fidelity" path that can't get through is no more honest than a proxy path that can). If you go through a pool, pin the exact model version: a floating alias returns 503, and that whole month's sample is gone
Both legsThe version only ever moves on an explicit decision, and every leg must move together. Changing the version = changing the ruler, and it makes the new report incomparable with old reports already sent out

The cost has to be written into how the report is worded: when sampling through a proxy pool, region and the model's minor version can't be independently verified — print this sentence in the methodology note.

Whether to re-freeze the baseline this month: one mechanical test

check the update date item by item

no

yes

Monthly recheck of baseline Top 20

≥6 items updated after freezing

Don't re-freeze, keep comparing in this segment

Re-freeze this month, open a new segment

Old segment stays historical only

Write old vs new X/20 as not comparable

Figure: Whether to re-freeze the baseline this month rests on one mechanical test only — never a gut feeling

Why this test is needed: put "the denominator doesn't change all quarter" (7.1) together with "the baseline isn't frozen during a window of concentrated shelf rewrites", and a gap opens up — the baseline freezes in some month, and afterwards most of the Top 20 pages get rewritten in bulk, so the denominator has already come apart from the shelf, yet discipline says you still can't change it. The test in the figure is what plugs that gap (not measured; the mechanism is clear).

Three points the figure can't show: "updated" only counts the visible update date on the URL, compared against the last freeze date; the old and new segments' X/20 must not be drawn as one single line on a chart; the monthly report for the re-freeze month must state exactly which URLs triggered it.

Two ways to freeze the baseline when the shelf has predictable rewrite windows

Some shelves have bulk-rewrite windows you can see coming: the same batch of sources gets rewritten around a particular point in time, and a rewrite = the candidate set gets swapped out (not measured; the mechanism is clear). Freeze the baseline inside the window, or let the baseline and the retest fall into two different phases, and the ruler is effectively broken for that month — any rise or fall is entirely produced by the window, not by anything you did.

Option 1 · Avoid itOption 2 · Record one per phase
Applies when: the window can be avoidedApplies when: it can't be avoided, e.g. the window covers half the year
Schedule both the baseline freeze and the retest outside the windowMeasure once before the window and once after
If a measurement point lands on the window, move it earlier or later — never move the work scheduleRecord two noise bands separately: normal-phase band / window band
Every comparison lands in the normal phaseAfter that, compare same phase to same phase only, never across phases
Figure: When you hit a bulk-rewrite window, pick one of two options and write it into the kickoff document — never choose on the day of the retest

Whichever option you choose, write the phase you're in at the time of freezing into the first line of the baseline-freeze file, and label the phase on every line of the monthly comparison; when numbers from different phases are shown side by side, state plainly "different phases, not comparable". Why this must be fixed in advance: getting this wrong doesn't throw an error — it just makes the whole month's report point the wrong way. Just like a misjudged door, the costs are asymmetric.

Example (e-commerce) During a major sale, prices, stock status, platform ranking and buyer search terms all change at once, and both seat rulers' noise bands blow out. Normal phase = outside the two weeks before a major sale and outside the two weeks after one; for any month where the retest window and the baseline window fall in different promotional phases, always write "cannot be determined (cross-promotional-phase)" — you may write neither a rise nor a fall.

Different confidence, different printing

Low-confidence work still gets done — it just has to be flagged when printed, carrying the qualifying sentence exactly as given. Which of this chapter's statements get printed, and which stay internal only:

StatementPrinted or notHow to print it
This month's measured noise bandPrintedThe value goes on page 1 of the monthly report; before the first measured value is obtained, seat counts are only listed side by side, with no verdict
Completing the five profiles benefits ChatGPT's local answers (from a public partnership)PrintedState it on the basis of "free, done in passing" — never promise an effect
Improving position is cheaper than winning a new slotPrintedCarry the qualifying sentence exactly: cross-industry pooled basis, not verified in this local industry
Fallbacks split by content formPrintedCarry "does not include Singapore, about 75% confidence"
The rate at which a review is quoted word for word into an answer (external live search, a very small sample)Not printed externallyUsed only to decide direction of work; the monthly report prints your own measured values for the M questions, split into recommendation-type / non-recommendation-type lines, and the two lines must be shown together
The share of the 90-day increase that lands on brand questionsNot printedA tabletop projection, not measured, internal reference only (the "do not say" line from holding position in 7.4)
An intent where "0 businesses are named, and someone is already fighting for the citation slot" (external live search, one-off)The specific figure is not printedUsed only for internal ranking; the report prints the named-business count and citation-slot ownership from your own baseline numbers, with the retest date and engine stated

The reviews row requires the two lines side by side because printing only the recommendation-type line would lead people towards asking for reviews, which the strictly regulated side cannot do at all. The last row carries one more red line: ❌ no material may ever say "no one is doing this" or "competition is near zero".

The percentage, sample and source behind each row are in → General Edition A.2 Evidence for picking targets, writing pages, off-site and retests; the general rule for how to flag a number and what stays for your eyes only is in → General Edition D.1 Number discipline: how to label numbers, and what stays internal.

Back to contents · GEO Playbook: General Edition

This chapter is published under a CC BY 4.0 licence · © Canlah AI. To republish or adapt it, credit “Canlah AI · GEO Playbook” and link to this page.

A condensed version for AI assistants is on GitHub, and the Markdown version of this chapter can go straight to an AI assistant. The quick guide and full-book downloads are in the downloads section. The measurements behind the numbers in this book are on the dataset page (CC BY 4.0).