# Appendix D · How to write numbers, and the glossary

> GEO Playbook · General Edition v1.0 · Canlah AI · CC BY 4.0 · Web page: https://canlah.ai/playbook/glossary/
> Markdown edition for AI assistants, same content as the web page. Figures are code blocks (wireframe / mermaid / bars / steps / split); "→" links open the matching section on the web, and the same URL with .md is its Markdown edition.

## D.1 Number discipline: how to label numbers, and what stays internal

**What you'll do in this section**: label every number you plan to publish with its source and confidence, sort material into two kinds (can go public / internal only), and never write a single month's swing up as a long-term conclusion. When you're done, you'll have a self-check list that every number goes through before it goes out.

### Which label a number gets

```split Figure: every number gets exactly one of four labels, each with its own fixed wording
Label || How to write it
External data with a source || Write the source and sample size directly
Inferred: derived from external data, not directly verified || "(inferred, XX% confidence)"; the percentage must be filled in
Self-set number: projection, post-mortem, a self-set threshold || "(self-set number, XX% confidence)" + mark item by item whether it may be said publicly
Confidence below 80%, in any form || Label it in the same sentence; never in a footnote or a separate note
```

External data with a source is written like this: "84% of AI citations come from **earned media** (news coverage alone is 27%), while **paid placements and advertorials are only 0.3%** (Muck Rack, 25M links / 17 industries)".

### Five hard rules

**Hard rule 1 · Never hide a known, material uncertainty.** Any number used to judge whether something worked must have its uncertainty printed on the same page. Hiding "the seat-count noise floor is unknown" while announcing results with seat counts = lying.

**Hard rule 2 · A single-run, single-engine, n ≤ 3 observation must never be used as a result or as acceptance evidence.** It can only serve as directional evidence of the kind "this factual error is there however many times you ask", and the sampling spec must be written on the same figure.

**Hard rule 3 · The data point below is written only here, once in the whole book, and not a word of it may be changed**; any other chapter that needs it points here and does not write it out again:

> Getting onto a list that AI already cites and ranking first on it = visibility **+16.5pp** and a position **0.8–1.8 places** earlier in the answer (Peec, 5.7M data points / nearly 200,000 answers / 8 engines; **+16.5pp is from a B2B SaaS sample; for emerging categories it is +13.4pp; not separately validated in Singapore dental, aesthetics or legal. Use it to order the work, never to promise any result**).

**Hard rule 4 · Scouting numbers ≠ your current numbers.** A number from a **generic retrieval API, with the region not locked, on a single engine, in a one-off run** can only be used to order your own work and decide where the effort goes; it must never be said publicly as your current state. **Any number calculated from it (for example, "expected N replies per batch") is equally barred from external use.** For the current list of which numbers count as scouting numbers, see （→ 通用版 A.1 机制、门与人这一层的证据）. Lifting the ban needs all three conditions at once; if even one is missing, the ban still applies: ① retested on the OpenAI leg (direct connection) + the Gemini model leg, with the target region locked, on your own frozen question pool; ② the retest date and engine are stated (including the number of passes and the region parameter); ③ what you cite externally is **your own number after the retest**.

**Hard rule 5 · Any long-term decision based on a single month's swing needs two consecutive months moving in the same direction before it is written into the discipline.** A single month's rise or fall may only trigger "does this task go on this month's schedule or not":

```mermaid Figure: a single month's swing must pass three checks, then move in the same direction for two consecutive months, before it is written into the long-term discipline; if any check cannot be answered, log it for this month only and do not schedule it
flowchart LR
  m["This month shows a swing"] --> c1{"Did the last similar swing recover?"}
  c1 -->|Cannot tell| stop["Log this month only, don't schedule"]:::warn
  c1 -->|Can tell| c2{"Crawling changed, or answer surface?"}
  c2 -->|Cannot tell| stop
  c2 -->|Can tell| c3{"Same-period total moved the other way?"}
  c3 -->|Cannot tell| stop
  c3 -->|Can tell| c4{"Same direction two months running?"}
  c4 -->|No| month["Decide this month's schedule only"]
  c4 -->|Yes| ok["Write into the long-term discipline"]:::hl
```

What each of the three checks rules out: ① it recovered last time = this is a swing, not a structural change; ② the source is still being read, it's just not making it into the answer = the answer surface narrowed, the source hasn't failed, and the work still has value; ③ the denominator changed, which makes a structural change read as an absolute decline. Log all three, item by item, on the same page. **You may not** write long-term conclusions such as "we're dropping this source / this leg is dead".

> 🔴 **The sentence "an AI answer usually has five to eight names" must never appear anywhere in this book.** It has no source, and it contradicts the rule "record exactly as many businesses as are actually named; never estimate". Use the number measured in your own baseline instead:
> "I asked K times and counted T company names in total that AI named. **That is X.X businesses per answer on average, fewest A, most B. This is what I counted on this batch of questions, not an industry figure.**"

### What material stays internal

```split Figure: one test only: if the material names a third-party organisation, it stays internal
Material || Can it be shown externally
Seat shelf map, weekly competitor seat report / competitor board in the monthly report || No, internal only
The "why AI trusts them and not you" comparison table || No, internal only
The parts of sentence-by-sentence before-and-after screenshots that show a third-party name || No; may be shown externally once names are redacted
Your own pages-on-the-list count, profile screenshots, before-and-after site comparisons || Yes (no third-party names)
Facts pages, price pages, thick pages, business profile text, outreach material || Yes; these are your advertising, so check them against your side's bans first
```

Internal only = must not be posted, must not be forwarded, must not be used for any promotion. Print the "Internal reference material" line, word for word, in the footer of this material (burned into the PDF or image itself, not written in the body of an email) and add an `-INTERNAL` suffix to the file name. For the footer wording and the three approved lines you may say publicly, see （→ 通用版 B.2 口径句与话术）. For the bans by side that apply to the two "Yes" rows of the table above, see （→ 通用版 5.2 三侧速查（一）：先判身份；价格、促销与赠送、结果数字） and （→ 通用版 5.3 三侧速查（二）：证言与评价、比较、榜单、头衔、外链、FAQ 与图注）.

**Why this matters most**: regulated industries commonly have a clause along the lines of "no comparing with or disparaging peers", and some also govern promotion that involves a third party. The moment material containing peer comparisons is put up at the front desk or sent out through a public channel, it runs into this kind of clause, and it often takes very little for a named peer to complain. This is the general pattern; to find exactly which clause applies on your side, check （→ 通用版 5.3 三侧速查（二）：证言与评价、比较、榜单、头衔、外链、FAQ 与图注）.

> **⚠️ The most common slip-up**: the first time you show someone, in person, what AI currently says about you, that screen often holds a recommendation-type answer naming a peer.
> **Only point at the screen and talk it through: no screenshot, no re-saving, nothing put into a PDF, no copy kept.**

## D.2 Glossary

**What you'll do in this section**: when you're not sure what a term means, look it up here instead of guessing from instinct. This is the book's only set of definitions; no chapter may set up its own.

The full definition of each of the following terms appears only here; every other chapter refers back to it without repeating the explanation:

| Term | Definition |
|---|---|
| **Names** | The total number of business names in the body of one AI answer. **Measured and recorded, never estimated** |
| **Seat / seat count** | The ones among these names that belong to one particular business. Your seat count = the number of times you were named across this period's N question-and-answer runs. **"Seat" and "answer-level times named" are one and the same number, not two measures.** ⚠️ **The seat count scales up linearly with the number of passes**, so it can only be compared at the same number of passes, on the same engine and with the same question pool. The number of passes is frozen for the whole quarter; to add more passes you must start a new baseline and write, on the same page, "not comparable with the previous baseline" — **the old and new lines must never be added together or plotted on the same trend chart** |
| **Listed as a source** | AI lists a given web page as a source underneath its answer. Log each occurrence; it is **never used as acceptance evidence** (during the baseline period this is mostly 0 or 1 — not enough resolution) |
| **Pages on the list X/20** | Of the source pages in the frozen baseline Top 20, the number where your name appears **on the page itself**. **The denominator 20 stays fixed for the whole quarter.** It measures off-site work, not AI behaviour |
| **On the list and cited X/20** | Of the pages on the list, those that **were actually cited this month**. **On the list ≠ cited** — AI may cite that page without picking up your line |
| **New listings** | Third-party URLs newly listed this quarter that are not in the baseline Top 20. **Not counted in pages on the list**; logged on their own line, each marked for whether it has been cited in an AI answer this quarter. One that has not been cited is logged only as work done, never as a result figure |
| **Noise band ±N seats** | The range (max − min) of the seat count when the same batch of questions is run 3 more times within the same week. **Floor fixed at ±1 seat.** A monthly change in the seat count **is only allowed to be written as "increased" when it is larger than the noise band** |
| **Model-API visibility baseline** | The seat and source distribution measured by the two model-API legs. In reports, the two legs may only be called "Model-API visibility baseline · OpenAI leg" and "Gemini model leg" — **never written as "what users see in ChatGPT" or "real user visibility"** |
| **Gemini model leg** | The leg run through the Gemini model API. **Gemini API ≠ AI Overviews ≠ AI Mode — three different products.** For AI Overviews / AI Mode visibility, **always use Search Console generative AI report impressions instead** |
| **Web control leg** | The leg where, once a month, a person manually runs a fixed set of 5 questions, 1 pass per question, on **the ChatGPT web app** (logged out or in a temporary chat, with an IP in the target market). **Calibration only, never acceptance evidence**: it does not go into the seat count, does not go into pages on the list, and is never added together with any ruler |
| **Named-set overlap / source-domain overlap** | On the same batch of 5 questions, the intersection ÷ union of the web leg's and the API leg's **named-business sets** and **source-domain sets**. **In any month where either falls below 50%, the report must print the non-extrapolation sentence unchanged** |
| **Coarse-screen top 10** | In the entry-ticket coarse screen, copy the source URLs listed under each question's answer into a table of its own, deduplicate, and take the top 10 by number of appearances: **three questions, three tables, never merged**. **n = 1, single engine, done by hand, finished within half a day**, serving only the one judgement of "is it worth doing". ⚠️ **It is not a denominator and is never used to report any number externally**; merge the three into one table and the next step, counting "reachable enough" table by table, can no longer be carried out |
| **Baseline Top 20 (= account-level Top 20)** | After the baseline has run its full rounds, combine the answers for **every intent**, and take the top 20 URLs by URL-level times cited from the raw `domains_cited`, **frozen, unchanged for the whole quarter**. This is the **only** valid denominator for "pages on the list X/20". ⚠️ **It does not judge any single intent** — a shared list always gives the same count, so using it to judge a single intent will either "rule them all out together" or "rule none of them out" |
| **Intent-level top 10** | Take the top 10 URLs by times cited, **separately for each intent, from that intent's own answers**. The method: add an intent-number column to the source summary table and group by intent to take the top 10 (the raw data already carries intent information; only how it is grouped when summarised changes). **The abstain line, check 2 of picking targets, and layer 3 of not-moved triage always read this table**, never the account-level one |
| **Intent cluster** | A group of buyer questions that are phrased differently but get essentially the same answer. **One cluster, one page**; variants go into H2s on the same page. Building a separate page for every variant only dilutes the signal |
| **Annual opportunity value** | **Price per order × gross margin × monthly capacity cap × 12.** Use it to rank buyer types, which sets the priority for the question pool and the pages |

The terms below have their full definition in their own chapter; here you get one line plus a pointer:

| Term | One line + pointer |
|---|---|
| Strictly regulated side / lightly regulated side / unregulated side | The three sides, decided by whether you need a licence or registration to open, whether price can be written as a range, and how far testimonials and peer comparisons are banned; for how to decide, see （→ 通用版 0.2 先判你属于哪一侧） |
| The six add-on switches A (agency liability)–F | Six switches that don't change the side, only add one more hard gate on top of it: A agency liability, B legally required fields, C regulated products, D spans two sides, E peer comparison banned, F referral commissions banned; for how to decide, see （→ 通用版 0.2 先判你属于哪一侧） |
| Content form (professional / consumer) | The second, independent decision after you've decided the side; it decides whether a page's weight goes on process or on tiered prices; for how to decide, see （→ 通用版 5.9 内容形态、FAQ、中文页与外语页） |
| Own sentence / market sentence / official-basis sentence | The own sentence is for the ChatGPT leg: the subject is you, with an exact number; the market sentence is for the AI Mode leg: a range + variables + a source. The strictly regulated side and switch E (peer comparison banned) each replace the market sentence with an official-basis sentence. For the full spec, see （→ 通用版 5.6 两条腿，两种句子） |
| The three labels | "Statute text", "Conservative line (not statute text)" and "Original text not obtained" — the tier of evidence behind a compliance sentence; for how to label and the stop-and-escalate rules, see （→ 通用版 0.4 合规句的三档标签与停笔规则） |
| The six page-type families | The 46 cited page types, grouped into six families by how buyers phrase their questions: the Price family, the Selection and reputation family, the Questions and rules family, the Entity and product facts family, the Primary sources family, and the Commentary family; for the full map, see （→ 通用版 5.1 页型总图：46 种、六个家族、三档证据、三侧开放） |
| Split day | The day you fix the door is the whole project's split day; every page also has its own split day, which is its launch date; for how to decide, see （→ 通用版 2.6 改门：nosnippet、robots 四步合并、WAF 白名单）, （→ 通用版 2.8 每页上线闸） |
| Baseline saved | All three are saved: the baseline Top 20, the noise band measured by the same-week retest, and the 36 brand-six runs; for the full spec, see （→ 通用版 4.4 网页对照腿与冻结基线（全书唯一完整规格）） |
| Weeks used across the book | D1–D6 are for the audit and freezing the baseline; W1 starts from the day you fix the door (the split day); the D90 settlement falls around week 13; see （→ 通用版 0.1 全书一句话与 90 天翻书顺序） |
| The four identities | How the same URL looks through four paths: a 200 in your access logs, Bing's Live Test, GSC's "crawled" status, and ChatGPT quoting it back; for how to read it, see （→ 通用版 2.3 闸二：实抓的四身份（只读））. The launch gate for every page has a separate check called "the four HTML copies match", which checks whether the source code fetched by four different user agents is the same version — that is not the same thing as the four identities |
| The four states | Already present / reachable / to ask / not reachable, judged per URL (against the coarse-screen top 10 during the coarse screen, against the intent-level top 10 after the baseline). **"To ask" is only allowed to exist at the coarse-screen stage** (counted as reachable there, with its count logged separately); at the baseline stage it must be cleared to zero; for the full criteria, see （→ 通用版 4.6 四态、两层分母与弃权线） |
| The brand six questions | A fixed 6 questions × 3 rounds × 2 engines = 36 runs / month, frozen for the whole quarter. **Changing the set of questions = changing the ruler**: the factual-error counts of the month before and the month after can no longer be compared; for the questions themselves and what to do after you find an error, see （→ 通用版 3.5 品牌六问与发现错误之后） |

```mermaid Figure: each of the three source lists answers one question only; anything that doesn't state which one is rejected
flowchart LR
  s["Want to cite a source list"] --> q{"Stated which list it is?"}
  q -->|No prefix| back["Rejected as an error"]:::warn
  q -->|Coarse-screen top 10| a["Only judges if worth doing"]
  q -->|Baseline Top 20| b["Sole denominator: pages on the list"]:::hl
  q -->|Intent-level top 10| c["Judges single intents and abstain line"]
```

**Mix them up once, and the basis for the whole quarter's verdicts is void**: the coarse-screen top 10 is an n = 1 snapshot taken in half a day; the baseline Top 20 is one shared list that cannot judge a single intent. So anywhere "Top 20" appears it must carry a prefix (baseline / account-level — the same list), and anywhere "top 10" appears it must state whether it's the coarse screen or intent-level. **Writing it bare is always rejected as an error.**