# Chapter 5 · Writing pages (extra): producing hundreds of pages, from studying a competitor to launching in batches

> GEO Playbook · General Edition v1.0 · Canlah AI · CC BY 4.0 · Web page: https://canlah.ai/playbook/at-scale/
> Markdown edition for AI assistants, same content as the web page. Figures are code blocks (wireframe / mermaid / bars / steps / split); "→" links open the matching section on the web, and the same URL with .md is its Markdown edition.

## 5.D1 The seven steps at a glance, and what to tell clients

**What you'll do in this section**: see the order of the seven steps for producing hundreds of pages in one go, understand how this chapter divides the work with the earlier writing-pages sections, and get an approved line on how far you can go when you talk to clients.

```steps Figure: the seven steps are a one-way chain: each step's output is the next step's input; skip one and the mistake gets copied into hundreds of pages
Step 0 | Pick the study target | Measure 30+ commercial queries; find the one that wins on on-site form
Step 1 | Take the architecture apart | Line by line, by archetype; a verification round tries to overturn each rule; the result is a spec + lint
Step 2 | Correct with measurement | Use citation data to revise the detail rules; where they conflict, measurement wins
Step 3 | Turn topics into a list | Five keyword-mining tracks in parallel, clustered into one main query per page
Step 4 | The three-pass pipeline | Write the draft, localise, adversarial review, each pass clears lint
Step 5 | Source audit | Trace every number in data-report pieces to its source
Step 6 | Launch in batches and retest | 30–50 pages a week, judged on day 14 / 30 / 45
```

Step 6 is followed by a monthly loop: add a new batch of keywords every month, have a person review the finished pages, then release them in batches — see 5.D10.

**How this divides the work with the earlier sections.** Which page type to pick, what each side may write, and the general rules for every page are still governed by （→ 通用版 5.1 页型总图：46 种、六个家族、三档证据、三侧开放） through （→ 通用版 5.10 页型共用件：数字六项检查、无公开价、计价单位、日期与 schema）. This chapter handles one thing only: when you need to produce dozens to hundreds of pages at once, how to make every single one follow the same rules without drifting. The first batch of landing pages still clears its gates one page at a time under （→ 通用版 5.4 三十天写页顺序与三道机械闸）; this chapter is for after that, for when you need to cover the long tail — it does not replace the ranking gate in 5.4.

**How far you can go when you talk to clients.** What this method has proven is the "production" half: in one replication project on a business's own blog, it was used to finish every page, and every page cleared its gates. The other half — "results", meaning rankings and AI citations after launch — has no retest data yet. So tell the client only "we produce using this method, and keep-or-cut is judged by the day 14 / 30 / 45 retests", never "this method has already been proven to earn citations".

> **Example** (a GEO agency's own blog) Using these seven steps, it finished 323 articles × en / zh / zh-tw (293 new articles + 30 rewrites), all of them passing the architecture gate.

## 5.D2 Pick the right study target: the competitor that wins on on-site form

**What you'll do in this section**: use three checks — measured rankings, domain age and backlinks — to filter down to one competitor whose "way of winning is copyable", write down which layer of keywords it wins on, which layer it loses on, and whether the page type it wins on is one your side can build. When you're done, you'll have a study target and one sentence stating which part of it, and only that part, you will copy.

```mermaid Figure: a young domain with near-zero backlinks still taking first place shows it wins on on-site form — only this kind is worth studying, and only study the layer it wins
flowchart LR
  q1{"Measured in the top 10?"} -->|No| no1["Do not study it"]
  q1 -->|Yes| q2{"Young domain and near-zero backlinks?"}
  q2 -->|No| no2["Wins on domain/backlinks, not copyable"]:::warn
  q2 -->|Yes| yes["Wins on on-site form, copyable"]:::hl
  yes --> scope["Write down which layer it wins, which it loses"]
  scope --> q3{"Is the page type it wins on open to you?"}
  q3 -->|Not open| swap["Study the skeleton, swap page type by side"]
  q3 -->|Open| go["Next step: take the architecture apart"]
```

**How to check the three items.** For rankings, use a SERP API such as Serper to measure 30 or more commercial queries, set `gl` to your target market, and see who is in the top 10; for domain age, check whois; for backlinks, look at third-party mentions, directory listings and inclusion in best-of lists. The criterion is not "whose content is better" but "who wins on your queries, and whose way of winning is copyable".

**Why only "young domain + near-zero backlinks + first place" is worth studying.** Domain age and backlinks are the two things you cannot copy in the short term. Once you rule those two out, if it still takes first place, the only thing left that can explain the ranking is on-site form — architecture, title, first sentence, tables — and these can be taken apart line by line. The reverse is also true: an old rival that wins on 20 years of domain age and backlinks will not reach its ranking even if you copy its writing — do not study it.

**See clearly where it loses.** A study target does not necessarily win on every layer of keywords. If it only wins on ultra-long-tail queries stacked with multiple modifiers, and never makes the top 10 on head terms, then what you are copying is a "long-tail matrix play", not "the ability to win head terms". This sentence decides which layer of keywords topic selection (5.D5) covers.

**Check availability first for the page type it wins on.** For the page type the study target wins its rankings on, check the availability column in 5.1 for your side first. If it is not open to you, only study its skeleton and where its sentences sit, and swap the page type to the one your side can build under the same intent, per （→ 通用版 0.3 哪些页型对你有限制、改做哪一型）.

> **Example** (a GEO agency) The study target's domain was 6 months old with backlinks ≈ 0, yet it took 5 #1 rankings, and Google AI Overview copied its first sentence word for word. But it only won ultra-long-tail queries stacked with 3 or more modifiers (best + industry + location + year); it did not make the top 10 on a single head term, and every page that took a ranking was a list page — exactly the type the strictly regulated side can use least. Evidence: see Appendix A.4.

**A hard rule: the study target is never named.** The study target is only used to study form; its name never goes into any article, list or comparison page, and no "us vs. that company" page gets built either; comparisons against a tool category or an approach are fine. This rule goes into a hard lint block: a banned-name list, plus a ban on any `<our-brand>-vs-*` slug.

## 5.D3 Take the architecture apart until you can write from it

**What you'll do in this section**: take every article of the study target apart into rules, line by line and grouped by archetype, then send a verification round to check those rules against the original articles and try to overturn them. When you're done, you'll have two files — a writing spec and a lint gate — neither one optional.

```mermaid Figure: every rule taken apart goes through the verification round, and only the rules that hold up enter the spec; whatever in the spec can be machine-checked goes into lint, whatever cannot goes to review
flowchart LR
  src["Full text: find llms-full.txt first"] --> arch["Vertical: line by line, per archetype"]
  src --> cross["Horizontal: across every article"]
  arch --> v{"Verification round against the original"}
  cross --> v
  v -->|Overturned| drop["One batch's habit, kept out of the spec"]:::warn
  v -->|Holds up| spec["Writing spec"]
  spec -->|Machine-checkable| lint["Lint gate"]:::hl
  spec -->|Not machine-checkable| rev["Adversarial review, article by article"]
```

**Find `llms-full.txt` first.** A competitor writing for AI citation often exposes the full Markdown of every article in this one file, so you get everything in one go without crawling page by page.

**What each teardown covers:**

| Who takes it apart | Output |
|---|---|
| Vertical: group by archetype, one agent per group reads every article line by line | title formula, H2 skeleton, opening sentence, paragraph rhythm, table usage, FAQ, Method block, where the business mentions itself, verbatim boilerplate |
| Horizontal: a separate agent looks across every article | the title and slug system, recurring fixed blocks, Chinese/English localisation differences, the technical layer (schema / llms / hreflang / publishing cadence) |
| Verification round: 3 agents | check against the original articles, with the sole aim of overturning the rules the first two teardowns produced |

**Why the verification round cannot be skipped.** The person doing the teardown sees the habits of one batch of articles, and a competitor's writing can change from batch to batch. Skip verification, and "one batch's habit" gets written up as "a global iron rule", then copied into hundreds of pages by the pipeline.

> **Example** (a GEO agency) The verification round corrected 9 rules; among them, Oxford commas, abbreviations, summary length and bold usage genuinely differed from batch to batch in the study target.

**What each of the two files is for.** The spec is what people and agents write from; its content is "archetype × title formula × body skeleton × shared blocks × style rules". Lint is what the machine enforces: everything in the spec that can be machine-checked is written as a check item, and an article that has not cleared lint does not count as done. At a scale of hundreds of pages, no one can check formatting article by article, so nothing that can be handed to a machine is left for a person.

**Lint must compare titles by word root, not by substring.** "agencies" and "agency", "tracking" and "track" count as the same word. Compare by substring, and writers trying to avoid a repeat will twist the title into something awkward (e.g. "Agency and Service Providers").

## 5.D4 Correct with measurement: the competitor sets the skeleton, measurement sets the details

**What you'll do in this section**: take the measured citation data and go through every detail rule in the spec, change what conflicts with measurement, and add a qualifier to two of them. When you're done, every detail rule in your spec is clearly marked as either "the competitor's habit" or "what measurement showed".

```split Figure: each of the seven measured findings changes one detail rule in the spec; the skeleton follows the competitor, but when a detail conflicts, measurement wins
What measurement showed || What the spec changed to
Copied sentences are almost always the first sentence of a paragraph || Every H2's first sentence must stand alone as an answer, carrying a number
About 73% of copied facts sit in the first 30% of the page || Lint checks the first 30% has at least 8 numbers (an empirical figure)
Tables get copied header row and all || Table headers carry a verdict word: Best for, Winner, etc.
ChatGPT looks for the primary source, AI Mode for second-hand pages || Every page carries two sentences: your own fixed price + the market range (unregulated side only)
Price-comparison and vs. pages are cited most; the competitor has neither || Add two archetypes: price comparison and comparison
Cited pages ran 62–9,000 words || Word count is only a hint; only list pages keep a floor
When the summary number and the table disagree, AI uses the table || The summary number must equal the table
```

**Why measurement wins when there's a conflict.** The competitor's writing only tells you "this is how it happens to write", not "this is the writing AI most likes to cite". It ranks because its whole form works together; pull out any one detail rule on its own and that rule may not be why it gets cited. The correction draws on teardowns of cited pages across several industries and both AI legs; evidence: see Appendix A.3.

**Qualifier 1: "two sentences per page" only holds on the unregulated side.** The two legs look for different pages, so a page needs two kinds of sentence (see （→ 通用版 5.6 两条腿，两种句子）); but whether the "market range" sentence can be written on a regulated side, and what to swap it for, follows the price row in （→ 通用版 5.2 三侧速查（一）：先判身份；价格、促销与赠送、结果数字） and the comparison row in （→ 通用版 5.3 三侧速查（二）：证言与评价、比较、榜单、头衔、外链、FAQ 与图注） — check your own column, then check your industry edition. Table headers carrying a verdict word such as "Winner" or "market range" should likewise be checked against 5.3 first on a regulated side.

**Qualifier 2: "at least 8 numbers in the first 30%" is an empirical figure.** It is based on that ~73%: of the 75 copied sentences that could be located, about 55 fell within the first 30% of the body. The denominator is the 75 that could be located out of 526 total citations, not all 526; the positions were estimated by eye, so treat them only as a strong trend. The "8" is lint's threshold, not a number computed from this data set. Do not tell a client "put 8 numbers in the first 30% and you'll get cited".

**The rest do not need a separate rulebook.** First sentences of paragraphs, table column headers, and word count set by page type: the book already has general rules for these in （→ 通用版 5.7 所有页型都成立的七条） and （→ 通用版 5.8 结论块、密度闸、出处闸）, and the spec simply cites them. The two new archetypes, price comparison and comparison, follow the blueprints for the Price family page types and for （→ 通用版 5.15 对比页（pt04）） respectively; whether the comparison page is open on your side also needs checking against 5.1 first.

## 5.D5 Topic selection: turn every head and long-tail keyword into a list with one main query per page

**What you'll do in this section**: run five keyword-mining tracks in parallel, cluster the results into "one main query per page", then run three passes to catch what's missing. When you're done, you'll have an article list: one main query per row, a group of near-synonym secondary keywords, and one archetype.

```mermaid Figure: the five keyword-mining tracks cannot see each other and each mines on its own; after pooling and clustering, three more passes fill the gaps before it becomes the list
flowchart LR
  s1["SERP expansion + autocomplete"] --> pool["Keyword pool"]
  s2["Reverse-engineer competitor sitemaps"] --> pool
  s3["Own GSC: queries with impressions"] --> pool
  s4["Stacked matrix + competition spot-check"] --> pool
  s5["Competitor names, how-to, definition topics"] --> pool
  pool -->|Cluster| cl["One main query per page"]
  cl --> fix["Review agent: fill gaps, merge dupes, fix titles"]
  fix --> human["Person checks each competitor title"]:::hl
  human --> vs["Re-sweep price-comparison and vs. terms"]
  vs --> list["Article list"]
```

**How each of the five tracks mines:**

| Track | How it mines | How it filters |
|---|---|---|
| SERP expansion | Seed keywords × Serper's relatedSearches + PAA; Google autocomplete (seed + a–z + market name / best / vs / cost / tools) | — |
| Competitor slugs | Crawl the sitemaps of competitors and tool vendors, and reverse-engineer from each blog slug the keyword it is targeting | — |
| Own GSC | Queries with impressions already recorded — the most genuine signal of demand | — |
| Stacked matrix | [best/top] × category × industry × market × year × optional engine | Keep only combinations a real person would search; sample 60+ to test competitiveness (green / amber / red) |
| Name and definition topics | `{brand} review / alternatives / pricing / vs`, how-to, definition topics | Use autocomplete to drop terms nobody searches for |

**The two hard rules of clustering.** Near-synonyms go into secondary; one intent gets only one page. This is the same rule as （→ 通用版 5.5 页面单位：一个意图簇一页）, and it matters even more at scale — otherwise, across hundreds of pages, you end up competing with yourself for the same keyword.

**Why the human title check cannot be skipped.** Clustering and deduplication are both done by an agent, and a wrongly deleted item throws no error. Having a person go through the competitor's title list line by line and ask "do we have a matching page?" is the last chance to catch a missing page.

> **Example** (a GEO agency) 1,476 keywords ended up as 293 new articles + 30 rewrites. During deduplication, an article that matched the study target 1:1 was wrongly deleted; the human title check recovered 9 articles in total.

**Why price-comparison and vs. keywords need a separate sweep.** The clustering agent tends to file them as "opinion pieces", when they should actually be price-comparison pages or comparison pages. Get the archetype wrong, and the three downstream pipeline passes will write the whole page on the wrong skeleton.

**Check every row's archetype against 5.1's availability column for your side before you start writing.** In the stacked matrix, combinations starting with best/top map to list pages, and vs. keywords map to comparison pages; both types have restrictions on regulated sides.

## 5.D6 The production pipeline: three agent passes per article, plus automated lint gates

**What you'll do in this section**: give every article a "write → localise → adversarial review" pipeline, with lint attached after every pass, and use the strongest model for review. When you're done, you'll have a production line that can run dozens of articles in one go and resume after an interruption.

```mermaid Figure: every article runs through the three passes independently, each pass followed by a lint stop; real problems are caught almost entirely by the third pass, so it must not be downgraded
flowchart LR
  card["Topic card"] --> en["Write EN"]
  en --> g1{"EN lint all clear?"}
  g1 -->|Fails| en
  g1 -->|Clears| loc["Localise into zh + zh-tw"]
  loc --> g2{"Lint all clear in all three languages?"}
  g2 -->|Fails| loc
  g2 -->|Clears| rev["Adversarial review, fixed on the spot"]:::hl
  rev --> g3{"Lint still all clear?"}
  g3 -->|Fails| rev
  g3 -->|Clears| nxt["Into source audit and launch batch"]
  rev -->|If downgraded| miss["Real problems go uncaught"]:::warn
```

**What each pass reads, and what model it uses:**

| Pass | What it does | Model | Gate |
|---|---|---|---|
| Write EN | First read the spec, the template for the same archetype, **your own reference article for the same archetype** (one already-live article assigned per archetype) and the measurement report, then research public sources and write | Sonnet | The EN part of lint fully clears |
| Localise | Write zh + zh-tw, adapting rather than translating literally, sections and numbers matched one to one | Sonnet | Lint fully clears (all three languages) |
| Adversarial review | Check the 11 items lint cannot check, fix them on the spot | **Opus, never downgraded** | Lint still fully clears |

**Why writing does not read the competitor's original text.** First, the pipeline needs to run in the cloud and cannot depend on a local archive of the competitor's pages; second, reading the competitor's original text lets its sentences leak into the draft. The study target's form has already been captured in the spec and lint, so writing only needs to look at your own reference article for the same archetype.

The 11 review items include: the subject and qualifiers of the first sentence, the business's own limitations, competitor wording, absolute dates, table notes, FAQ, a closing line telling readers to verify for themselves, the first sentence of each paragraph being the answer, and summary = table.

**Why review cannot be downgraded.** Every rule the spec cannot lint-check is pushed onto this one pass; lint can only check format, never facts or logic. Almost every real problem in the pipeline is caught by review.

> **Example** (a GEO agency) Problems review caught: a press release's first sentence attributed an award to the company itself when a different entity had actually won it, contradicting the page's own table; a summary number that did not match the table; a price with no currency stated; a number cited with no source; and the competitor described as "does not support a feature", which should have read "not found on the public page".

That last kind of mistake steps straight onto the red line on lying: not found is not the same as does not exist. Review corrects it to the standard set in （→ 通用版 5.31 交稿：每页五步、撒谎红线与自检）.

**On a regulated side, review has one more job.** First decide the side, then go through the banned-word and required-field checks for your column under 5.2–5.3, item by item. When you hit one of the three stop-and-escalate situations in （→ 通用版 0.4 合规句的三档标签与停笔规则）, stop and escalate to the client's compliance officer for written sign-off; carry each compliance sentence's label over exactly as the manual gives it, and never let the agent fill in a label or upgrade one on its own. **A written sign-off resolves only the three stop-and-escalate situations; it never turns something banned as Statute text into something publishable.**

**How it runs.**
- Every article runs through the three steps independently, none waiting on another: use a pipeline, no batch barriers, never make a fast article wait for a slow one.
- 45 articles per batch, 2–4 pipelines running in parallel.
- Support skipping by stage (e.g. `stages: ['zh','review']`): after an interruption, only the missing steps are run; steps that already cleared their gates are not redone.

## 5.D7 Six traps we hit: build the safeguards in before you start

**What you'll do in this section**: before you start the run, install one safeguard for each of six traps. When you're done, your pipeline configuration has six new settings: model split, writing in segments, GSC cross-check, slug freeze, source tracing and batching.

```split Figure: each of the six traps has a safeguard you can install before you start; fixing it after something goes wrong always costs more
Trap || Install before you start
Running the strongest model everywhere in parallel blows through the weekly quota || Use Sonnet for writing and localising — saves about two-thirds of the calls
About 5% of Chinese localisation gets blocked by content filters || Switch to writing in segments; rewrite whatever gets blocked
A new page targets the same keyword as an already-ranking page || Cross-check against GSC before launch; a new page must not compete with an old page already ranking for that keyword
Rewriting an already-indexed old page || Keep the slug unchanged; a keyword that already has impressions must not be lost
Raw results of an agent's live measurement never saved to disk || Trace every number back to its source before launch (5.D8)
Hundreds of pages on the same template all go live at once || Launch in batches; make every page's list, market and evidence different
```

**Quota can only be saved on the first two passes.** Real problems are caught by review (5.D6), so only writing and localising get downgraded; review keeps the strongest model.

**Where the impression keywords stay when rewriting an old page.** A keyword that already has impressions in GSC must remain in one of: title, first sentence, an H2 or the FAQ. The rest of the constraints on rewriting an old page are in 5.5.

**Why saving to disk gets missed.** When an agent writes a data-report piece it measures things live — crawling, calling an API — but the raw result is left only in the chat transcript. A number you publish must be reproducible, so this kind of draft must clear 5.D8.

**Why every page needs to differ.** With hundreds of pages on the same template, the risk in content at scale is exactly that the pages look alike. The batching cadence is in 5.D9.

## 5.D8 Source audit: trace every number in data reports to its source

**What you'll do in this section**: give every data-report piece and press release two agents (one traces every number to its source, the other re-checks a sample and is suspicious by default), then give the whole piece one of three verdicts: publish as is / fix then publish / hold back. When you're done, you'll have one more step that every such piece must clear before publication, so a number an agent measured live but never saved to disk cannot go straight out.

```mermaid Figure: every number is first tagged with one of six source classes, then a second agent, suspicious by default, re-tests 5 or more; only after both steps are done does the whole piece get a verdict
flowchart LR
  doc["Data report, press release"] --> audit["Audit: tag every number's source"]
  audit --> tag{"One of six source classes"}
  tag -->|Measured live| raw["Raw result never saved to disk"]:::warn
  tag -->|No traceable source| none["Find a checkable source, or delete"]:::warn
  tag -->|Other four classes| kept["Record where the source is"]
  raw --> chk["Re-check: re-test a sample of 5+"]:::hl
  none --> chk
  kept --> chk
  chk --> v{"Verdict for the whole piece"}
  v -->|Holds up| ok["Publish as is"]
  v -->|Needs a fix| fix["Fix then publish"]
  v -->|Cannot be fixed| hold["Hold back"]
```

**The six source classes.** From the writing record, the audit agent files every number on the page into one of: third-party citation, measured live by the agent, own evidence base, official website page, calculated, no source.

**Why this is a separate step.** Adversarial review can catch the odd sourceless number (5.D6), but it checks writing article by article, not every number's source; and the raw result of an agent's live measurement is left only in the chat transcript (5.D7, trap five). Data reports and press releases have many numbers from mixed sources, so for this kind of piece the step is mandatory.

**Why the re-check is suspicious by default.** The re-check agent does not trust the audit agent's tags: it takes a sample of 5 or more, re-measures each one or re-opens its source, and looks for itself.

**What to do with a sourceless number.** Follow the red line on lying in （→ 通用版 5.31 交稿：每页五步、撒谎红线与自检）: if a checkable source can be found, add it; if not, delete the number. Never leave an unchecked number in, and never give one a source no one has actually opened and read.

## 5.D9 Launch in batches: from launch day to the keep-or-cut decision on day 45

**What you'll do in this section**: decide how many go out per batch and which types launch first, complete two tasks on launch day, then retest at three checkpoints — day 14, 30 and 45 — to decide keep or cut. When you're done, you'll have a timetable running from launch to verdict, so hundreds of pages are never released all at once, and never released with no one watching afterward.

```steps Figure: from launch to the keep-or-cut verdict, a batch of pages passes through only five action points; still outside the top 30 on day 45 means cut losses
Every week | Launch a batch | 30–50 pages; list, price-comparison and comparison pages first
D0 | Launch day | GSC URL check + IndexNow
D14 | First retest | Serper: rankings on the first 3 result pages + whether AI cites it
D30 | Second retest | Measure again with the same ruler
D45 | Keep-or-cut | Still not in the top 30: merge or take down
```

**Check your side first before deciding which types launch first.** Comparison pages and price-comparison pages are the commercial page types cited most in measurement (5.D4), and list pages match the best/top combinations in the stacked matrix (5.D5). All three types have restrictions on the regulated sides: check the availability column in 5.1 for your side first, and for any type that is not open to you, put the type your side can build for the same intent into the first batch instead.

**Why the two launch-day tasks cannot be skipped.** They are not a retest — they only make sure the page can be indexed. Skip them, and a later retest that finds "not indexed" gets misread as "not cited".

**Both retests use the same ruler.** Day 14 and day 30 use the same tool and the same set of queries, and both check the first 3 pages of results and whether AI Mode / AIO cites the page; switch mid-way and the two sets of numbers are no longer comparable (（→ 通用版 7.5 量具纪律：换尺子就不可比）). This ruler measures a new page's organic ranking and whether AI cites it, and it is not the same ruler as the monthly one in （→ 通用版 7.1 复测的产出与量具） — do not merge the two sets of numbers, and do not compare them to each other.

**Where a merge goes.** A page still not in the top 30 on day 45 is merged into the one page for its intent cluster, following the rule in （→ 通用版 5.5 页面单位：一个意图簇一页）; if there is nothing to merge it into, take it down.

**Two protections for syncing at launch.**
- An already-live article that a person has edited directly is never overwritten by the next sync: before syncing, compare against the content version recorded in the last batch, and if they do not match, keep the person's edited version.
- A link pointing to an article not yet live is removed first and added back automatically once the target goes live; every article keeps at least 2 related-reading links.

**Judging results is not just about rankings.** Success is judged by whether the AI's answer cites you and whether leads come in. Both of these only have data once the retest checkpoints arrive — this chapter stops at "production", as stated in 5.D1.

## 5.D10 Add new keywords every month: new keywords come in, a person reviews before anything goes out

**What you'll do in this section**: turn "find new keywords → write → launch" into a loop that runs itself once a month, but keep one human review step right before "launch". When you're done, you'll have a task that opens a PR automatically on the 1st of every month, so new keywords never pile up unattended and never go out unreviewed.

```steps Figure: one round a month — the machine finds keywords and writes, the person only reviews; only what passes review enters the weekly launch batch
1st of the month | Incremental keyword mining | Related searches + PAA + autocomplete, deduplicated by word root
Same day | Pick 5–6 articles | Drop keywords already covered, one main query per page
Next | The same pipeline | Write the draft, localise, adversarial review, all three languages clear lint
Open a PR | Wait for human review | Not auto-merged
After review | Launch in batches | Handed to the weekly launch task
```

**Why it is not auto-merged.** Finding keywords and writing can both be handed to a machine, but "is this keyword worth doing, should this page go live" needs a person to look — get it wrong once across hundreds of pages and the mistake gets copied across the whole batch. The monthly volume is capped at 5–6 articles, a load a person can actually review.

**The cost of incremental keyword mining.** One round is roughly twenty-odd search-API calls: related searches, "people also ask" and autocomplete each run once, deduplicated by word root, then keywords already covered by an existing page are dropped.