Every sentence in the main text that says "see Appendix A for evidence" ends up here: a one-line conclusion / a number / the sample and timing / the source / the confidence level and whether it can be said publicly, one row per item. This appendix does not repeat the main text's rules; it only gives the row of data that sits behind each rule.
A.1 Evidence for mechanics, the door and identity
What you'll do in this section: check the original basis for any number in the mechanics, door or identity layers — sample size, source, confidence, and whether it can be said publicly.
searchVIU measured the live fetches of five systems — ChatGPT, Claude, Perplexity, Gemini and AI Mode: not one of them reads JSON-LD, and hidden Microdata or RDFa are ignored the same way, even when that piece of information exists nowhere else on the page. Measured, may be said publicly; it can only be used to say "facts must be written out, one field per line, in visible HTML" — it does not follow that "schema is useless".
Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 control pages (2025-08 to 2026-03). After adding JSON-LD, AI Overviews (AIO) citations changed −4.6% — the only significant result, and negative; the other two platforms showed no significant effect. The only conclusion is that "schema needs one half-day job to seal it once, and it does not go on every page's checklist"; it does not follow that "schema is useless": it works through Google's Knowledge Graph, an indirect path.
Whitespark, 540 queries / 3 US cities / 6 industries (including dentists, medical clinics, personal injury lawyers). Can be used to say "cost questions are the question type with the highest AI coverage, not just one of them"; US sample, must not be written as a Singapore number.
Steady Demand, 1,487 queries / 50 metros / 14,472 citations. May be said publicly, but the study states outright that it did not separately quantify the citation contribution of business hours, the service list, photos, Q&A or posts — the profile is an entry ticket, not a lever, and must not be used in place of other actions.
A Singapore brand's D2C category page, raw-fetched with a retrieval-crawler UA. Two more numbers from the same page: 460,655 bytes of HTML, only 2,182 characters of visible body text left after stripping tags (the original test record gives body text as five parts in ten thousand; bytes and characters are different units, not recalculated here). Measured, 1 case on a single page — directional only.
| Conclusion | Number | Sample and time | Source | Confidence · can it be said publicly |
|---|---|---|---|---|
| Mainstream AI retrieval crawlers execute zero JS; the only exceptions are Gemini and Applebot | 500M+ fetches sampled, no evidence of JS execution; none of GPTBot/OAI-SearchBot/ChatGPT-User/ClaudeBot/Claude-User/PerplexityBot/Bytespider/Meta-ExternalAgent execute JS | 2024-12 | Vercel × MERJ | Measured, may be said publicly; "69% of AI crawlers don't execute JS" is a mis-cited second-hand figure — do not use it |
Taking llms.txt out of the prediction model made citation-frequency prediction more accurate, not less | No specific improvement figure given | Sample size not stated | SE Ranking (XGBoost model); Google's 2026-05-15 AI optimisation guidance states outright that it isn't needed | Backed by both an official document and a third-party study, may be said publicly; "this is the biggest fake GEO lever of 2026" is this book's own judgement, not the study's wording |
| Video view count / likes / subscriber count are almost uncorrelated with citation frequency | Correlation coefficient about −0.03 | 1.7M data points | Otterly | Measured, may be said publicly (the full set of video evidence is in A.2) |
| GPTBot is a training crawler, not a retrieval crawler; blocking it does not affect whether you get cited | 88.2% of sites that block GPTBot are cited anyway | — | BuzzStream (other sources in A.5) | Measured, may be said publicly |
| Applebot-Extended is only a training opt-out switch and does not affect the Applebot retrieval leg; all three Anthropic bots obey robots.txt | — | — | Apple's About Applebot; Anthropic's official documentation | Official documentation, may be said publicly |
| At least eight retrieval sources; local answers mainly go through Foursquare + Yelp + Labrador-local, medical questions through Labrador-medical | OpenAI's own index, Labrador (split into vertical sub-indexes such as general / local / medical / legal / news / shopping), Bright Data, Oxylabs, a third SERP-fetching path, Yelp, TripAdvisor, Microsoft Web IQ, and two internal pipelines; OpenAI and Foursquare publicly announced a partnership in 2024-12, 100M+ POIs / 200+ countries | 2026-05-21 to 07-21, captured through ChatGPT's own SSE result_source field | Peec | Measured, may be said publicly; IndexNow only feeds Bing and Yandex, it is not a switch for the ChatGPT leg |
| OpenAI has signed a deal with Yelp, and Apple Business has broad coverage — claiming the five business profiles adds two more trustworthy sources | OpenAI signed with Yelp on 2026-07-23; Apple Business covers 200+ countries including SG | — | Compiled internally (no URL given in the original) | 70% confidence, no URL given in the original; may be said publicly, but only in the "free, done in passing" wording, never promising a result |
| Buyers rarely click the links in an AI summary; once they have a name, they search for it separately | About 1% of visits click through | 2025, 900 US users | Pew | US sample, may be said publicly but must carry a qualifier |
| Most AI citations come from earned media; news coverage alone accounts for 27%; paid placements and advertorials are only 0.3% | 84% / 27% / 0.3% | 25M links, 17 industries | Muck Rack | Cross-industry aggregate, may be said publicly with a qualifier (not separately validated for the three Singapore industries); the correct reading is "84% is other people writing about you on their own initiative, not you buying a placement" — it must not be read as "84% comes from third-party sites" and then used to put effort into directories, paid listings and listicle submissions, which is exactly that 0.3% bucket |
| Even when the buyer searches a brand name, owned content still accounts for a low share | 23% | Sample basis not detailed in the original | Omniscient | May be said publicly; the denominator differs from the 84% in the row above — the two numbers must not be subtracted or compared side by side |
| Page-type distribution is an industry variable, not a general rule (example: healthcare) | Across all industries: articles 23.7%, listicles 19.6%, comparison pages 1.87 citations per search (top of the whole table), category 5.2%, profile 3.8%; healthcare/medical breakdown: articles 54%, listicles 0%, category 5.2% | 25,337 citations | deltaV | Measured, may be said publicly; the healthcare breakdown is an example only and must not be applied to other industries |
| Scouting numbers can only order the work; they must not be said publicly as the current state | Recommendation-type / non-recommendation-type review-copying frequency 67%/0% (n=15); second-opinion type: 0 businesses named at answer level, 6 in citation slots; the per-tier competitive conclusions for Chinese and Indonesian; the item on Invisalign's official tiers replacing reviews; the off-site page-type share group | Generic search API, region not locked, single engine, one-off | Internal measurement, 2026-09-21 etc. | Lifting this restriction requires all three conditions at once: ① already retested on your own frozen question pool using the OpenAI direct leg + Gemini model leg, region locked to SG ② the retest date and engine are stated ③ what you cite publicly is the retested number; any number derived from these is equally not for external use (general rule in → General Edition D.1 Number discipline: how to label numbers, and what stays internal) |
Google's own wording confirms the AI Overviews gate is Googlebot + nosnippet, not Google-Extended | "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." "a page must be indexed and eligible to be shown in Google Search with a snippet" "robots.txt directives for Googlebot is the control" | Read on 2026-09-21 | Google's crawler documentation + AI features and your website | Official documentation, verbatim; may be said publicly |
| Google's official documentation lists Merchant Center and Business Profile in its AI Overviews / AI Mode best-practices checklist, and states outright that no extra schema is needed | "Checking that your Merchant Center and Business Profile information is up-to-date" is on the checklist; "You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add." | Read on 2026-09-21 | Google, AI features and your website | Verified (official wording), may be said publicly; the line "AI Mode / AI Overviews / Gemini's product answers all draw from the same Shopping Graph" is an internal-pipeline detail that has not been verified — use it only to order the work, never as Google's official position |
| Wikipedia is the single largest node in ChatGPT's citation set, and Wikidata is its structured anchor | Top-10 source share about 47.9%; 7.8% of total citations out of 680M citations; Ahrefs 2026-07 tracked 8.9% | Profound (680M citations); Ahrefs 2026-07 | Profound; Ahrefs | Measured, may be said publicly; Wikidata's admission criteria require one of three conditions, and the relevant one requires that the subject "can be described using serious and publicly available references"; entry is not automatic |
| Pure directory listing is the page type with the lowest citation density in the measurements; the single biggest effect is the entity-strength multiplier at the person layer | category 5.2% + profile 3.8% (in healthcare: listicles 0%, articles 54%); the single biggest effect is an 8x entity-strength multiplier (Wikidata record, business-media coverage, structured disclosures); independent businesses with only Yelp reviews and a Google profile appear only unreliably, and 78% of independent local businesses have an AI citation share of roughly zero | deltaV, 25,337 citations; 5WPR, 320+ prompts / 10 metros / 5 engines | deltaV; 5WPR | Measured, may be said publicly; a single third-party study, US sample, not separately validated for the three Singapore industries — when cited, the qualifier must be written in the same sentence, and it must not be used to promise any result or amount |
| On the regulated sides, almost all persuasion is banned at the organisation layer, but not at the person layer (example: healthcare) | The regulator explicitly allows practitioners to publish their qualifications, areas of practice, practice arrangements and contact details | — | The industry regulator's own rule text (rule text and numbering in the industry edition's appendix) | The original record says "explicitly allowed." Both internally and externally, be clear that "allowed ≠ unregulated": the industry's nine advertising standards (factual / accurate / verifiable / not extravagant / not misleading / not sensational / not persuasive / not comparative / not disparaging) still apply to personal profiles; laudatory terms have their own separate ban list from the regulator |
| Identity fixes are the only lever where "what AI says often changes within 7–14 days of the fix"; writing pages, off-site and video all pay back on a scale of months | 7–14 days | Internal tabletop observation / internal review | Internal | 70% confidence, not measured, never state the scale publicly |
| The one recommendation-type question that didn't copy from reviews had its answer body copy from the manufacturer's official tier + the regulator's specialist register instead (example: dental) | 4 of 6 recommendation-type questions copied reviews and star ratings; the one that didn't copy reviews copied the manufacturer's official tier (Black Diamond / Platinum Elite) + the regulator's specialist register | n=15, single engine, 2026-09-21 | Internal measurement | 85% confidence; before a region-locked retest, use for internal ordering only, not in external messaging; this is a mechanism observation, not compliance clearance; on the strictly regulated side, do not put manufacturer tiers on your own pages (the conservative line) |
Industry rule-text sources: where the main text marks something "Statute text" and the rule belongs to a particular industry's regulator (for example, the rule that disclosing payment does not exempt you from liability, for list pages whose title is itself laudatory, or the four items explicitly allowed in a practitioner's personal profile), the rule text, its number and its URL are kept in that industry edition's appendix, not in the General Edition; the main text always writes "rule text in the industry edition's appendix", never pointing to this appendix.
A.2 Evidence for picking targets, writing pages, off-site and retests
What you'll do in this section: check the original basis for the ratios used in picking targets, the web control leg, the basis for annual opportunity value, where off-site effort goes, the video correlation coefficient, the noise band and the confidence list.
PRWeb, 1,120 queries / 7 industries / 13 cities. Question-pool wording can produce a fake zero visibility; the question-pool ratios are therefore fixed: with a location ≥60%, price words ≥30%, bare category words ≤10%. Measured, may be said publicly.
Searched on 2026-09-21, region not locked to SG, single engine, one-off, 80–85% confidence. This is for internal ordering; before it goes into external material, retest it region-locked to SG on your own frozen question pool; before launch, spot-check 3 questions from your own pool. The bar is not whether you're on a list — it's whether you have a Chinese-language landing page.
Internal figure, 80% confidence. Probe count: running the full set from the start is 30 × 5 × 2 = 300; two stages is 60 + (12–18) × 5 × 2 = 180–240, plus one round of same-week retesting. The cost: dead-pile questions only ever get an n=2 conclusion, which can never be reported externally; a question wrongly knocked out gets no second chance for the whole quarter. The dead-pile criterion is an operating convention, not a measured conclusion — what backs it is the direction seen across our own 200 tests (educational-type questions almost never name a business), 70% confidence; the red line for external wording is that you may only write "no business named in the coarse screen", never "zero visibility".
Searched on 2026-09-21, n=15, single engine, 85% confidence. Before a region-locked SG retest, use this only for internal ordering and to decide the direction of the work — not as a promise and not in external messaging; how this prints in the monthly report is in the confidence list at the end of this section.
D1–D30: why each step moves rankings
| Time | Why it moves rankings (with sources) |
|---|---|
| D1 morning · the door's five technical gates | A closed door = everything after it is multiplied by zero. A ×0/1 switch, cause and effect certain, zero cost; gate 0 comes first because it's the only direct evidence, and because access permissions take the longest to wait for |
| D1 afternoon · the work order | The three inputs — side, acceptance tier, and annual-opportunity-value ranking — are already in hand; this is filling in a form, not research |
| D2–D3 · the shelf map and comparison table | Turns probe data you've already paid for into tasks, without spending on new probes |
| Week 1 · entity and person pages | When machines identify the wrong business, all page work returns zero or even negative; the person page also doubles as the landing point for bylined off-site contributions and video credits |
| Weeks 1–2 (in parallel) · the four Group B steps | Institutional and government source share rose from about 1/6 to nearly 1/3 (Otterly / Axios, ChatGPT's 2026-08 source reshuffle), and effort put in converts to being on the list with almost no loss — it doesn't depend on a third-party editor's approval |
| Weeks 1–2 (in parallel) · the first off-site round works only the "already present" class | Putting content into a container that's already being cited pays off on a scale of days to weeks (inferred, 65% confidence, not measured) |
| Weeks 2–4 · the thick-page main line | Price-related queries trigger AI Overviews >80% of the time, hybrid intent 97% (Whitespark); comparison pages get 1.87 citations per search, top of the whole table, and this is counted by page type — a self-built page earns it just the same (deltaV, 25,337 citations); Chinese price-type / scenario-type / second-opinion-type questions are 80% empty, the only open field within 90 days where you can pick up seats starting from zero |
| Within 14 days of each page launching · a long video narrated by the named expert | Correlation coefficient between YouTube mentions and AI Overviews visibility: 0.737 (backlinks are only 0.218, Ahrefs 2026); YouTube accounts for 23% of citations in Google AI answers (5W Citation Share); 94% of YouTube AI citations go to long-form videos, and 78% of videos with timestamps get cited repeatedly (Otterly, 1.7M data points) |
| Weeks 2–3 · bylined contributions to earned media | 84% of AI citations come from earned media (journalism alone is 27%), while paid placements and advertorials are only 0.3% (Muck Rack, 25M links / 17 industries) |
| Every Monday, weeks 2–4 · the weekly competitor seat report | You can see week by week which cells competitors are fighting for; it also doubles as next week's task list — whoever pushed you out, go and take the source they're standing on |
| Weeks 3–4 · cut the second off-site round to a minimum | Every hour freed up converts to pages; Group A is a structural gap in Singapore: staking a whole line of work on third-party inventory that cannot be created comes up empty every quarter, automatically |
| One day at the start of every month · the web control leg | The citation distribution between the two API legs and the web version is a systematic skew, not noise: overlap with authoritative-media rankings is 45.5% on the web vs 27.3% on the API; public-broadcaster sources are 34.6% on the web vs 12.2% on the API; a single domain can shift from position 1 on the web to position 61 on the API (welt.de); running 5 rounds shows the same size of skew as running 50 (University of Hamburg + Leibniz Institute, 2025-11, 24,000+ answers / five weeks) |
| D30 · the monthly review checks the door and identity unconditionally, once each | These two layers fail silently — a page that's never even being read, or AI identifying the wrong business, can easily coexist with a seat count that's rising for some other reason |
The full basis for the video leg
Beyond "video view count / likes / subscriber count are correlated with citation frequency at about −0.03" (Otterly, 1.7M data points), already given in A.1: the correlation coefficient between YouTube mentions (title / transcript / description) and AI Overviews visibility is 0.737 (Ahrefs 2026; backlinks are only 0.218 — the strongest single correlation of any factor measured, more than three times that of backlinks); YouTube accounts for 23% of citations in Google AI answers (5W Citation Share); 94% of YouTube AI citations go to long-form videos, 78% of videos with timestamps get cited repeatedly, and description length has r=0.31 (Otterly, the same batch of 1.7M data points); by platform, Perplexity 38.7% + Google AIO 36.6% — the two together account for three-quarters of YouTube citations, and the gain doesn't land in ChatGPT's seat count. 0.737 is a correlation coefficient, not causation: the mechanism hypothesis is that the YouTube page is already in the cited set (the container is already inside the candidate set, so words written on it have a chance of being copied) — causation has not been measured.
Two numbers that must not be used side by side
| Number | What it measures | The only place it may be used | Where it must never be used |
|---|---|---|---|
| Ahrefs 43.8% (ChatGPT citations from best-of blog posts, 750 queries / 26,283 URLs) | Share — how much of all citations the listicle class accounts for | Deciding whether to work a category at all (whether listicles belong in this quarter's scope) | Setting single-page priority, deciding how many letters to send, or putting it in the same table as page-type density |
| deltaV page-type density (comparison 1.87 citations/search; category 5.2%, profile 3.8%; in healthcare: listicle 0%, articles 54%; 25,337 citations) | Density and within-industry distribution — within a category you've already decided to work, which single-page type is more likely to be cited | Ordering single pages within a category you've already decided to work | Ruling out an entire category; applying it across industries |
If the two appear side by side in the same table, that is a mistake: send it back for rewriting.
Reddit and off-site source shares
Reddit's share of ChatGPT answers fell ≥73% in 2026-08 (daily average 497 → 132), while ChatGPT's total citations actually rose 3.5% over the same period; but a crash of the same size in 2025-09 later fully recovered, and ChatGPT is still reading Reddit (about a quarter of pages fetched during the study period were Reddit threads) — what stopped is putting it into the answer, not reading it (public tracking by Otterly / Axios + the same batch's fetch-side records). Criterion: only put it on the task list when this month's share is > 2%; at ≤ 2%, just log it; only write it into the regular cadence once it's been > 2% for two months running. In the same batch, institutional and government source share rose from about 1/6 to nearly 1/3 (same Otterly / Axios tracking, ChatGPT's 2026-08 source reshuffle).
Off-site timing and hit rates
- Putting content into a container that's already being cited pays off on a scale of days to weeks; a new self-built page, 2–8 weeks (inferred, 65% confidence, not measured).
- Paid listings and directory submissions: a free editorial slot, 2–4 weeks; a paid slot, 4–8 weeks (both inferred, not measured).
- Bulk outreach reply rate 15–25%; from a placement going live to a change showing up in AI citations, 2–6 weeks (inferred, 55% confidence, not measured). These two numbers are used only to schedule internal work hours.
- The time for identity fixes to pay off (7–14 days, 70% confidence, not measured) is in A.1.
Page length and differences by side
- Cited pages average 2,290 words, 78% over 1,000 words (Trakkr, 1,465 cited pages). For the length distribution broken down by page type, see rule 7 of the seven rules in A.3.
- In AI Mode, professional services win with educational content, consumer services win with review volume (PRWeb, 1,120 queries / 7 industries / 13 cities, does not include Singapore, 75% confidence).
Example (home repair): among plumbers and electricians, the winners had 337% more reviews: 546 vs 125.
This winning mechanism cannot be turned into "so the strict side should go and ask for reviews too": on the strict side, reviews from consumer-type buyers can only be observed passively.
The statistical basis for retests
If you run 20 samples for one intent, a Wilson interval only lets you judge "it moved" once it is named 10/20 times; yet a baseline of 0/10 already has an interval upper bound of 27.8%. Run it on this basis and week 13's default output is bound to be "not demonstrated", whether or not anything actually moved. The fix isn't to loosen the statistics — it's to change what counts as the acceptance measure: judge by firm, countable numbers (seat count, pages on the list, factual-error count) plus a measured noise band; the Wilson three-state test is internal reference only, used to trigger the not-moved triage, never a conclusion (rule in → General Edition 7.1 What a retest produces, and the rulers).
Question breakdown: the searches ChatGPT actually ran
- Sample: 225 stored ChatGPT answers (gpt-5.5 API with web search, Singapore location); no new queries were sent: 101 multi-industry (2026-09-14, 17 industries), 60 law firm (2026-09-29), 64 dental and aesthetics (2026-09-29). From each answer's web-search log we took the searches the model actually ran, then used Google's top 10 (Serper, gl=sg) as a proxy, looking up the original question and each search once.
- How many searches: for "who's best" questions, the median per answer was 16 multi-industry and 11 dental; for information questions, 8.5 multi-industry and 6 dental.
- Which search found the cited pages: 97 "who's best" answers cited 743 URLs; 49 (7%) were in the top 10 for the original question, 363 (49%) in the top 10 for one of AI's searches. Placebo control: swapping in the same number of searches from another question in the same industry hit only 46/736, against 314/736 for the question's own searches (the denominator 736 counts only answers that could be paired with another "who's best" question in the same industry; the question's own side is cut to the same number of searches, which is why it is 314, not the 363 above). A single search is no more accurate than the original question (on average 0.57 vs 0.51 cited URLs in each top 10); the searches win by being many and coming from many angles.
- Which words were added: of 995 multi-industry searches for "who's best" questions (one search can fall into several categories), 373 were about credentials and official sources, 238 price, 208 a year, 105 reputation, and 247 used
site:to limit the search to one website (about 150 of them to one business's own website); 79 (8%) still carried best or top. In dental, AI swaps the patient's words for specialist terms (deep bite becomes orthodontist, periodontist). - Dental and aesthetics, measured separately on 8 procedures × 4 phrasings (2026-09-29, each question asked of ChatGPT 2 times and AI Mode once): for a broad "best" question, 50/71 of the clinics named had their own website cited in the same answer, and 36/62 of the names overlapped with third-party lists for the same procedure (rough control: of the clinics in Google's top 10 for the same procedure's price and detail questions, 27/83 are also on these lists, and only 9/40 once the clinics AI named are removed); for a narrow "best" question (best plus a specific condition or group of people), ChatGPT 16/16 and AI Mode 8/8 gave a list, 76% of the clinic pages cited were deep pages such as doctor pages and procedure pages, and asking the same question twice gave a list overlap of 0.46, against 0.28 for broad "best".
- Limits: Google is only a proxy, and ChatGPT's search back end is not Google; only the ChatGPT API version was measured, so the searches AI Mode breaks a question into can't be seen; Singapore and English only, and each question was sampled on one day only; all of it is correlation, and there is no intervention experiment on whether adding pages gets you onto the list. Use it to set the categories and the build order of the question-breakdown table; never promise externally that you will get onto the list (main text in → General Edition 4.7 Actions by state, the ten question types and this quarter's slots).
Confidence list: what to print and what not to
| Statement | Confidence | How to mark it |
|---|---|---|
| The noise band for seat count is measured every month (3 same-week retests); it is no longer unknown. Before the first measured value is in hand, seat counts are only listed side by side, never judged | Measured | Print, the noise band figure must be on page 1 of the monthly report |
| "Completing the five profiles → ChatGPT's local answers benefit" (from a public partnership) | 70% | Print, on the "free, done in passing" basis; no result is promised |
| "Improving your position on a list is cheaper than getting onto a new one" | Cross-industry aggregate basis; +16.5pp is from a B2B SaaS sample, not validated in the three local industries | Print, carry the qualifier unchanged |
| Splitting industries into professional and consumer services (each follows different fallback items) is based on PRWeb's 1,120-query sample, not including Singapore | About 75% | |
| Reviews get copied word for word into the answer body: recommendation type 67% (4 of 6), non-recommendation type (price / scenario / process / second opinion) 0% (0 of 9) (searched through a generic search API on 2026-09-21, n=15, SG not locked, single engine) | 85%, very small sample | Not printed externally. Its use is to set the direction of the work: thick pages target price / scenario / process questions, and the review-copying frequency on that class of question is 0 — so thick pages are not paired with any review action, and the substitute for regulated industries is registration status, technical-standard certification and checkable-fact sentences. What the monthly report prints is your own measured values for your own M questions, split into two lines, recommendation type / non-recommendation type. The two lines must be printed side by side — printing only the recommendation-type line would lead people towards asking for reviews, something regulated industries simply cannot do |
| "60–70% of the 90-day increase lands on brand questions" | 70%, tabletop projection, not measured | Not printed, internal reference |
| For one intent, "0 businesses named at answer level; several local pages are already fighting for citation slots" (searched through a generic search API on 2026-09-21, SG not locked, single engine, one-off) | 85%, one-off sample | Do not print the specific figure. For internal ordering only; what the report prints is the named-business count and citation-slot ownership from your own baseline numbers, with the retest date and engine stated. No material may say "no one is doing this" or "competition is near zero" |
The gain from getting onto a list: the one wording for the whole book
This data point may only be written this one way anywhere in the book; quote it word for word:
Getting onto a list that AI already cites and ranking first on it = visibility +16.5pp and a position 0.8–1.8 places earlier in the answer (Peec, 5.7M data points / nearly 200,000 answers / 8 engines; +16.5pp is from a B2B SaaS sample; for emerging categories it is +13.4pp; not separately validated in Singapore dental, aesthetics or legal. Use it to order the work, never to promise any result).
A.3 Page-type measurements (1): two legs, seven rules, the nine tier-A types
What you'll do in this section: check the basis for the two-engine round (526 citations, 92 pages), how often the seven rules hold, how the count for the 46 types adds up and the disputed calls, and the cited samples for each tier-A type.
What the two probes are called: "the two-engine round" = the 2026-09-23 five-industry probe, 60 questions, the ChatGPT and AI Mode legs, 526 citations, 92 pages taken apart one by one; "the ChatGPT-only round" = 5 industries × 10 questions, ChatGPT leg only, 241 citations.
How labelling works: measured n/N (N is how many pages of this type were taken apart, n is how many of them match), measured, 1 case (backed by only one cited page — follow it, but know it's a single case), not measured (inferred or an external rule, with no cited page to check against). Two limits on the measured basis: the seed keywords all come from Google Autocomplete (gl=sg); only the first 600 characters of each answer were stored, so for the second half of an answer there's no way to match each citation to its sentence.
Denominator: of the 526 citations, only 75 could be traced to the exact sentence copied on the original page (the rest either failed to fetch, or the copied sentence was a multi-paragraph rewrite with no single source). Below, any figure like "73%" or "about 55/75" has this 75 as its denominator, not the full 526; each item's position (e.g. "about 1% in" or "about 53% in") is a visual estimate based on how far down the page it sits, not character-level coding — treat it as a strong trend, not a precise measurement to put in a client report.
Tables are the carrier copied from most often; AI copies even the column headers straight across (examples of copied headers: "Part-time | Full-time | MOE", "Best for | Finish | Price"); on 6 of the 16 dental pages cited, what got copied was a table row. A table placed later on the page still gets copied (a government fee table at the 45–51% mark).
The two legs differ: ChatGPT looks for the source, AI Mode for second-hand summaries (measured in five industries)
| Industry | What ChatGPT mainly cites | What AI Mode mainly cites |
|---|---|---|
| Dental · aesthetics | Government and public institutions 76% (38/50: MOH fee benchmarks, CPF, SingHealth, SMC) | Commercial sites such as clinics 75% (39/52); for aesthetics "who's best" questions, goes through Google cards and booking product pages |
| Family law | .gov.sg 77% (46/60; judiciary 21 times, ask.gov 11 times) | Commercial sites such as law firms 69% (36/52) |
| Tuition · education | Price questions cite peers' price-list pages; "who's best" cites organisations' own websites (about 35 of 44 URLs); judgement-type questions cite only government and media, 9/9 | Price-list pages, lists, Reddit |
| B2B SaaS | Official pricing pages about 53%, plus about 7 citations of the help centre, then writes its own conclusion | YouTube, Reddit and LinkedIn together 40%, the rest split between peer comparison pages and lists |
| E-commerce · consumer goods | Product pages (PDP) 54%, category pages 21%, government pages 13% (used as the basis for a judgement) | Google Shopping cards 44%, YouTube 15%. URL overlap between the two legs is 0 |
The denominator for each percentage is that engine's citation count for that industry's questions; site type is classified by a domain-matching rule (recomputable from the public dataset, stats.md 3.1).
Inference 2: "who's best" intents are not won by writing pages, measured in five industries
When buyers ask "who's best / best clinic / best lawyer / best tuition centre", neither leg cites a business's own "best of" article. For where to put the effort instead, see → General Edition 5.6 Two legs, two kinds of sentence.
| Industry | Where ChatGPT goes | Where AI Mode goes |
|---|---|---|
| Aesthetics, dental (n=1 question) | Public-hospital specialist pages, clinic procedure pages | 2 Google searchviewer cards + 4 booking product pages |
| Family law (n=1) | Doyle's ×2, Legal500, Chambers, ask.gov, judiciary | Law firms' own websites, SassyMama, the Law Society directory, a map card |
| Tuition (5 questions) | Organisations' own websites, about 35/44; zero citations of lists | Mainly lists (singaporetuitionteachers, sethlui, tutorcity) |
| E-commerce | PDP 54% | Shopping cards 44%; for "vs" questions, YouTube 3/3 |
| SaaS ("best CRM") | Official pricing pages 4/4 | Reddit, YouTube ×2, business.com, the Slack blog |
How often the seven rules hold
- What gets copied is almost always the first sentence of a paragraph: legal 15/18, and the other four industries can all be matched to some paragraph's first sentence. Counter-example: on a law firm's prenup-agreement page, the H2's first sentence "is a complicated matter" wasn't copied — what got copied on the same page was the FAQ entry that gives the conclusion directly. The strongest rule in the whole sample.
- The core fact falls in the first 30% of the page: about 55 of the 75 locatable citations (73%); by industry: dental 12/16, legal 11/18, tuition 7/7, SaaS about 15/19, e-commerce about 10/15. Between the H1 and the first H2: 32/75 (43%); the first or second H2's opening sentence, or the first table: 23/75 (31%); the middle or later section: 17/75 (23%); in the FAQ alone, rarely. Position is a visual estimate, as noted at the start of this section.
- Whatever can be made into a table, make it a table: frequency in the first figure at the start of this section.
- The title carries a year: dental price blog posts 8/8, tuition price pages 7/8, SaaS editorial pages 12/13, e-commerce lists 5/6; exception: official pricing pages 0/7 carry a year; pages of the One question, one page type add one only when the answer changes over time (not measured).
- A statement H1: cited clinic pages 10/11, tuition price pages 7/8, SaaS official pages 8/8; pages with ≥3 question-form H2s: only legal 3/18, SaaS 4/13, government process pages 0/4.
- The FAQ is rarely the source of the copied sentence: pages with an FAQ — dental 11/16, SaaS 17/21, tuition price pages 7/8, legal 8/18; only about 6 cases had the copied sentence come directly from the FAQ; in most, either the body didn't have it or the FAQ just restated the conclusion.
- A precise number matters more than a clean layout: mindstretcher's structure is messy, but its $272.50 is precise to the cent; a clinic's price page has 40 H2s, but the first one is a price list; singaporelegaladvice hasn't been updated in 4 years — all of these pages still get cited.
Three more rules that show up by page type:
- Word count is not a criterion: cited pages range from as short as 40 words (ask.gov), 62 words (a law firm's price list), 260 words (Talenox), up to as long as 9,000 words (MissLobang); the median is about 1,300–1,500 words; pages in the 1,800–3,000 word range: legal 2/18, tuition 2/8, e-commerce 3/11.
- Price, results and specs must be in the first-screen HTML: hard numbers on tuition-organisation pages, 6/6 readable straight from plain HTML. Counter-examples: axiom's S$420/month cannot be found in the homepage HTML, so ChatGPT had to look elsewhere; zoho switches currency by IP; mi.com's air-purifier-4 static HTML has no price.
- Outbound links and a byline are not conditions for being cited: dental pages with an authoritative outbound link, 1/16; legal, 2/18; a named individual author: legal 2/18, tuition price pages 1/8, SaaS official pages 1/8, PDP 0/5. Exception: medical pages — a named doctor in the page header, 6/12; a byline is worth doing there.
46 types: how the count adds up, and the disputed calls
The two-engine round (two legs, buying-decision questions) 9 types
New in the ChatGPT-only round (ChatGPT leg only, non-decision questions) +33 types
Absorbed from the external teardown +5 types
Merged as a duplicate (the teardown's first-party data report page folded
into the Statistics source page as its self-built variant; not a new type) −1 type
──────────────────────────────────────────────────────────────────────────────────────
Final 46 types
46 is a count that comes from the data: the two-engine round's 60 seed questions were all piled onto buying-decision questions (how much / who's best / A vs B / can it be done); the ChatGPT-only round swapped the phrasing for "what is X", "how to … step by step", "what percentage", "what to bring", "is it legal", "near <MRT station>", "is X good" and "complete guide" — and the page-type count immediately jumped from 9 to 42. The number of page types isn't an industry trait; it's a function of how much question-phrasing ground you cover.
Two pages count as the same type only when all three merge rules hold: ① the triggering questions belong to the same family ② the position and form of the copied sentence on the page belong to the same family (a single first-screen sentence / a table cell / an H3 paragraph's first sentence / a chart data label / a footer NAP string / a numbered rule) ③ both can be written from the same block blueprint. Three disputed calls:
- Type 13, Step-by-step procedure page ≠ type 7, Eligibility and process page: type 7 answers whether you meet the conditions, how many times and how long, and what gets copied is the condition sentence; type 13 answers what to do at each step, which form to submit and how much the official fee is, and what gets copied is a cell in the
Step | Resulttable. - Type 17, Subsidy and limit rules page ≠ type 7: what gets copied is the amount cell in the limits table, not a condition sentence.
- Type 22, Collected FAQ page ≠ type 6, One question, one page: type 6 is one question per URL, with the first sentence giving a yes/no; type 22 packs 4–20 questions onto one page, and in the measurements the same page was copied from three different Q&A blocks, separately.
The ruler bias of the ChatGPT-only round: because it only ran the ChatGPT leg, the 33 new types it added skew towards the Primary sources family, and page types that lean on second-hand summaries can't be measured — this is a bias from the ruler, not the full picture (the reason and the retest plan are in A.4). Tallied by page type (formal run and rerun combined; not a citation count): 228 for the new page types, 63 for the nine known types.
The nine tier-A types: cited samples
n is written as "two-engine round → ChatGPT-only round".
| Page type · n | Sample URL | Where it was copied from / which leg |
|---|---|---|
| ① Single-service price page · 6 → 2 | A dental clinic's Invisalign cost page | The first sentence after the H1 (about 1% in) + the first sentence of H2 "Invisalign or braces?" (about 53% in), cited by both legs; another example: a law firm's price list, 4 lines in all, cited by ChatGPT for a Chinese-language question. For this type, 5/6 were cited by ChatGPT only |
| ② Price guide page · 16 → folded into 21 full-guide citations | A dental clinic's crown-cost blog post | Right under the first H2, a Crown type | Cost (before GST) | Suitable for table, cited by both legs, more by AI Mode; another example, tutorbee, is in the flagged samples below |
| ③ Official pricing page · 31 → 14 | https://monday.com/pricing | Copied from three places — the price card, the footnote, the FAQ; ChatGPT 27 : AI Mode 4, and AI Mode only cited it for brand-price questions (HubSpot price questions 3/3) |
| ④ Comparison page · 20 → 1 | https://ventureharbour.com/hubspot-vs-salesforce/ | The "The short version" and "So which should you choose?" paragraphs, cited by AI Mode; AI Mode 17 : ChatGPT 3, and 2 of those 3 ChatGPT citations were third-party lab pages (RTINGS, TechRadar) |
| ⑤ List page · about 29 → 2 | https://www.misslobang.com/article/best-robot-vacuums-singapore-2026 | "S$999.90 as of mid-September", cited by ChatGPT. AI Mode about 23 : ChatGPT about 6, and all 6 of those ChatGPT citations were third-party (G2, TechRadar, MissLobang, Doyle's ×2, Legal500); these 4 list domains, 4/4, state their methodology or a check date in the title or first screen |
| ⑥ One question, one page · 23 → 6 | https://ask.gov.sg/sgcourts/questions/clywqg7te006bq97hezb0sc5l | The whole page is 40 words in two sentences; the first sentence, "There is no legal requirement…", cited by ChatGPT. Government-page version: ChatGPT 20 : AI Mode 3; the 4 commercial-site single-question-page URLs (one law firm, one divorce-lawyer keyword site, two dental clinics) were all cited by AI Mode (n=4) |
| ⑦ Eligibility and process page · 29 → 7 | https://www.judiciary.gov.sg/family/understand-requirements-getting-divorce | The table of statutory facts × explanation × when you can apply, paraphrased row by row by ChatGPT; the same page was cited under 4 different questions; citation counts by domain are in the second figure in this section |
| ⑧ Remedy and second-opinion page · 2 → — | A dental clinic's failed-implant page | The TL;DR box before the first H2 (the first 144 words), AI Mode's first citation; another example, an aesthetic clinic's page, was used by ChatGPT as its safety criterion. 2 of 2 answers discussed safety first; the legal-industry version was not measured |
| ⑨ Entity anchor page · multiple pages → 10 | https://axiomeducation.com/ | The results card (at the 10–11% mark) cited number by number by ChatGPT; fees are not in the homepage HTML. PDPs were only cited by ChatGPT; the person profile page itself was not measured — what was measured was the byline in the page header (medical 6/12, legal 2/18, tuition price pages 1/8, PDP 0/5) |
Three flagged samples:
- A range sentence on a single-service price page (a clinic's Invisalign page): "Invisalign at [the clinic] costs S$3,900 to S$7,900, depending on how complex your case is" — a range: this breaks the rules on the strict side. Do not copy this wording.
- A law-firm page that writes its own price and a market range in the same H2, cited by both legs (a law firm, measured, 1 case): "At [the firm], our fixed fee for an uncontested divorce ranges from $1,690 to $2,890." and "An uncontested divorce in Singapore generally costs between $1,500 and $3,500 (Source: Singapore Legal Advice)." — the rules do not say directly whether a law firm may put its own fees and a market range in the same place; the closest is the rule against comparing fees with other lawyers, so on the stricter reading we advise against copying it (Conservative line (not statute text)). This case only proves that both legs will cite this sentence shape (the copied sentences on three other cited pages have the same shape) — it is not a wording you may copy.
- A price-guide page's by-level price table (tutorbee): the card at the top says Primary $35–$55, the table says $25–$35 — AI skipped the card and used the table.
A.4 Page-type measurements (2): tier B and C samples, and the AI Mode retest list
What you'll do in this section: look up, type by type, the cited samples and n for each of the 33 tier-B types; tier C's statement that nothing has been measured; the priority order for retesting on the AI Mode leg; and the three things you must record alongside every run.
Statement of basis: tier A's data comes from the two-engine round (526 citations, 92 pages). Tier B comes from the ChatGPT-only round, run on the same day (2026-09-23): 5 industries × 10 questions, ChatGPT leg only (gpt-5.5 + web_search, Singapore location), 241 citations. Each type's n is a page-by-page count for that page type and combines the formal run with one extra run, so it is not the same as the number of citations. Tier C has no measurement of our own: its specs are taken, unmodified, from three files in an external teardown (the teardown target is a GEO agency's public blog). Every n = 1 finding is flagged, and it is directional only, never a settled conclusion.
The question was in Chinese, asking what HIFU is and what its side effects are. This is the only self-built commercial page that ChatGPT cited repeatedly in the ChatGPT-only round (every other commercial citation was a price guide): one page supplied several answer fragments (3 of the 7 citations came from this one page; we have no comparison data for the hit rate of single-question pages, so no multiple is given).
The 33 tier-B types: n and key evidence
| Page type | n | Key evidence |
|---|---|---|
| 10 Regulatory obligations and penalties page | 38 | IRAS AIS's obligation sentence on the first screen was copied: "AIS employers are required to submit … by 1 Mar each year. Late submission may lead to a fine of up to $5,000."; for B2B software, the ten questions were run twice (the formal run and one extra run), and 59 of the 113 citations combined (52%) landed on .gov.sg; figure in → General Edition 5.23 Tier-B questions and rules types (1): regulatory obligations, definitions |
| 11 Store / branch page | 30 | A dental clinic branch page's "Sun: 9.30am – 1pm" was copied from the end of the page; brizosystem's address line was copied from 90% of the way through the body text |
| 12 Definition page | 19 | medical 10, legal 5, training 2, SaaS 2; all 10 citations of a US hospital's health library landed on the first sentence under an H3 question, and all were conditional sentences; in e-commerce, "HEPA vs true HEPA" got 0 citations, because ChatGPT never called web_search |
| 13 Step-by-step procedure page | 14 | judiciary's Step | Result table was reordered into the answer; the $56 official fee was copied exactly |
| 14 Preparation and bring-list page | 12 | 3 of NDCS's items, grouped by patient category, were copied; the same question run twice gave 0 citations the first time and 8 the second, so stability is poor |
| 15 Legislation text page | 12 | all 8 citations of sso.agc.gov.sg carried a ProvIds provision-number anchor; sso.agc itself returned 403 and could not be retrieved |
| 16 Schedule and deadline page | 9 | commonwealthsec's date range on the first screen was copied; in one run, a single question cited 9 URLs of this type to cross-check them |
| 17 Subsidy and limit rules page | 8 | MOH CHAS's limits (up to $830 a day, $240–$5,290 for surgical items; paraphrased) came from cells in the itemised limits table |
| 18 Regulator guidance PDF | 8 | in PDPC's 62-page PDF, the Chinese question hit the line in the question-style table of contents that was phrased the same way |
| 19 Statistics source page | 8 | the parallel restatement sentence in PubMed's CONCLUSIONS section (77.6%) was copied, not the raw data in the RESULTS section. This type overturns "core facts go in the first 30%": for statistical questions, ChatGPT digs all the way to the end of the page (the SingStat sentence sits 85% of the way through, the IMDA chart at 53%) |
| 20 Register / approved list page | 7 | line 79 of the IRAS AIS PDF (at 84%) and the disclaimer at 2% were stitched into a single answer sentence; in another case, a vendor's row sat at 94% of the way through a table (ch5:2260) |
| 21 Vendor legal terms page | 7 | the list of country names in the definitions section of HubSpot's DPA (at 17%); the jst-asia-pacific commitment sentence (at 73%) |
| 22 Collected FAQ page | 6 | see the bars figure above |
| 23 Official replies and speech records | 5 | the whole sentence in mlaw's parliamentary reply giving 4,150 of 6,220 cases, or 66%, was copied |
| 24 Third-party single-business review | 4 | tutorly's at-a-glance fee table and its Pros/Cons were copied; rtings' static HTML is only 181 words of navigation, so it was listed as a citation but had no sentence that could be copied |
| 25 Sentiment, forum and news pages | 4 | of the 5 citations for the question asking whether a dental chain is any good, 2 came from news and Reddit; neither page could be retrieved (the Reddit and mothership pages were both blocking pages) |
| 26 Policy hub page | 4 | IRAS's question-style anchor structure = the AI answer's section structure; IMDA's three section pages are JS-rendered and could not be retrieved |
| 27 Parameter and rate basis page | 3 | CPF's "It was gradually raised and has reached $8,000 in 2026." was copied; the calculator tool page itself got 0 citations |
| 28 Review aggregate page | 3 | birdeye's first-screen "4.2/113 reviews" was copied; reviews.io's first-screen 4.9 does not match its JSON-LD 4.85 |
| 29 Self-built reputation and credentials page | 3 | a law firm's "[The firm]'s own family-law page shows 4.9 stars, with 3,455 Google reviews" was accepted, with the answer noting that it came from the firm's own page |
| 30 Verification and lookup page | 3 | SDC's statement that every dentist must register before practising was copied; mlaw's register lookup page had 0 words copied and was only attached at the end of the answer as an exit link |
| 31 Help centre / how-to page | 3 | the list of objects in the first paragraph of HubSpot's KB, at 13–14%, was copied; the "quotes do not sync" negative list, at 25%, was copied |
| 32 Device and product regulatory documents | 3 | NEA's line that side effects are mostly temporary and ease in about a week was copied almost verbatim; the clinic's own self-written side-effects paragraph got 0 citations |
| 33 Third-party directory listing | 2 | the first run copied its Sunday hours as 9am–9pm; in the second run of the same question, ChatGPT instead flagged it as unreliable and said to call to confirm (n = 1, directional only) |
| 34 Change notice / old-vs-new page | 2 | judiciary's "replaces" sentence was copied; talenox's "View-only mode" transition-period clause was copied |
| 35 Integration / marketplace listing | 2 | HubSpot's listing page is JS-rendered with 0 words of body text; the object list that actually got copied sits in the first paragraph of the KB page on the same site |
| 36 Case-law page | 2 | a 689-word summary page and the 16,152-word full text were cited side by side; the holding sentence sits directly under the H1 |
| 37 Category list page | 2 | mothercare's whole sentence (169 models, S$112 to S$3,045, 9 brands; paraphrased) was copied; another site's whole page is JS-rendered, with 0 words of static body text |
| 38 Calculator page (n = 1) | 1 | sgschoolkaki's tiered conclusion on the first screen and its comparison table were copied; the calculator component itself got 0 citations |
| 39 Trust centre / compliance proof page (n = 1) | 1 | trust.hubspot's "SOC 2 Type II … June 23, 2026" was copied; not one sentence of the same vendor's 3,657-word marketing version was copied |
| 40 Misconception page (n = 1) | 1 | healthhub's "When you buy and make payment for a vape, it is considered a purchase, which is also prohibited." and "Effective 1 May 2026, anyone caught owning, using and/or buying is liable to a fine up to $10,000" were copied |
| 41 Verdict-first page (n = 1) | 1 | shopback's prices (S$1,299 non-HEPA / S$1,349 HEPA) were copied, with the source credited as ShopBack, not Dyson's own site |
| 42 Time-limited promotion page (n = 1) | 1 | a dental chain's nett-price sentence (scaling + polishing + intraoral 3D scan for S$109 nett at participating clinics; paraphrased) was copied; the page itself was blocked by Cloudflare and could not be retrieved |
The six n = 1 items in the table (38–42, plus the single observation in type 33) are directional only, not settled conclusions.
The four tier-C types: nothing measured
| Page type | n | Note |
|---|---|---|
| 43 Buyer's selection framework | 0 | the teardown target's version of this archetype cannot be found anywhere in the Google US top 30; the ChatGPT-only round had no matching question |
| 44 Concept pillar page | 0 | adjacent to the ④ Comparison page and type 12, the Definition page, but with a different intent; to be verified |
| 45 News and policy explainer | 0 | for rule-type questions, almost 100% of citations land on the original .gov.sg page; explainers on commercial sites get 0 citations |
| 46 Market observation page | 0 | the teardown target's own 5 Type=Market articles are the only internal-link dead ends on the whole site (0 links out), and they are not even in the /blog index |
The teardown target's own pages of these types do not even rank organically. Its raw ranking data shows 8 of 31 non-brand terms in the top 10 and 5 at #1, all of them list pages; its tool-comparison pages and single-competitor review pages cannot be found in the Google US top 30. The teardown target measured organic rankings and we measured AI citations: numbers from the two rulers cannot be combined.
Three conclusions that recur across page types: ① AI does not click calculators. For "calculator" or "how do I calculate" questions, ChatGPT did not cite a single interactive calculator tool, and this repeated across three industries. It cites the static table next to the calculator, the tiered conclusion sentence on the first screen and the authoritative source for each parameter in the formula, then does the arithmetic itself (shared by types 17 / 27 / 38). ② ChatGPT looks for the source, AI Mode for second-hand summaries: full evidence in A.3. ③ The teardown target's market-observation articles are internal-link dead ends: see type 46 in the table above.
Waiting for the AI Mode leg
During the ChatGPT-only round, the AI Mode leg's data interface (SerpApi) had run out of quota, so the AI Mode leg was not run at all. All 33 tier-B types therefore have evidence from the ChatGPT leg only. "Only the ChatGPT leg has been measured" does not mean AI Mode does not cite them; it only means they have not been measured. The known direction of the bias: ChatGPT looks for source pages and AI Mode for second-hand summaries, and in e-commerce the two legs' URL overlap was 0. The 33 newly added types skew systematically towards the Primary sources family, and page types that lean towards second-hand summaries could not, by construction, show up in the ChatGPT-only round.
| Priority | Page type | Expected difference | Why we expect this |
|---|---|---|---|
| High | 24 Third-party single-business review | AI Mode should be markedly higher than ChatGPT (only 4 citations in the ChatGPT-only round) | In the two-engine round, the ④ Comparison page was AI Mode 17 : ChatGPT 3 and the ⑤ List page 23 : 6; reviews and lists are AI Mode's home ground |
| High | 28 Review aggregate page, 29 Self-built reputation and credentials page | AI Mode may go to Google cards / map cards instead and not cite these two types | In the two-engine round, for "who's best" questions AI Mode went to searchviewer cards + booking product pages |
| High | 25 Sentiment, forum and news pages | AI Mode should be higher (all 4 citations in the ChatGPT-only round depended on ChatGPT digging to page 3 / page 23 of Reddit) | In the two-engine round, AI Mode's Reddit + YouTube + LinkedIn made up 40% combined for SaaS |
| High | 12 Definition page | The cited pages may change wholesale: ChatGPT cites the source, while AI Mode may switch to peers' long educational articles | The cleanest comparison point for the source-vs-second-hand split |
| Medium | 37 Category list page, 41 Verdict-first page | AI Mode may switch to citing Google Shopping cards and YouTube | In the two-engine round, AI Mode in e-commerce: Shopping cards 44%, YouTube 15% |
| Medium | 14 Preparation and bring-list page | Stability itself needs retesting (the same question run twice gave 0 → 8 citations) | Whether checklist pages win citations swings wildly |
| Medium | 38 Calculator page | Test whether "AI does not click calculators" holds on both legs | It repeated across three industries in the ChatGPT-only round (n = 3 questions), but all on the ChatGPT leg |
| Medium | 19 Statistics source page, 23 Official replies and speech records | AI Mode may use media paraphrases instead of the primary source's own figures | If so, "build your own statistics page that states n = and its basis" needs a different way to win on the AI Mode leg |
| Low | 10 Regulatory obligations and penalties page, 15 Legislation text page, 18 Regulator guidance PDF | Both legs are expected to cite them, with little difference | Legal facts have only one source; a second-hand summary cannot replace it |
| Must test | 43 Buyer's selection framework, 46 Market observation page (tier C) | The ChatGPT-only round cannot measure them, by construction | Both are second-hand-summary / discursive forms, so only the AI Mode leg can give the first piece of evidence. If type 46 still gets 0 citations on the AI Mode leg, it can be closed out and not built |
Whatever question you run, record these three things alongside it:
- The full answer text: both probe rounds stored only the first 600 characters. Next round, save the whole answer; otherwise, in a long answer, "which sentence got copied" can only be guessed.
- Rerun the same question: strong swings were observed for type 14 (checklists) and types 24/28 (reviews). Run each question ≥2 times and record the difference in citation counts between the runs.
- Fill in the pages that could not be retrieved: the ChatGPT-only round flagged 7 places as "could not be retrieved" (sso.agc.gov.sg 403; a dental chain's page blocked by Cloudflare; ishopchangi.com / ecosystem.hubspot.com / IMDA section pages JS-rendered; PubMed's cookie wall; G2/Capterra 403; mothership/Reddit blocking pages). Next round, render them in a browser before breaking them down; this is exactly why type 42's format spec is currently empty.
A.5 List of sources
What you'll do in this section: look up, entry by entry, the original sources for the evidence and rules used across the whole book: full name, URL, date retrieved.
This section lists only general sources: consumer-protection and price-transparency guidelines, do-not-call guidelines, destination-country advertising law, platform documentation and data research. Each industry regulator's own advertising regulations, FAQs and register lookup pages belong in that industry edition's appendix, not here. Where the original gives no URL or date retrieved, write "not given" or "not stated" as it stands; do not fill it in.
Consumer protection and price-transparency guidelines
- CCCS · Guidelines on Price Transparency:
cccs.gov.sg/consumer-protection/legislation-and-guidelines/guidelines-on-price-transparency(date retrieved not stated in the original) - The CPFTA statute (Consumer Protection (Fair Trading) Act):
sso.agc.gov.sg/Act/CPFTA2003(date retrieved not stated in the original)
Do-not-call guidelines
- PDPC · Do Not Call Registry and Your Business:
pdpc.gov.sg/overview-of-pdpa/do-not-call-registry/business-owner/do-not-call-registry-and-your-business(date retrieved not stated in the original)
Destination-country advertising law
- China's Advertising Law (2021), English translation:
chinalawtranslate.com/advertising-law-2021/(date retrieved not stated in the original) - SAMR's Enforcement Guidelines on Absolute Terms in Advertising (no URL given in the original; date retrieved not stated in the original)
Platform documentation
- Gartner Peer Insights · Incentives FAQ:
gpivendorresources.gartner.com/en/articles/6812506-incentives-faqs(date retrieved not stated in the original) - Gartner Peer Insights · Review Sourcing FAQ:
gpivendorresources.gartner.com/en/articles/6812574-review-sourcing-faqs(date retrieved not stated in the original) - OpenAI merchant feed spec:
developers.openai.com/commerce/(date retrieved not stated in the original) - OpenAI merchant portal:
chatgpt.com/merchants(date retrieved not stated in the original) - Google, AI features and your website: eligibility for AI Overviews and AI Mode, Merchant Center and Business Profile listed as best practice, no extra schema needed; original wording in A.1 (no URL given in the original; read on 2026-09-21)
- Google's crawler documentation (what
Google-Extendedis for, and the original wording "does not impact a site's inclusion in Google Search"; see A.1); Google's 2026-05-15 AI-optimisation guide (llms.txtnot needed) (no URL given for either; date retrieved not stated in the original) - Apple, About Applebot:
Applebotserves Spotlight / Siri / Safari retrieval,Applebot-Extendedis only a training opt-out switch (no URL given in the original; date retrieved not stated in the original) - Anthropic's official documentation:
ClaudeBot/Claude-SearchBot/Claude-Userall respect robots.txt (no URL given in the original; date retrieved not stated in the original)
Data research (sample sizes, time periods and qualifiers are all in the matching entries in A.1–A.4; this is only an index and does not repeat the numbers)
- Citation density and page types: deltaV, 25,337 citations (A.1)
- Earned media share: Muck Rack, 25M links / 17 industries (A.1)
- Zero JS execution: Vercel × MERJ, 500M+ fetch samples (A.1)
- Share of price queries that trigger AI Overviews: Whitespark, 540 queries (A.1)
- The gain from getting onto a list: Peec, 5.7M data points / nearly 200,000 answers / 8 engines (A.2)
- Skew between the citation distributions of the web leg and the API legs: University of Hamburg + Leibniz Institute, 2025-11, 24,000+ answers / five weeks (A.2)
- Video correlation coefficients: public correlation studies, namely Ahrefs 2026, 5W Citation Share and Otterly's 1.7M data points (A.2)
- The rest, which have only the researcher's name and no URL in the original: searchVIU, Ahrefs (JSON-LD tracking, Wikipedia share), Steady Demand, 5WPR, Pew, Omniscient, SE Ranking, BuzzStream, Profound, Peec (
result_sourcecapture) (all in A.1); PRWeb, Trakkr, Ahrefs (listicle share), Otterly / Axios (Reddit and source-reshuffle tracking) (all in A.2)