WHITEPAPER · Part 7 of 13Full text

Chapter 4 · Visibility as a Distribution: A Measurement Standard for a Non-Deterministic Environment

Marketing in the Agent Era · Canlah AI · a Singapore SEO + GEO agency

≈ 15 min read · 2026-08

You are reading the complete text — not a summary. All 13 chapters are published in full, free, no registration. The PDF is a print edition of the same content.

0.103Jaccard overlap of cited sources: API layer vs browser layer
0.45–0.59day-to-day brand-list similarity within a single layer
503complete answers archived in one client audit

4.1 A single measurement is worthless

AI answers are stochastically generated. Ask the same prompt twice and the brand list may differ. The industry broadly agrees on this — and then stops there. Most vendors acknowledge it and go on selling “rankings” derived from single measurements.

The best available evidence comes from a study covering three German-language verticals between 24 January and 20 March 2026 (A−, preprint): day-to-day Jaccard similarity for brand mentions was 0.45–0.59 (consumer electronics 0.557, sporting goods 0.453, telecoms 0.589); at the source level it was less stable still, at 0.34–0.42.

Put plainly: ask the same question today and tomorrow, and about half the brand list changes, while two-thirds of the cited sources change.

That study excluded finance and real estate because brand recognition rates fell below 70%. The exclusion is itself worth noting: in some verticals, even “was the brand mentioned at all” is hard to determine stably.

Same question, today vs tomorrow: about half the brand list changes (day-to-day Jaccard similarity)
A "ranking" taken from one measurement: by the next day about half the brand list has changed; the source level is less stable still, with two-thirds of the cited sources changing. The study excluded finance and real estate because brand recognition rates fell below 70%.

4.2 Canlah’s own measurement: two quantities the industry has conflated

In one client audit we archived 503 complete responses; the API-layer samples among them were OpenAI 240 + Gemini 240 (227 each retrieval-enabled), covering 60 “engine × layer × question” combinations, comprising 3,821 citation instances and 409 unique domains. Collection cost was USD 86; secondary analysis added nothing.

We had set out to find a cost-saving threshold — “run until the source set converges, then stop”. The result overturned that design.

Metric Behaviour within 8 rounds Rounds required
Brand mention rate Unnamed layer 0/262, named layer 192/192, identical at k = 1 / 2 / 3 / 8 1 round (take 2 as a failure buffer)
Cited source set At round 8 the median group was still adding 1 new domain; baseline coverage only 0.88; only 36% of groups added nothing in that round No affordable number of rounds buys it

Stopping-rule results (recall relative to that group’s full 8-round union): a fixed 2 rounds gives median domain recall of only 50%; reaching 95% recall requires 5.6 rounds, saving just 30% against the current 8. No threshold achieves both economy and fidelity.

From this follows the core methodological claim of this report.

The industry has blended two quantities of entirely different character into a single score. “Was I mentioned?” is cheap and immediately stable. “Who was cited?” is expensive and never converges. Any product selling a single AI visibility score is averaging a quantity that has converged together with one that never has.

This has direct value for buyers. Brand mention rate can be measured cheaply and frequently (1–2 rounds suffice), so it is worth running as a continuous monitoring metric. The complete picture of who is being cited cannot be enumerated within any reasonable budget; any product claiming to deliver a complete citation league table is reporting a truncated sample, and usually does not disclose where the truncation falls.

Two quantities blended into one score: one stabilises in 1 round, one never converges in 8
"Was I mentioned?" is cheap and immediately stable; "Who was cited?" is expensive and never converges. Any product selling a single AI visibility score averages a quantity that has converged with one that never has. Limitation (4.3): 0/262 and 192/192 are both saturated extremes; this dataset contains no middle-case sample (brands with mention rates around 30–60%), so "one round is enough" holds only for the saturated case and must not be extrapolated.

4.3 The limitation of that claim (which must appear alongside it)

The two figures supporting “one round is enough” — 0/262 and 192/192 — are both saturated extremes. The brand is entirely invisible in unnamed category queries and invariably mentioned in named ones. At the extremes, “more rounds change nothing” is close to a trivial result.

What a scoring model actually has to serve is the middle case — brands with mention rates around 30–60% — and this dataset contains no middle-case sample at all. Until that calibration is filled in, “one round is enough” holds only for the saturated case and must not be extrapolated.

We state this limit in the body rather than in an appendix, because extrapolating a conclusion into an unmeasured range is exactly the behaviour this report criticises throughout.

How we contain it: until the mid-range calibration is filled in, for brands whose mention rate sits in the middle of the range we calculate the required sample size from the binomial distribution and set the number of rounds accordingly, rather than applying “one round is enough”. Every mention rate in every report is given as a range, never as a point value. In other words, the cost of this limit is borne by us as a higher sampling cost; it is not passed on as false precision in the client’s report.

4.4 Cross-layer divergence: a quantity larger than day-to-day variation

In a separate audit (n = 12 questions), we put the same queries simultaneously to the API layer and the browser layer, and compared the source sets each cited:

Mean Jaccard 0.103 (median 0.105 · SD 0.046 · range 0.028–0.170 · 95% CI [0.074, 0.132]).

The two layers are reading almost entirely different documents.

This dataset also supplies a self-refuting example. On the surface, brand-hit agreement between the layers is 83.3%, which looks reasonable. Split it apart: the 8 questions where both layers reported no hit agree 100% (conveying no information at all), while the 4 questions carrying real signal agree 50% — the same as a coin toss. The 83.3% is diluted by the empty questions.

Two quantities that must be kept distinct and never placed on the same axis:

What it measures Value
German-language study (A−) Within one layer, across days Brands 0.45–0.59; sources 0.34–0.42
Canlah (first-party) Across layers (API vs browser), same day Sources 0.103

The correct combined statement is: the magnitude of cross-layer divergence is roughly three times that of day-to-day variation. It must not be stated as “Canlah measured lower agreement than the academic studies” — that compares two different quantities as though they were one.

Practical implication: for any assertion delivered to a client about “what AI says about us”, the headline evidence must be a browser-reproducible interface screenshot. The API is appropriate for scale, trend and relative share; any assertion the client can falsify in their own browser in one step must pass through the browser layer. In one real case (the walk-through in Chapter 3) we observed the same brand on one grounded API channel across 3 dates producing a consistently wrong answer (stating the business had not opened), while the browser layer produced the correct answer over the same period; a second API channel did not make that error — which is why “which model, which layer, which phrasing” has to be stated every time.

How an 83.3% agreement rate is diluted by empty questions
Brand-hit agreement between the API layer and the browser layer, n = 12 questions: a headline number that looks reasonable collapses to a coin toss once split apart. In the same dataset, mean Jaccard of the cited source sets across the two layers is 0.103 (median 0.105 · SD 0.046 · range 0.028–0.170 · 95% CI [0.074, 0.132]) — the two layers are reading almost entirely different documents; this is a different quantity from Figure 1's day-to-day 0.45–0.59 and must never be placed on the same axis. Headline evidence delivered to a client must be a browser-reproducible interface screenshot.

4.5 A proposed measurement standard

On this basis we propose four minimum standards, and publicly commit Canlah’s delivery to them.

  1. Pre-register the prompt panel. Publish the full prompt set and generation rules before measurement begins. Common industry practice is for the panel to be written by people who already know the answer, which is methodologically equivalent to setting your own exam.
  2. Set n separately for mention rate and for source set, and report them separately. Do not composite them into a single score.
  3. Report variance, not just point values. Give a repeat-measurement interval for every metric.
  4. Headline evidence is browser-level, with timestamped screenshots and complete source lists; API results are labelled separately as “API-layer reference”.

None of these requires proprietary technology; any competitor may adopt them. We publish them because in a category where measurement is unreliable, reproducibility is the only defensible differentiation.

4.6 A real pre-registered prompt panel (Singapore maths-olympiad tuition)

Of the four standards above, the first is the easiest to promise and the hardest to check. So here is a panel from a real audit, reproduced as it was used: a DeepAudit deliverable dated 29 July 2026 (the masked sample was archived on 1 August) for a maths-olympiad training centre in Singapore (the client’s name, domain and address are masked as ▓▓▓▓▓▓; competitors appear as “Competitor-A▓” and so on). The probe ran 36 prompts across 6 layers, 6 per layer, sampled twice each on OpenAI and Gemini — 144 calls (142 succeeded, 2 failed; the failures are archived too). The table lists 24 of them: the first 20 rows excerpted in the deliverable’s Appendix B, none dropped, reproduced as printed, plus the 4 brand-named prompts from the notarised samples in Appendix C, so that every layer has at least two. The probe language was English; wording is verbatim. Every figure below is Canlah first-party; every funnel mention rate is API-layer sampling, not browser-reproduced, and under standard 4 of 4.5 it counts only as “API-layer reference”.

Generation rules (the four questions behind standard 1 in 4.5)

  1. Who writes it? Not the client, and not an analyst who already knows the answer. A model (Gemini) synthesises both the layer structure and the prompt wording from the client’s business profile, a vertical skeleton and a rules file. Code then hard-checks the output: regulated verticals must carry a trust-verification layer; a business with a physical location must carry a local layer, an online-only one must not; the brand and competitor layers are never dropped; banned templates are filtered; placeholder syntax is enforced. Anything non-compliant is pruned, and if pruning cannot fix it the whole synthesis falls back to static templates, with the fallback recorded. Of the 36 prompts, 12 (33%) are rewrites of real user search data (11 from Google Autocomplete, 1 from Google People Also Ask), each with its original phrase and a source link; the other 24 are synthesised per buyer-journey layer. The deliverable’s table labels all 24 “template synthesis”, because the rendering layer only distinguishes “rewritten from a real query” from “everything else” — it does not separate model synthesis from static templates. The deliverable states “Funnel structure synthesized based on business profile customization”; we did not open this run’s synthesis record to confirm the fallback flag, so that point is unverified.
  2. Is the answer known before it is written? No. The panel is generated before probing starts and archived with the model name, prompt hash and raw response under _source/funnel_synthesis/ in the deliverable. Autocomplete and PAA come from Google, which does not know which brand the AI will name. A third-party methodology critique (B) notes that the industry’s usual panels are written by people who already know the brand under test, so the conclusion is embedded before measurement begins. This rule exists to answer that.
  3. What share carries no brand name? By layer structure, 24 of 36 (67%) contain no brand at all (S1–S4); 12 name the client or a competitor (S5–S6). Only the former count towards the visibility score. In the named layers, the AI repeating the question already counts as a “mention” — a tautology — so they carry zero weight and are reported for reference only. The 24 is “6 per layer × 4 layers”; the 20 S1–S4 prompts visible below do carry no brand name, and the 4 we cannot see (the last 4 of S4) are unverified.
  4. How is the market kept separate? At the time, this panel passed two gates: every query ran in a gl=sg context, and candidates carrying a US anchor (“USA / United States”) were rejected before entering the pool — that rule went into the code on 25 July, before this probe. What did not exist yet was a stricter gate. An incident recorded on 7 August: a clinic client’s gl=sg pool contained a same-named clinic in a small English town, plus candidates like “… london” and “… meaning in urdu” — Google Autocomplete completes on the seed string; gl only biases ranking, it never excludes a foreign entity. After that incident the code gained a deterministic geo/language gate: any foreign market place-name (London, Sydney, Kuala Lumpur…), language name or non-target script is rejected, only the audited market’s own place-names (“Singapore” and its districts) pass, and rejections are counted and disclosed, never dropped silently. That gate was not applied retroactively to this 29 July panel. As far as the table shows, no foreign entity got into this panel — but that was luck, not a gate.
# Layer Intent Prompt (verbatim) Source (deliverable’s own label)
1 S1 Category awareness category · recommendation Which center offers the best math olympiad training in Singapore? Autocomplete: best math olympiad training singapore
2 S1 category · recommendation What are the top options for math olympiad training in Singapore? Autocomplete: math olympiad training singapore
3 S1 category · recommendation What is the best math olympiad training center in Singapore for… People Also Ask: What is the best tuition in Singapore?
4 S1 category · question Is math olympiad training worth it for primary school students in… Autocomplete: is math olympiad worth it
5 S1 category · question How to choose a reliable math olympiad training program in Singapore? Autocomplete: how to train for math olympiad
6 S1 category · recommendation What are the most recommended math olympiad training classes in… Autocomplete: math olympiad training for kids
7 S2 Compliance qualification verification question · trust How to check if a math olympiad training center is MOE-registered in Singapore? Template synthesis
8 S2 question · trust Are all math olympiad training providers in Singapore legally registered with the Ministry of Education? Template synthesis
9 S2 question · trust How to verify the official qualifications of math olympiad training… Template synthesis
10 S2 question · trust What safety and educational standards must tuition centers offering… Template synthesis
11 S2 question · trust Where can I find the official list of registered math olympiad… Template synthesis
12 S2 question · trust How to avoid unaccredited education agencies when booking math… Template synthesis
13 S3 Geo-regional exploration local Where can I find top-rated math olympiad training near ▓▓▓▓▓▓? Autocomplete: math olympiad training near me
14 S3 local Are there any reputable math olympiad training classes in ▓▓▓▓▓▓, Singapore? Autocomplete: math olympiad classes near me
15 S3 local Which math olympiad training center near ▓▓▓▓▓▓ has the most… Autocomplete: mathematical olympiad training center near me
16 S3 local Looking for high-scoring math olympiad training preparation near… Autocomplete: math olympiad preparation near me
17 S3 local What are the best math competition training options near ▓▓▓▓▓▓? Autocomplete: math competition classes near me
18 S3 local Best physical tuition centers for math olympiad training around… Autocomplete: maths olympiad classes near me for kids
19 S4 Competition achievement proof question · evidence Which math olympiad training center in Singapore has the highest SASMO or NMOS medal win rate? Template synthesis
20 S4 question · evidence What are the proven competition results of math olympiad training students in Singapore? Template synthesis
21 S5 Brand-name direct query brand Is ▓▓▓▓▓▓ worth it for high-level math competition prep? Appendix C notarised sample (source not listed)
22 S5 brand What do parents say about the teaching quality at ▓▓▓▓▓▓? Appendix C notarised sample (source not listed)
23 S6 Competitor-name comparison comparison Competitor-C▓ vs ▓▓▓▓▓▓ — which center is better for Olympiad math? Appendix C notarised sample (source not listed)
24 S6 comparison Should I switch my child from Competitor-A▓ to ▓▓▓▓▓▓ for advanced training? Appendix C notarised sample (source not listed)

“…” is the deliverable’s own table truncation; the full sentence sits in the deliverable’s raw/ archive and we do not complete it here. Rows 7, 8, 14, 19 and 20 take their full sentence from the Appendix C notarised samples; the rest are as printed in Appendix B. Row 3’s original phrase is “What is the best tuition in Singapore?” — a rewrite from a broader category question. We print it; we do not defend it. 20 of these 24 carry no brand name; the mask in “near ▓▓▓▓▓▓” covers an address, which under the masking rules is not a brand name.

Why the layers look like this

What this panel does not yet do


This report is free to read in full. Want the PDF edition for forwarding and archiving? Get it at the whitepaper home →