Marketing in the Agent Era · Canlah AI · a Singapore SEO + GEO agency
AI answers are stochastically generated. Ask the same prompt twice and the brand list may differ. The industry broadly agrees on this — and then stops there. Most vendors acknowledge it and go on selling “rankings” derived from single measurements.
The best available evidence comes from a study covering three German-language verticals between 24 January and 20 March 2026 (A−, preprint): day-to-day Jaccard similarity for brand mentions was 0.45–0.59 (consumer electronics 0.557, sporting goods 0.453, telecoms 0.589); at the source level it was less stable still, at 0.34–0.42.
Put plainly: ask the same question today and tomorrow, and about half the brand list changes, while two-thirds of the cited sources change.
That study excluded finance and real estate because brand recognition rates fell below 70%. The exclusion is itself worth noting: in some verticals, even “was the brand mentioned at all” is hard to determine stably.
In one client audit we archived 503 complete responses (OpenAI 240 + Gemini 240, of which 227 each were retrieval-enabled), covering 60 “engine × layer × question” combinations, comprising 3,821 citation instances and 409 unique domains. Collection cost was USD 86; secondary analysis added nothing.
We had set out to find a cost-saving threshold — “run until the source set converges, then stop”. The result overturned that design.
| Metric | Behaviour within 8 rounds | Rounds required |
|---|---|---|
| Brand mention rate | Unnamed layer 0/262, named layer 192/192, identical at k = 1 / 2 / 3 / 8 | 1 round (take 2 as a failure buffer) |
| Cited source set | At round 8 the median group was still adding 1 new domain; baseline coverage only 0.88; only 36% of groups added nothing in that round | No affordable number of rounds buys it |
Stopping-rule results (recall relative to that group’s full 8-round union): a fixed 2 rounds gives median domain recall of only 50%; reaching 95% recall requires 5.6 rounds, saving just 30% against the current 8. No threshold achieves both economy and fidelity.
From this follows the core methodological claim of this report.
The industry has blended two quantities of entirely different character into a single score. “Was I mentioned?” is cheap and immediately stable. “Who was cited?” is expensive and never converges. Any product selling a single AI visibility score is averaging a quantity that has converged together with one that never has.
This has direct value for buyers. Brand mention rate can be measured cheaply and frequently (1–2 rounds suffice), so it is worth running as a continuous monitoring metric. The complete picture of who is being cited cannot be enumerated within any reasonable budget; any product claiming to deliver a complete citation league table is reporting a truncated sample, and usually does not disclose where the truncation falls.
The two figures supporting “one round is enough” — 0/262 and 192/192 — are both saturated extremes. The brand is entirely invisible in unnamed category queries and invariably mentioned in named ones. At the extremes, “more rounds change nothing” is close to a trivial result.
What a scoring model actually has to serve is the middle case — brands with mention rates around 30–60% — and this dataset contains no middle-case sample at all. Until that calibration is filled in, “one round is enough” holds only for the saturated case and must not be extrapolated.
We state this limit in the body rather than in an appendix, because extrapolating a conclusion into an unmeasured range is exactly the behaviour this report criticises throughout.
How we contain it: until the mid-range calibration is filled in, for brands whose mention rate sits in the middle of the range we calculate the required sample size from the binomial distribution and set the number of rounds accordingly, rather than applying “one round is enough”. Every mention rate in every report is given as a range, never as a point value. In other words, the cost of this limit is borne by us as a higher sampling cost; it is not passed on as false precision in the client’s report.
In a separate audit (n = 12 questions), we put the same queries simultaneously to the API layer and the browser layer, and compared the source sets each cited:
Mean Jaccard 0.103 (median 0.105 · SD 0.046 · range 0.028–0.170 · 95% CI [0.074, 0.132]).
The two layers are reading almost entirely different documents.
This dataset also supplies a self-refuting example. On the surface, brand-hit agreement between the layers is 83.3%, which looks reasonable. Split it apart: the 8 questions where both layers reported no hit agree 100% (conveying no information at all), while the 4 questions carrying real signal agree 50% — the same as a coin toss. The 83.3% is diluted by the empty questions.
Two quantities that must be kept distinct and never placed on the same axis:
| What it measures | Value | |
|---|---|---|
| German-language study (A−) | Within one layer, across days | Brands 0.45–0.59; sources 0.34–0.42 |
| Canlah (first-party) | Across layers (API vs browser), same day | Sources 0.103 |
The correct combined statement is: the magnitude of cross-layer divergence is roughly three times that of day-to-day variation. It must not be stated as “Canlah measured lower agreement than the academic studies” — that compares two different quantities as though they were one.
Practical implication: for any assertion delivered to a client about “what AI says about us”, the headline evidence must be a browser-reproducible interface screenshot. The API is appropriate for scale, trend and relative share; any assertion the client can falsify in their own browser in one step must pass through the browser layer. In one real case we observed the same brand across 3 engines × 3 dates producing a consistently wrong answer at the API layer (stating the business had not opened), while the browser layer produced the correct answer over the same period.
On this basis we propose four minimum standards, and publicly commit Canlah’s delivery to them.
None of these requires proprietary technology; any competitor may adopt them. We publish them because in a category where measurement is unreliable, reproducibility is the only defensible differentiation.