The AI Visibility Market Has a Measurement Problem—and a Sampling-Depth Gap
AI visibility tools sell a score built from one answer per prompt per day. Mention rates need intervals, and cited-source sets need far deeper sampling.
The AI visibility market has a measurement problem: in the four tool plans reviewed on September 23, 2026, the headline score rests on one answer per prompt, per engine, per day. That answer is not stable. The same prompt sent to the same engine twice can return a different brand list, a different order and different cited pages. A SparkToro and Gumshoe study published on January 27, 2026 sized the problem, and Canlah AI audit archives show what repeated sampling does and does not fix.
The sharper issue is sampling depth: how many answers are collected per question, engine and layer before a rate is reported. Public tool plans describe coverage as prompts multiplied by engines multiplied by days, yet rarely state how many answers stand behind one day’s number. Brand mention rate and the cited-source set also behave differently under repetition, so blending them into one score hides which part is measured and which part is estimated. A visibility score is one sample from a distribution, not the distribution itself.
Key findings
- Lists are not stable: in SparkToro’s January 27, 2026 study, 600 volunteers ran 12 prompts 2,961 times; two responses had less than a 1 in 100 chance of returning the same brand list.
- Two quantities, one score: whether a brand is mentioned converges quickly under repetition, while the set of cited sources keeps growing. A single composite averages a converged quantity with an unconverged one.
- Canlah AI’s own gap: in an archived client audit, a fixed two rounds recovered a median 50% of the cited domains that eight rounds found, and Canlah AI’s September 7, 2026 self-audit requested a floating Gemini alias instead of pinning a model, although every response logged its exact version.
- A five-layer record: a defensible AI visibility number needs a pre-registered panel, a stated sampling plan, raw evidence, an interval and a comparable retest.
What the SparkToro study measured
SparkToro founder Rand Fishkin and Patrick O’Donnell of Gumshoe.ai asked 600 volunteers to run 12 recommendation prompts through ChatGPT, Claude and Google AI Overviews, using Google AI Mode when no Overview appeared. The volunteers submitted 2,961 responses, each prompt running 60 to 100 times per tool.
The study reports less than a 1 in 100 chance that ChatGPT or Google’s AI returns the same brand list in any two of 100 responses, and roughly a 1 in 1,000 chance of the same order. Ranking position behaved close to random.
Appearance rate across many runs still held meaning: one digital marketing agency appeared in 85 of 95 Google AI responses to the same prompt. The unit of measurement is therefore the share of repeated answers naming a brand, not its position in one answer.
Sources: SparkToro, AIs are highly inconsistent when recommending brands or products, January 27, 2026
The study is vendor-affiliated: a co-author works for an AI tracking company, and its volunteers used US consumer interfaces rather than API collection.
Why a single snapshot misleads
The output is a lottery with weights
Recommendation answers are drawn from a weighted pool, so a brand with a true 30% chance of being named is absent from most single answers and present in some. One observation cannot separate a 30% brand from a 5% brand or a 60% brand.
Daily cadence is not repetition
Thirty answers collected one per day mix two sources of change: sampling noise within a day and genuine drift across days, including model updates.
Interfaces read different documents
API answers and browser-rendered answers are separate measurement layers, so a tool that samples one layer and reports the result as “what ChatGPT says” transfers a claim the buyer cannot reproduce in a browser. Across 12 questions in a Canlah AI audit, the two layers’ cited sources shared a mean Jaccard overlap of 0.103.
Branded questions inflate the rate
A question that names the brand returns the brand almost every time. When branded and open questions share one denominator, the score measures recognition, not discovery.
What public tool plans disclose about sampling depth
Across the four products reviewed on September 23, 2026, public pricing pages express depth as prompt allowance, engine count and cadence; the number of repeats within a collection run was not stated on any reviewed page.
AI visibility tool plans, public pricing and help pages reviewed September 23, 2026.
| Provider | Published tracking unit | What the published plan does well | Repeats per prompt per run | What a buyer should verify |
|---|---|---|---|---|
| Peec AI | 50, 150 or 350 prompts on three models, daily; answers = prompts x models x days | States the answer arithmetic openly, so the denominator can be computed before signing | Not stated on reviewed page | Whether one day’s figure for one prompt and model rests on a single answer |
| OtterlyAI | 15 prompts at $29, 100 at $189 and 400 at $489 per month, daily | Publishes allowance and price at every tier, so depth per dollar is comparable before purchase | Not stated on reviewed page | Whether a larger allowance buys more questions or more runs per question |
| Profound | Trial of 50 prompts across ChatGPT, Gemini and Google AI Overviews, daily for seven days | Names the engines in the trial, so the engine mix is known before signing | Not stated on reviewed page | Whether seven days separates model drift from same-day sampling noise |
| Semrush AI Visibility Toolkit | 25 tracked prompts at $99 per month; score reported relative to the category median | Supplies a comparison baseline outside the brand’s own history, which a raw rate omits | Not stated on reviewed page | How the category median is built and whether it moves when the competitor set changes |
Source: provider pricing and help pages reviewed September 23, 2026. Published tracking units are public positioning, not independently tested capability.
“Not stated on reviewed page” means the provider’s public materials did not describe that field when reviewed on September 23, 2026. It is not proof that repeats or intervals are absent from the provider’s internal process.
Sources: Peec AI pricing; OtterlyAI pricing; Profound pricing; Semrush AI Visibility Toolkit help page
Peec AI’s pricing page gives the clearest arithmetic: 50 prompts on three models for 30 days produce 4,500 answers, yet each daily figure for one prompt and model still rests on one answer. Semrush is the only one of the four that scores a brand against a category median rather than against its own history, which gives a baseline outside its own dashboard.
A first-party view from Canlah AI
Canlah AI’s September 7, 2026 self-audit of canlah.ai sent 39 English buyer questions for the Singapore market to two API engines, three runs each: 32 open questions and seven that named the brand. Of 234 planned answers, 233 succeeded and 216 were grounded. On 177 grounded answers to open questions, Canlah AI was named 7 times, a 4.0% mention rate with a 95% Wilson interval of 1.9% to 7.9%. ChatGPT produced all seven mentions, 7 of 84 or 8.3%; Gemini produced 0 of 93.
Three runs per question limit what a single question can show: its rate can only read 0%, 33.3%, 66.7% or 100%. The runs were also collected in a compressed window with no spacing, and the Gemini leg requested a floating model alias instead of a pinned model; every response still logged its exact version (gemini-3.5-flash-lite). This is a measurement gap in Canlah AI’s own process, not a client result.
A separate archived client audit tested depth directly: it stored 503 complete answers, including 240 each from OpenAI and Gemini across 60 engine, layer and question groups, with 3,821 citation instances and 409 unique domains. Collection cost was USD 86. The evidence shows which sample size each measure needs; it does not show how either relates to revenue.
Mention rate and cited sources need different sample sizes
In that client audit, brand mention rate did not move with depth. Open questions returned 0 of 262 mentions and branded questions 192 of 192, with the same result at one, two, three and eight rounds. The cited-source set behaved differently. At round eight the median group was still adding one new domain, and only 36% of groups added nothing in that round.
A fixed two rounds recovered a median 50% of the domains found across all eight. Reaching 95% recall required 5.6 rounds, a saving of only 30% against eight, so no stopping rule delivered both economy and completeness.
The mention result carries its own limit: both figures are saturated extremes, and the archive holds no brand in the 30% to 60% range, where extra rounds matter most, so “one round is enough” holds only for the saturated case. Citation share and mention rate are different instruments, not two readings of one gauge.
How many prompts and runs a visibility number needs
The required depth depends on the rate being measured and the difference a buyer wants to detect, so there is no universal prompt count. For a brand observed at 30%, 20 answers give a 95% interval of 14.5% to 51.9%. Sixty answers narrow it to 19.9% to 42.5%, 100 answers to 21.9% to 39.6% and 300 answers to 25.1% to 35.4%.
Those intervals set the smallest change a report can honestly call movement: with 20 answers, a shift from 30% to 45% sits inside the noise.
Cross-layer evidence adds a second constraint. The 12-question audit puts a 95% interval of 0.074 to 0.132 around that 0.103 overlap, and brand-hit agreement was 83.3% overall but 50% on the four questions with any signal. Claims about what a buyer will see belong to the browser layer; API samples suit scale and trend.
A five-layer measurement record
Canlah AI proposes a five-layer record for any AI visibility figure a budget depends on.
The five-layer record, as applied in Canlah AI’s September 7, 2026 self-audit.
| Layer | What to record | Decision it supports |
|---|---|---|
| 1 Pre-registered panel | Question set, generation rules, branded and open split, market and language, fixed before collection | Whether the result was measured or written into the questions |
| 2 Sampling plan | Runs per question per engine, spacing, separate n for mention rate and for source discovery | Whether the depth matches the precision claimed |
| 3 Evidence record | Full answer, cited URLs, timestamp, pinned model version, API or browser layer, failed runs | Whether any number can be traced to raw answers |
| 4 Interval report | Rate, denominator and interval per engine and layer; branded results reported separately | Whether an observed change exceeds sampling noise |
| 5 Comparable retest | Same panel, same model versions where available, same n, dated before and after | Whether an action is associated with a change, without claiming causation |
Source: Canlah AI reporting standard, September 23, 2026. It is a proposal for buyers to demand, not an industry standard.
A panel without a sampling plan produces a precise-looking number of unknown precision, and a plan without pinned versions produces a retest that cannot be compared. A larger dashboard is not the same as a deeper sample.
What buyers should test
- Ask how many answers stand behind each reported rate per engine per day, and how many runs each question receives per collection.
- Require an interval, or the raw counts needed to compute one, beside every mention rate.
- Check that branded and open questions are reported with separate denominators.
- Confirm which layer each figure comes from, API or browser, and whether model versions are pinned.
- Request a positive control: one question where the brand should certainly appear, run through the same pipeline.
Treat those five answers as the test, not the score. A figure that cannot be re-run from a raw record on a stated date is a claim, not a measurement.
Where Canlah AI fits
Canlah AI is a Singapore-based SEO + GEO agency that measures and improves brand visibility inside ChatGPT, Gemini, Google AI Overviews and Google AI Mode, reporting AI visibility as re-verifiable ranges with timestamped evidence.
It is most relevant to brands that need a visibility number defensible enough to fund work against, not to every company seeking a cheap daily dashboard. The company’s own archive shows the limit as of September 23, 2026: its mid-range calibration is unfilled, and its self-audit ran with an unpinned Gemini version. That is the layer on which AI visibility providers are competing in 2026.
Methodology and limitations
SparkToro figures are vendor-affiliated research based on volunteer submissions from US consumer interfaces, not an independent benchmark. Tool plan details come from public pricing and help pages reviewed on September 23, 2026; no paid accounts were used, and internal sampling practices may differ from public descriptions. Canlah AI figures come from its September 7, 2026 self-audit archive and from anonymised client audit archives summarised in its 2026 whitepaper. The datasets are not a matched experiment. The intervals above are binomial illustrations, not estimates of any named brand.
Frequently asked questions
What is the AI visibility measurement problem?
The AI visibility measurement problem is that a score presented as a fixed fact is one sample from an output that changes between runs. In SparkToro’s January 27, 2026 study, two responses to the same prompt had less than a 1 in 100 chance of returning the same brand list. A rate measured across many repeated answers still carries meaning, while a rank or a mention read from one answer per day does not.
Are AI visibility tools accurate?
No. A single daily reading is not accurate in the sense most buyers assume. AI visibility tools can count answers correctly, but one answer per prompt per day samples an unstable output, so the figure carries an interval that was not stated on any of the four pages reviewed on September 23, 2026. Accuracy improves when runs are repeated, branded and open questions get separate denominators and model versions are pinned.
How many prompts are needed to measure AI visibility?
There is no universal number, because precision depends on total answers, not prompts alone. For a brand near 30%, 20 answers give an interval of roughly 15% to 52%, while 300 answers narrow it to roughly 25% to 35%. Canlah AI’s own self-audit used 32 open questions, three runs each on two engines, and still reported an interval of 1.9% to 7.9%.
Related articles
FREE WHITEPAPER
Marketing in the Agent Era
All 13 chapters public — no email wall. Includes an original dataset on the agent-readiness of 50 cross-border DTC brands.
Read it free →KEEP READING
- Food Media Is Becoming ChatGPT's New Evidence for Recommending Singapore Restaurants Perspectives
- From Snapshot to Distribution: Why AI Visibility Reporting Is Competing on Evidence You Can Re-Run Perspectives
- GEO Is Becoming an Agent-Run Agency Service: How Standardization Is Reshaping AI Visibility Retainers Perspectives