Skip to main content
Perspectives 2026-09-23 Updated 2026-09-23 · 10 min read

From Snapshot to Distribution: Why AI Visibility Reporting Is Competing on Evidence You Can Re-Run

AI visibility reporting is moving from one-off snapshots to evidence a buyer can re-run. What SparkToro, Profound and Evertune published, and what to test.

Layered dark cards receding into depth, one edge lit indigo

AI visibility reporting is being judged on evidence a second party can re-run, not on a single snapshot score. SparkToro’s January 2026 study found that ChatGPT and Google’s AI almost never return the same list of brands twice for the same prompt. Profound published an experiment in July 2026 arguing that one run per prompt per day is enough for its visibility score. Evertune states that it samples every prompt 100 times per model. These are vendor and analyst positions, not independent audits, but they show the direction: a single answer is treated as one draw from a distribution.

The buyer’s question is no longer whether a vendor shows a visibility score. It is whether the evidence behind it can be re-run: the same questions, engines, interfaces and definitions, a stated number of runs and an archive of raw answers and cited URLs. In multilingual markets each engine and language adds its own variance. A screenshot is one component of AI visibility evidence, not the evidence itself.

Key findings

  • Lists are unstable: SparkToro and Gumshoe.ai collected 2,961 runs of 12 prompts from 600 volunteers and reported a less than 1-in-100 chance that ChatGPT or Google’s AI returns the same brand list twice.
  • Mentions and sources converge at different speeds: Profound reports that ten daily runs cut citation-share noise by about 40% but visibility noise by only about 11%, so one run count does not fit both metrics.
  • Canlah AI’s own audit was a thin sample: on September 7, 2026, no open question produced a Canlah AI mention in every ChatGPT run, and only 10.6% of linked domains recurred across all three runs.
  • A re-run record is the unit of evidence: question set, interface, run count, raw archive, range and retest should travel with every number in an AI visibility report.

A single answer is one draw, not a measurement

SparkToro founder Rand Fishkin and Patrick O’Donnell of Gumshoe.ai asked 600 volunteers to run 12 prompts through ChatGPT, Claude and Google’s AI Overviews or AI Mode, for 2,961 runs in total. Each prompt was run 60 to 100 times. Nearly every response differed in brands, order and list length. The study puts the chance of two identical brand lists at less than 1 in 100 for ChatGPT and Google’s AI, and the chance of the same order at about 1 in 1,000.

Fishkin did not conclude that tracking is pointless. He wrote that visibility percentage across dozens to hundreds of prompts run multiple times is a reasonable metric, while any tool reporting a “ranking position in AI” is not. Gumshoe.ai sells AI visibility tracking, and the study discloses that conflict. It is a US consumer sample, not a benchmark for business-to-business questions or for Asian markets.

Sources: SparkToro, AIs Are Highly Inconsistent When Recommending Brands or Products

A snapshot of an unstable list can show a brand first on Monday and absent on Tuesday with nothing changed in the market. A report built on repeated answers describes a distribution, and only a distribution can be checked by someone else.

How vendors are pricing the sample

Stated sampling designs across the four products reviewed on September 23, 2026 span two orders of magnitude, from one run per prompt per day at Profound to 100 runs per prompt per model at Evertune.

Sampling designs published by four AI visibility vendors, pricing and documentation pages reviewed September 23, 2026.

Vendor Stated sampling design What the stated design does well Interface stated What a buyer should verify
Profound Runs every tracked prompt once a day; its July 2026 study compared 1 and 10 runs a day on 753 prompts across seven US platforms Publishes the experiment behind the choice, with prompt count, window and the measured noise reduction States that it captures answers from the browser rather than the API Whether a two-week window of thousands of prompts matches the buyer’s own prompt count
Evertune Samples every prompt 100 times per model across 11 or more AI models States a fixed repeat count per model, so the denominator behind each figure is known before signing Not found on reviewed page Which interface produced the 100 samples, and whether the raw answers can be exported
Peec AI Prices tracking in credits: 1 prompt on 1 model for 1 day equals 1 credit Publishes the credit arithmetic openly, so coverage can be computed from a plan before purchase Not found on reviewed page How many runs sit behind each daily data point, and what coverage the credit budget buys
Otterly.AI Daily tracking frequency on all plans, from 15 prompts on the Lite plan to 400 on Premium Publishes the prompt allowance at every tier, so depth per dollar is comparable across plans Not found on reviewed page Whether a daily reading is a single run, and how variance is reported

Source: vendor pricing and documentation pages reviewed September 23, 2026. Stated sampling design is public positioning, not independently tested capability.

“Not found on reviewed page” means the vendor’s public materials did not name that detail when reviewed on September 23, 2026. It is not proof that the capability is absent.

Sources: Profound, Is Once a Day Enough; Profound Answer Engine Insights; Evertune FAQ; Peec AI agency pricing; Otterly.AI pricing

Profound’s experiment is the most detailed public argument in the set. Over June 1 to 14, 2026, it reports that one run a day landed within about two percentage points of a ten-run reading for visibility, and that ten runs a day cut day-to-day noise in citation share by roughly 40%. The argument rests on a large prompt portfolio carrying the statistical weight, which a buyer with 30 questions does not have.

Why a dashboard number is not evidence on its own

Two vendors can publish the same visibility percentage from samples a hundred times apart, because Profound states one run per prompt per day and Evertune states 100 runs per prompt per model, yet the chart looks the same either way. Without the run count, a five-point movement cannot be separated from sampling noise. Without an interface label, an API answer cannot be checked against what a buyer sees in a normal browser session. Without raw answers and cited URLs, the score cannot be audited after the reporting period closes.

A composite score adds a fourth problem. Brand mention rate and the cited-source set behave differently under repetition, so a blend inherits the instability of the less stable metric.

What run-to-run variation looked like in Canlah AI’s own audit

Canlah AI ran its audit pipeline on its own site, canlah.ai, on September 7, 2026, with 32 open English buyer questions for Singapore and three runs per question. The OpenAI leg used gpt-5.5 with web search through the API. ChatGPT mentioned Canlah AI in 7 of 84 grounded open answers. Those mentions came from six questions, and none of them produced a mention in every run: four questions produced one mention in three runs, one produced two in three, and one had only a single grounded run.

The sources moved more than the mention. On 27 open questions answered with sources in all three ChatGPT runs, 284 distinct question-and-domain pairs appeared, and 30 of them, or 10.6%, appeared in all three runs. The mean overlap between any two runs of the same question was 24.9%. In browser-rendered Google AI Overviews, 26 questions had two completed renders. Three showed an Overview in one run and not the other, and the other 23 shared a mean of 43.8% of linked domains between their two renders.

This is a measurement weakness in Canlah AI’s own process, not a customer result. Three runs in a compressed window can show that the variance exists but cannot bound it. The Gemini leg also used a floating model alias, so its exact version was not recorded.

The evidence shows that mention and citation vary at different rates on this question set, not how large the variance would be over days, in other categories or on other engines.

Sources: Canlah AI, The Own-Brand GEO Audit Experiment

Mention rate and cited sources do not converge together

Canlah AI’s 2026 whitepaper describes one client audit of 503 archived answers, including 480 API samples across OpenAI and Gemini, in which the brand mention rate was identical at one, two, three and eight rounds while the cited-source set never settled. The mention result was 0 of 262 on unnamed questions and 192 of 192 on named ones at every round. After eight rounds the median question group was still adding a new domain, and a fixed two rounds recovered a median of 50% of the domains seen across all eight.

Both mention figures are saturated extremes, and the whitepaper states that the result does not extend to brands with mention rates between 30% and 60%. A separate 12-question comparison of API and browser answers to the same queries found a mean source overlap of 0.103.

Sources: Canlah AI whitepaper, Visibility as a Distribution

The re-run record

Canlah AI proposes a six-field re-run record as the minimum unit that turns an AI visibility number into reproducible evidence. It can live in a platform export, a spreadsheet or an agency report, as long as a second party can repeat the measurement.

The six-field re-run record, as proposed by Canlah AI on September 23, 2026.

Field What to record Decision it supports
1 Question set Frozen question wording, language, market and the split between branded and open questions, fixed before any answers are read Whether a change is in the market or in the questions
2 Interface and version API or browser, model version, location, account state and collection date Whether the result matches what a buyer sees
3 Run count per metric Separate run counts for mention rate and for the cited-source set Whether each number has enough draws to be read
4 Raw archive Full answer text, cited URLs, timestamps and failed or excluded runs Whether the score can be audited after the fact
5 Range Mention rate as an interval with its denominator, not a point value Whether a later movement falls outside ordinary noise
6 Retest Same fields rerun after a change, labelled as association unless a control design supports causation Whether an action coincided with a real shift

Source: Canlah AI reporting proposal, September 23, 2026. It is a standard for buyers to demand, not an industry standard.

A longer dashboard is not the same as a reproducible one. If any field is missing, the next report cannot be compared with the last.

What buyers should test

Six questions separate an AI visibility report that can be re-run from one that only looks precise.

  • Can the vendor state how many runs sit behind each number, per engine and per interface?
  • Are brand mention rate and cited-source share reported separately rather than blended into one score?
  • Is each rate shown as a range with its denominator, and are failed or excluded runs disclosed?
  • Are the raw answers and cited URLs exportable for the full reporting period?
  • Was the question set frozen before the answers were read, with branded questions kept out of the visibility score?
  • Does a retest use the same questions, interface and definitions as the baseline?

Ask any vendor, Canlah AI included, to produce the raw answers behind one reported number before signing, rather than accepting the chart that summarises them.

Where Canlah AI fits

Canlah AI is a Singapore-based SEO + GEO agency that measures and improves brand visibility inside ChatGPT, Gemini, Google AI Overviews and Google AI Mode, reporting AI visibility as re-verifiable ranges with timestamped evidence. Its engagements start from a locked buyer-question pool, keep raw answer archives and report mention rates as ranges.

It is most relevant to brands that need a re-runnable evidence trail across ChatGPT, Gemini, Google AI Overviews and Google AI Mode, not to every company seeking a low-cost daily tracker. Its own data also shows the current limit. The free instant check on canlah.ai returns a single-pass citability score, which is a snapshot by design, and its September 7, 2026 self-audit used three runs per question in one window. That is the gap on which AI visibility reporting is competing.

Methodology and limitations

Vendor sampling designs come from public Profound, Evertune, Peec AI and Otterly.AI pages reviewed on September 23, 2026 and were not independently tested. The SparkToro study was co-produced with a tracking vendor on a US consumer sample. Canlah AI figures come from its own September 7, 2026 audit archive of canlah.ai and from one anonymised client audit described in its 2026 whitepaper. The two Canlah AI datasets differ in category, engines, rounds and dates and were not combined. Neither contains a factual-accuracy field or a control group.

Frequently asked questions

What evidence should an AI visibility report include?

An AI visibility report should state the question set, engines, interface, market, dates and runs per question behind every number. It should separate branded from open questions, report mention rate as a range with its denominator and keep the raw answers and cited URLs so the result can be re-run.

Is an AI visibility dashboard screenshot enough evidence?

No. A screenshot shows one answer from one run in one interface. It illustrates what a buyer may see but cannot show how often that answer appears.

How many times should each prompt be run?

There is no universal number. Profound argues that one daily run across a large prompt portfolio is enough for visibility, while Evertune samples 100 times per model. Mention rate usually settles faster than the cited-source set, so the run count should be set per metric and disclosed.

References

Related articles

FREE WHITEPAPER

Marketing in the Agent Era

All 13 chapters public — no email wall. Includes an original dataset on the agent-readiness of 50 cross-border DTC brands.

Read it free →