The GEO Vendor Evidence Checklist: 12 Questions, 9 Red Flags, One Scorecard
What proof a generative engine optimization vendor should hand you — 12 procurement questions with weak vs strong answers, 9 red flags, and a scorecard you can run in one meeting. We score ourselves against it too.
Disclosure
Canlah AI sells the service this checklist evaluates. So the last section scores us against the same twelve questions, including where we fall short. If a vendor hands you a checklist that only their own product passes, that is itself a red flag — check ours against the criteria below.
TL;DR
A generative engine optimization vendor cannot control what a model says, so the only thing you can hold them to is the rigour of their measurement and the traceability of their evidence. Ask for: repeated sampling with reported frequencies (not single screenshots), per-engine breakdowns, a locked query set agreed before work starts, raw archived responses you can re-open, and a falsifiable definition of failure. The twelve questions below separate vendors who measure from vendors who narrate.
Why This Category Needs a Checklist More Than Most
In traditional SEO you could audit a vendor's claims yourself: open an incognito window, type the keyword, look at position four. Generative engines removed that check. Ask ChatGPT the same buyer question twice in fresh sessions and you will frequently get different brands and different citations — so a vendor can screenshot a good answer, and you cannot reproduce it, and neither of you is lying.
We measured how bad that instability is. In our own seven-engine test (August 2026), asking the same question across engines produced an average Jaccard overlap of 0.091 in English and 0.171 in Chinese between the source sets returned. In plain terms: two engines answering the same question agree on roughly one source in ten. Repeat runs on a single engine are steadier than that, but nowhere near stable. Independent analyses of AI Overviews point the same way — the overlap between what ranks in Google's top ten and what gets cited in AI answers has fallen from roughly three-quarters to somewhere between a fifth and two-fifths, depending on whose crawl you read.
That is the whole procurement problem in one paragraph. When the target moves this much, a single screenshot proves nothing and a frequency proves something. Every question below is a way of asking whether your vendor knows the difference.
The 12 Questions
Ask these in the pitch meeting. The middle column is what an answer sounds like when a vendor is narrating; the right column is what it sounds like when they are measuring.
| Ask | Weak answer | Strong answer |
|---|---|---|
| 1. How many times do you run each query? | "We check regularly." | A number, five or more per engine per cycle, with the cadence written into the contract. |
| 2. Do you report a frequency or a yes/no? | "You're visible in ChatGPT." | "Cited in 4 of 5 runs on ChatGPT, 0 of 5 on Perplexity, week of the 14th." |
| 3. Which engines, measured separately? | One blended "AI visibility score". | Per-engine columns, because engines cite largely different sources. |
| 4. Who writes the query set, and when is it locked? | Queries appear in the first report. | You approve the buyer-intent set before baseline; changes are versioned and disclosed. |
| 5. Do you measure a baseline before any work? | Work starts, numbers arrive later. | A dated pre-work baseline on the locked set — otherwise improvement is unfalsifiable. |
| 6. Browser-level or API-level probes? | Doesn't distinguish. | Says which, and why: consumer-facing answers are what buyers see; API results are a cheaper proxy that diverges. |
| 7. Can I re-open a specific answer you reported? | A cropped screenshot in a slide. | Archived raw responses with timestamps, retrievable per finding. |
| 8. Named or cited — which are you counting? | Uses "mentioned" and "cited" interchangeably. | Separates brand named in prose from your domain linked as a source; reports both. |
| 9. What sources are winning instead of us? | Only reports your own numbers. | A ranked list of the pages engines cite in your category — the actual work list. |
| 10. How would we know this programme failed? | "These things take time." | A pre-agreed failure condition and review date, in writing. |
| 11. How do you build off-site presence? | Vague "community seeding" or "authority building". | Named platforms, disclosed identities, no bought engagement — and they can explain the regulatory exposure of the alternative. |
| 12. What happens to the data if we leave? | Data lives in their dashboard. | You get the query set, the raw responses and the history — it is your measurement record. |
9 Red Flags
Any one of these is a reason to slow down. Three of them together is a reason to walk.
1. A guarantee of AI rankings or citations
No vendor controls model output. A guarantee is either a misunderstanding of the product or a sales device, and both should worry you. What can legitimately be guaranteed is the work: probe volumes, engine coverage, content output, placements.
2. A single "AI visibility score" with no per-engine breakdown
Blending engines that cite almost entirely different sources produces a number that cannot be acted on. If it goes down, nobody can tell you which surface moved.
3. Screenshots as the primary evidence format
One good answer, cropped, is the easiest artefact to produce and the least informative. Ask what the other four runs said.
4. Triple-digit uplift percentages attributed to "the Princeton study"
The peer-reviewed GEO study (KDD 2024) reports its strongest methods — adding statistics, quotations and cited sources — at roughly +41% on position-adjusted word count and about +28% on subjective impression, with keyword stuffing showing little to no improvement. Figures well above that, quoted casually, are usually fabricated somewhere upstream in the deck.
5. llms.txt sold as a citation lever
No major generative engine has confirmed using it, and Google has publicly said it does not. Publishing one is cheap and harmless; selling it as the reason you will get cited is not.
6. Undisclosed community activity
Manufactured consensus — accounts that read as ordinary users but are paid to advocate — is a disclosure violation in most jurisdictions and an explicit prohibition in China's 2026 GEO group standard. It is also the cheapest-looking line item that can become the most expensive.
7. No baseline, or a baseline measured after work started
Without a dated pre-work measurement on a locked query set, every subsequent number is unfalsifiable. This is the single most common gap we see in competitor reports.
8. Brand-name queries presented as visibility wins
If the question contains your brand name, the engine repeating it proves nothing. Visibility means appearing for questions asked by buyers who have never heard of you.
9. Reports that never contain bad news
Real measurement produces losses as well as gains — answers drift, competitors publish, engines swap retrieval sources. A reporting history with no declines is a reporting history that is being curated.
The Scorecard
Score each of the twelve questions: 2 if the vendor answers with specifics you could verify, 1 if the intent is right but the mechanism is vague, 0 if the answer is a claim rather than a method. Twenty-four points available.
| Total | Read it as |
|---|---|
| 19–24 | A measurement practice. Negotiate on scope and price, not on evidence. |
| 12–18 | Real capability, immature instrumentation. Workable if the gaps become contract terms. |
| 6–11 | A content shop with an AI label. You will not be able to tell whether it worked. |
| 0–5 | Narration. Ask for a pilot with a locked query set before any retainer. |
Our Own Answers, Including the Weak Ones
Running the checklist on ourselves, in the same format we would want from anyone else. Where we score well: our probes run in real browser sessions rather than only through APIs, we report frequencies per engine rather than a blended score, every reported answer is archived with a timestamp and can be re-opened, and we publish original measurement openly — our agent-readiness dataset of 50 cross-border DTC brands is public under CC BY 4.0 with a DOI (10.5281/zenodo.22103178), including the finding that unvalidated checks overreport readiness by 6.4 points because single-page-app soft-404s return 200 for everything.
Where we are honest about limits: sampling is a disciplined proxy, not omniscience — we cannot see every conversation a buyer has with an engine, and we say so rather than implying full coverage. Cadence is daily or weekly by tier, never "real-time", because it is not. And no engagement comes with a citation guarantee, for the reason at the top of the red flags list.
If you want to see what the evidence layer actually looks like before you buy anything, the fastest route is to run a baseline on your own domain, or to try the manual version yourself with our step-by-step guide to checking whether ChatGPT recommends your brand. How we scope and price against a measured gap is set out on the pricing page, and if you are still building a shortlist, our ranked guide to Singapore GEO agencies applies these same criteria to the market, us included.
Frequently Asked Questions
What evidence should a GEO vendor provide?
At minimum: a query set you approved before work began, a dated pre-work baseline, per-engine results reported as frequencies across repeated runs, archived raw responses you can re-open per finding, a separation between brand mentions and domain citations, a ranked list of the sources currently winning your category, and a written failure condition with a review date. Screenshots and a blended "AI visibility score" are not evidence.
How do I verify a GEO agency's case study?
Ask for the query set and the run count behind the headline. Then ask what the same queries returned in the runs that were not shown. A case study built on repeated sampling can answer both questions immediately; one built on a lucky screenshot cannot. Also check whether the queries contained the client's brand name — brand-name questions inflate results without proving discovery.
Can a GEO vendor guarantee AI citations?
No. Generative engines are probabilistic systems that no agency controls, and retrieval sources change continuously. A vendor can contractually commit to the work — probe volumes, engine coverage, content produced, placements earned — and to reproducible measurement of the outcome. Treat any guarantee of placement or ranking inside AI answers as a reason to ask for the measurement protocol behind it.
Is a single AI visibility score useful?
Not on its own. Engines cite substantially different sources for the same question — in our own seven-engine test the average overlap between engines' source sets was 0.091 in English — so blending them hides the only actionable information: which specific surface moved, and why. Ask for per-engine numbers, and ask what the score would look like if one engine were removed.
Take the Checklist Into Your Next Pitch
Print the twelve questions, score every vendor you meet, and keep the sheets side by side. The exercise costs one meeting and reliably separates the vendors who will be able to tell you whether it worked from the ones who will tell you a story about it. If you want the baseline that makes those answers checkable, request an AI visibility audit or write to admin@canlah.ai.
FREE WHITEPAPER
Marketing in the Agent Era
All 13 chapters public — no email wall. Includes an original dataset on the agent-readiness of 50 cross-border DTC brands.
Read it free →