WHITEPAPER · Part 6 of 13Full text

Chapter 3 · What GEO Can Actually Change, and What It Cannot

Marketing in the Agent Era · Canlah AI · a Singapore SEO + GEO agency

≈ 13 min read · 2026-08

You are reading the complete text — not a summary. All 13 chapters are published in full, free, no registration. The PDF is a print edition of the same content.

45GEO studies covered by the July 2026 critical review
99.4% vs 3.32%brand recall: query names the brand vs unnamed category query
r = +0.395Reddit presence — the strongest observed correlate of discovery

3.1 The strongest counter-argument first

On 15 July 2026, a critical review covering 45 GEO studies reached this conclusion: no technique under review demonstrated a stable, longitudinal, cross-platform causal effect on organic discoverability or on downstream behaviour (arXiv:2607.14035, A−, single-author preprint).

We put that sentence at the top of this chapter rather than in a footnote, for three reasons.

First, any buyer with judgement will find it within a single search. Having it produced against you, and producing it yourself and then answering it, are two entirely different positions to argue from.

Second, it is substantially correct. Below we set out the one point on which we think it needs qualifying — but that is a scope qualification, not a rebuttal.

Third, the existence of that review is itself the justification for the measurement standard set out in Chapter 4 of this report. If the causal effect has not yet been stably demonstrated, then continuing to sell a “ranking improvement” product is dishonest.

So, with causality unproven, what is left that can responsibly be delivered? Three things, and all three are verifiable:

  1. Factual errors can be found and corrected. False statements about your brand in AI answers — that you are not open, that you are some other company with a similar name, that you are associated with third-party content you do not control — are objective factual errors, and judging them right or wrong requires no causal theory at all. In a real case we observed the Gemini grounded API channel stably returning “this outlet is not open” on three consecutive days, while the browser layer, over the same period, showed it open (4.6 stars, 543 reviews; Google Places’ authoritative record: 541). The full timeline, the control groups, and what the case does not prove are set out in the walk-through below. Fixing and verifying this class of problem does not depend on whether “GEO works”.
  2. Gaps are measurable, and they are not probabilistic. Whether your product page carries structured data is a binary fact, not a correlational inference. Every cell in the table in Chapter 0 can be reproduced on your own site.
  3. We have not been able to prove causality, so we use pre-registered controls instead. On any ongoing engagement we publish the prompt panel and the decision rules before we start, we hold out a control group that receives no optimisation, and we report the results in full to the client whether or not they meet the target. This is not a substitute for causal proof, but it turns “did the money you spent change anything” from a vendor’s assertion into a criterion both parties agreed in advance.

The honest product definition therefore is: we sell diagnosis, error correction and controlled experiments; we do not sell ranking promises. That scope is narrower than the industry’s customary claims, but every item inside it can be verified by you at the point of delivery.

Walkthrough · one factual error, end to end

Point 1 above says factual errors in AI answers can be found and corrected. Here is the one case we hold, laid out in full: how it was found, what the evidence looks like at each step, what we did, and what it does not prove. Brand, domain and address are masked the way our sample reports mask them (client entities → ▓▓▓▓▓▓). Our sample convention normally keeps media sources, but the single media source in this case would re-identify the client on its own, so it is masked too. Everything below is Canlah first-party. The raw JSON sits in our internal evidence bundle and can be opened file by file on engagement.

Setting. A Singapore restaurant group, ▓▓▓▓▓▓, runs a rooftop bar, ▓▓▓▓▓▓, which opened in the second half of 2024 and had been trading for over a year and a half by July 2026. While auditing the group on 10 July 2026 we asked one question three ways: “Which restaurants does this group run in Singapore?”

Timeline.

Date · time (SGT) Engine / channel Layer What we observed Evidence
07-10 18:40 Google Places API Authority (the only adjudicable layer) openNow=true, 4.6 stars, 541 reviews, primaryType=bar Raw JSON (Places fields)
07-10 18:54–18:56 Gemini 2.5 Flash + google_search grounding API 3/3 samples wrote “the upcoming ▓▓▓▓” / “set to open” — i.e. not yet open 3 raw JSON files with sentence-level groundingSupports
07-10 same batch OpenAI web-search channel API No “not open” in 3 samples; 2 of the 3 did not mention the bar at all 3 raw JSON files
07-11 23:49–23:52 Re-test matrix: 12 calls, 5 control groups API See “Five controls” below 12 raw JSON files + summary JSON
07-12 early hours Google AI Mode, logged-in real Chrome Browser Card verbatim: “4.6 (543) Bar” / “Open”; body text in the present tense innerText extraction (screenshots stayed on the browser-extension side; not archived in this case’s evidence bundle)
07-12 12:44 Gemini 2.5 Flash + grounding, same window, re-run API Still “upcoming”, and self-dated “as of a September 2024 report” 1 raw JSON file

543 versus 541 is two reviews of drift across two days — which, in turn, confirms the browser layer and the authority layer were looking at the same entity.

Five controls (07-11). We did not stop at “the API is wrong”; we took apart why.

Root cause (the evidence-anchored part and the inferred part, kept separate). In the three 07-10 raw files, groundingSupports anchor every “upcoming” sentence to the same retrieved chunk: a page on a mainstream local English-language outlet (one sentence is additionally anchored to the group’s own concepts page). On 07-11 we resolved the redirect links and confirmed the page is a pre-opening feature published on 29 September 2024; on 1 September 2026 we reopened the article and the sentence “the upcoming ▓▓▓▓, a Mediterranean-inspired rooftop cocktail bar and lounge” is still there, near-verbatim what the AI wrote. One limitation to state: the redirect links inside the JSON have since expired, so this step can no longer be reproduced from the raw files alone — only from our resolution log at the time plus the reopened article. We then crawled the group’s whole website and searched for upcoming|set to open|coming soon: zero hits — the AI was not fed “not open” by the client’s own site. But the concepts page gives this bar one line, “New chic rooftop dining lounge bar located at the ▓▓▓▓”, plus a logo (re-crawled 1 September 2026, wording unchanged); the bar’s own sub-site is thin; and that “New” line is the very text the AI co-cited. The most “authoritative” description of it on the open web is still that pre-opening piece, and the site offers nothing that could outrank it. This is a content vacuum, not a copy error.

Why the browser layer was right — this paragraph is our inference, based on the Gemini API documentation rather than on a measurement: the documentation states that the API’s google_search grounding reads the web-text index only, while Google Maps structured data (open status, rating, review count) is a separate tool, off by default and billed separately; the Gemini web app and AI Mode appear to fuse that entity layer. One screen, two pipelines, two answers — the phenomenon is measured; the pipeline explanation is inferred.

What we did. The 07-10 audit report prescribed a corpus-side fix: put a machine-readable “now open” narrative (present tense, hours, address) plus structured data on the brand’s own domain, so that fresh content has a chance to outrank the old piece; group E had already shown the domain itself is citable — what is missing is the content. What we have not done is the correction. We have not carried it out: on 1 September 2026 we re-crawled the group’s concepts page and its description of the bar is still the same “New …” line, and our internal archive holds no post-correction re-test of any kind. This case carries a diagnosis record only; the “correct → re-test” step is not part of this report, and we will not pass off group E or group C as that step — those changed the question or the model, not the corpus.

What this case proves.

  1. Judging a factual error needs no causal theory. Google Places’ open-status field is an adjudicable authority; an AI answer is either right or wrong against it.
  2. The error is stable, not noise. Same model, same channel, 10, 11 and 12 July: 3/3, 3/3, 1/1.
  3. The error comes from retrieved stale corpus, not from stale model memory (group B). That is what places the fix on the corpus side rather than “wait for the model to update”.
  4. The API layer and the browser layer gave opposite answers in the same window. Any “AI visibility monitoring” run on the API alone would hand the client either false comfort or false alarm on this fact.
  5. “The AI got it wrong” only means something with three qualifiers attached: which model, which layer, which phrasing. Change any one and the answer changes.

What this case does not prove.

  1. It does not prove that publishing content fixes it. The post-correction re-test has not been run; we do not know how long new content would take, or whether it can outrank a piece that has stood for more than 21 months.
  2. It does not prove the browser layer is ground truth. The same AI Mode screen listed as current a restaurant of the group’s that we have reason to believe is closed; that venue’s Places record was not archived in this case’s evidence bundle, so we do not adjudicate it — we only note that each layer is stale in its own way, and neither is the truth.
  3. It does not prove a causal effect for GEO. This is one brand, one fact, one channel; n = 1. What it shows is that this class of error can be traced to a specific source — nothing more.
  4. It does not generalise across engines. The OpenAI channel never said “not open” on the same question; its failure modes were omission and misattribution — a different fix, not the same prescription.

3.2 What that “40%” actually is

The widely cited claim that “GEO improves visibility by 40%” comes from work by Princeton and others in 2023, published at KDD ’24 (arXiv:2311.09735, A, the only peer-reviewed source cited in this report). The figure requires three qualifications.

  1. It measures citation share, not discoverability. The study’s experimental pipeline retrieves the top 5 sources via Google Search, then has a model synthesise a cited answer from them. In other words, whether the content enters the candidate set is a given precondition; optimisation acts only on whether it is cited once already in the set. This does not conflict with the review in §3.1 — the two measure different segments of the funnel.
  2. The model is from the GPT-3.5 era. The generation stage used GPT-3.5-turbo. As of writing, we have found no independent replication of its principal effects on mainstream 2026 models. This is the most important known gap in this report.
  3. “40%” and “37%” are different measures. The 37% figure comes from a subjective-impression score judged by a large model on Perplexity, on a 200-case subset; the gain under the position-adjusted word count (PAWC) measure is 9–22%. Any citation must state which measure is meant.

3.3 The equaliser effect: cite it in pairs or not at all

The most commercially attractive finding in that study is the so-called equaliser effect: a site originally ranked 5th raised its AI visibility by +115.1% through the single optimisation of citing sources.

This is widely quoted, and almost never accompanied by its other half: under the same experimental condition (all sources optimised), the site originally ranked 1st fell by 30.3%.

This report’s rule: +115.1% never appears alone. Telling only the first half describes a zero-sum redistribution as though it were incremental creation. The honest statement is that GEO reallocates citation share within the candidate set, systematically favouring mid- and lower-ranked participants and disadvantaging the leader. That is genuinely good news for challenger brands, but the mechanism is redistribution, not growth.

Both halves of the equaliser effect: the site originally ranked 5th +115.1%, the site originally ranked 1st −30.3%
‘Originally ranked’ means the position in Google Search results when entering the candidate set (the study’s pipeline retrieves the top 5 sources via Google Search) — not a ‘rank position in AI’. The rise and the fall come from the same experimental condition: GEO reallocates citation share within the candidate set — redistribution, not growth. +115.1% never appears alone.

3.4 Keyword stuffing produces negative returns — and this figure is correct

Table 5 in §6 of that paper shows keyword stuffing performing roughly 10% below baseline on Perplexity (PAWC 21.9 against a baseline of 24.1, i.e. −9.1%).

During the writing of this report, an automated research pass proposed that “the original paper contains no negative value; this should be changed to ‘almost no improvement’.” That correction was rejected on review — the original table does give a negative value, and the proposed “correction” was itself wrong. We keep this on the record because it illustrates the judgement in the preface: in this category, error propagates in both directions, including towards the more conservative. Returning to the primary source item by item is the only reliable method.

Keyword stuffing lands roughly 10% below baseline: PAWC 21.9 vs 24.1 (i.e. −9.1%)
Table 5 in §6 of that paper (Perplexity). The negative value is really in the primary-source table — an automated research pass proposed changing it to ‘almost no improvement’ and was rejected on review: a ‘correction’ towards the more conservative reading is error propagation too.

3.5 So what does work

Stack the qualifications in §3.1 to §3.4 together and the defensible conclusions that remain are far shorter than a typical GEO checklist, and the centre of gravity is entirely different.

Discovery and citation must be treated as separate problems. A study of ChatGPT provides the sharpest contrast available (A−, preprint): when a query names the brand, the model’s recall of it is 99.4%; in unnamed category queries, the brand is spontaneously surfaced in only 3.32% of cases. In the same study, on-site GEO scores were uncorrelated with discovery rate; what correlated positively on Perplexity was referring domain count (r = +0.319) and Reddit presence (r = +0.395).

A separate study covering roughly 252,000 trials shows that format-level rewriting alone has essentially no effect.

Three conclusions follow.

  1. On-site optimisation governs whether retrieved content gets cited; off-site assets govern whether it is retrieved at all. The latter carries the greater weight, and the overwhelming majority of GEO services sell the former.
  2. Pure formatting work (subheadings, Q&A blocks, schema) is not sufficient to change outcomes. It is necessary hygiene, not leverage.
  3. Third-party mentions and referring domains are currently the most strongly correlated actionable variables. This shifts GEO’s centre of gravity from content production towards PR and community — which is why the ninety-day sequence in Chapter 7 places off-site authority ahead of on-site work.

The honest conclusion of this chapter: what GEO can reliably change is whether already-retrieved content gets cited, and its relative share within the candidate set. It has not been shown to stably change whether content is retrieved at all, nor to produce durable downstream traffic. Any promise beyond that scope has no evidence behind it today.


Named vs unnamed: 99.4% recall when the query names the brand, 3.32% surfaced when it does not
ChatGPT study (A−, preprint). Discovery and citation must be treated as separate problems: in the same study, on-site GEO scores were uncorrelated with discovery rate; what correlated positively on Perplexity was referring domain count (r = +0.319) and Reddit presence (r = +0.395) — while the overwhelming majority of GEO services sell on-site work.
This report is free to read in full. Want the PDF edition for forwarding and archiving? Get it at the whitepaper home →