Chapter 3 · What GEO Can Actually Change, and What It Cannot

Marketing in the Agent Era · Canlah AI · a Singapore SEO + GEO agency

3.1 The strongest counter-argument first

On 15 July 2026, a critical review covering 45 GEO studies reached this conclusion: no technique under review demonstrated a stable, longitudinal, cross-platform causal effect on organic discoverability or on downstream behaviour (arXiv:2607.14035, A−, single-author preprint).

We put that sentence at the top of this chapter rather than in a footnote, for three reasons.

First, any buyer with judgement will find it within a single search. Having it produced against you, and producing it yourself and then answering it, are two entirely different positions to argue from.

Second, it is substantially correct. Below we set out the one point on which we think it needs qualifying — but that is a scope qualification, not a rebuttal.

Third, the existence of that review is itself the justification for the measurement standard set out in Chapter 4 of this report. If the causal effect has not yet been stably demonstrated, then continuing to sell a “ranking improvement” product is dishonest.

So, with causality unproven, what is left that can responsibly be delivered? Three things, and all three are verifiable:

  1. Factual errors can be found and corrected. False statements about your brand in AI answers — that you are not open, that you are some other company with a similar name, that you are associated with third-party content you do not control — are objective factual errors, and judging them right or wrong requires no causal theory at all. In a real case we observed the API layer stably returning “this outlet is not open” while the browser layer, over the same period, returned the correct information (4.6 stars, 541 reviews). Fixing and verifying this class of problem does not depend on whether “GEO works”.
  2. Gaps are measurable, and they are not probabilistic. Whether your product page carries structured data is a binary fact, not a correlational inference. Every cell in the table in Chapter 0 can be reproduced on your own site.
  3. We have not been able to prove causality, so we use pre-registered controls instead. On any ongoing engagement we publish the prompt panel and the decision rules before we start, we hold out a control group that receives no optimisation, and we report the results in full to the client whether or not they meet the target. This is not a substitute for causal proof, but it turns “did the money you spent change anything” from a vendor’s assertion into a criterion both parties agreed in advance.

The honest product definition therefore is: we sell diagnosis, error correction and controlled experiments; we do not sell ranking promises. That scope is narrower than the industry’s customary claims, but every item inside it can be verified by you at the point of delivery.

3.2 What that “40%” actually is

The widely cited claim that “GEO improves visibility by 40%” comes from work by Princeton and others in 2023, published at KDD ’24 (arXiv:2311.09735, A, the only peer-reviewed source cited in this report). The figure requires three qualifications.

  1. It measures citation share, not discoverability. The study’s experimental pipeline retrieves the top 5 sources via Google Search, then has a model synthesise a cited answer from them. In other words, whether the content enters the candidate set is a given precondition; optimisation acts only on whether it is cited once already in the set. This does not conflict with the review in §3.1 — the two measure different segments of the funnel.
  2. The model is from the GPT-3.5 era. The generation stage used GPT-3.5-turbo. As of writing, we have found no independent replication of its principal effects on mainstream 2026 models. This is the most important known gap in this report.
  3. “40%” and “37%” are different measures. The 37% figure comes from a subjective-impression score judged by a large model on Perplexity, on a 200-case subset; the gain under the position-adjusted word count (PAWC) measure is 9–22%. Any citation must state which measure is meant.

3.3 The equaliser effect: cite it in pairs or not at all

The most commercially attractive finding in that study is the so-called equaliser effect: a site originally ranked 5th raised its AI visibility by +115.1% through the single optimisation of citing sources.

This is widely quoted, and almost never accompanied by its other half: under the same experimental condition (all sources optimised), the site originally ranked 1st fell by 30.3%.

This report’s rule: +115.1% never appears alone. Telling only the first half describes a zero-sum redistribution as though it were incremental creation. The honest statement is that GEO reallocates citation share within the candidate set, systematically favouring mid- and lower-ranked participants and disadvantaging the leader. That is genuinely good news for challenger brands, but the mechanism is redistribution, not growth.

3.4 Keyword stuffing produces negative returns — and this figure is correct

Table 5 in §6 of that paper shows keyword stuffing performing roughly 10% below baseline on Perplexity (PAWC 21.9 against a baseline of 24.1, i.e. −9.1%).

During the writing of this report, an automated research pass proposed that “the original paper contains no negative value; this should be changed to ‘almost no improvement’.” That correction was rejected on review — the original table does give a negative value, and the proposed “correction” was itself wrong. We keep this on the record because it illustrates the judgement in the preface: in this category, error propagates in both directions, including towards the more conservative. Returning to the primary source item by item is the only reliable method.

3.5 So what does work

Stack the qualifications in §3.1 to §3.4 together and the defensible conclusions that remain are far shorter than a typical GEO checklist, and the centre of gravity is entirely different.

Discovery and citation must be treated as separate problems. A study of ChatGPT provides the sharpest contrast available (A−, preprint): when a query names the brand, the model’s recall of it is 99.4%; in unnamed category queries, the brand is spontaneously surfaced in only 3.32% of cases. In the same study, on-site GEO scores were uncorrelated with discovery rate; what correlated positively on Perplexity was referring domain count (r = +0.319) and Reddit presence (r = +0.395).

A separate study covering roughly 252,000 trials shows that format-level rewriting alone has essentially no effect.

Three conclusions follow.

  1. On-site optimisation governs whether retrieved content gets cited; off-site assets govern whether it is retrieved at all. The latter carries the greater weight, and the overwhelming majority of GEO services sell the former.
  2. Pure formatting work (subheadings, Q&A blocks, schema) is not sufficient to change outcomes. It is necessary hygiene, not leverage.
  3. Third-party mentions and referring domains are currently the most strongly correlated actionable variables. This shifts GEO’s centre of gravity from content production towards PR and community — which is why the ninety-day sequence in Chapter 7 places off-site authority ahead of on-site work.

The honest conclusion of this chapter: what GEO can reliably change is whether already-retrieved content gets cited, and its relative share within the candidate set. It has not been shown to stably change whether content is retrieved at all, nor to produce durable downstream traffic. Any promise beyond that scope has no evidence behind it today.


This report is free to read in full. Want the PDF edition for forwarding and archiving? Leave your email at /whitepaper.