Skip to main content
CANLAH AI
Try
DO IT YOURSELF — OR WORK WITH US

Doing GEO In-House vs Hiring an Agency: An Honest Comparison

Most teams should start in-house: checking your own brand name in ChatGPT once a month costs nothing and tells you something real. Hire an agency when the question widens beyond your brand name — because at that point a single ask is not evidence, and disciplined repeat-sampling, account isolation and evidence archiving become full-time work.

This page is written for the marketing lead who has already typed the company name into ChatGPT and is trying to work out what to do with the result. It argues both sides. If your situation is in the do-it-yourself section, do it yourself — we would rather lose the enquiry than sell you a retainer you do not need.

01 — THE SAMPLING TRAP

A Single Check Is Not a Finding

Almost every in-house AI-visibility effort starts the same way: someone opens an assistant, asks a question a customer might ask, and screenshots the answer. It is a reasonable instinct, and it is also the single most common way teams reach a confident conclusion that is not true.

The problem is not the question. It is that one answer carries almost no information about the underlying state, because generative systems do not return a fixed result. Three measurements make this concrete:

  1. 01
    Ask the same question twice, get the same answer less than 1% of the time.

    Not a paraphrase — an identical answer. Generative systems sample from a distribution on every call, so variation is the default condition rather than a malfunction. A screenshot captures one draw from that distribution.

  2. 02
    Repeat the same prompt and the set of brands named overlaps only about 45–59%.

    Roughly half the brand list changes between runs of the identical prompt. So the difference between “we were in there” and “we are gone” can be produced by nothing at all — no algorithm update, no competitor campaign, no change on your site.

  3. 03
    Which means a single check cannot support a decision.

    “The answer changed” is not a finding. “We went from zero out of five mentions to four out of five, measured under the same protocol both times” is a finding. The unit of evidence is a frequency across controlled repeats, never one screenshot.

This is why the yesterday-versus-today comparison your team has probably already run cannot be trusted in either direction. The reassuring reading and the alarming reading are both inside the noise band. Everything else on this page is about what it takes to get outside it — and whether that is worth paying for. If you want the underlying mechanics of how AI answers get assembled in the first place, our explainer on why brands go missing from ChatGPT covers it.

02 — DO NOT HIRE US

When You Should Do This In-House

There are four situations where hiring us would be a poor use of your budget. We would rather say so on a public page than discover it in month two of an engagement.

You only want to know how your brand name performs.

Branded queries are the easiest case and the least volatile. Ask an assistant about your company once a month, read the description it produces, and check it for factual errors. This is genuinely useful, takes minutes, and does not need an agency. Fix what is wrong on your own site and in your own listings first.

One market, one language.

Multi-engine, multi-market coverage is where the operational cost lives — different retrieval behaviour, different regional answers, different phrasings per language. If your buyers are all in one market asking in one language, a narrow in-house habit covers a surprising amount of it.

You have engineering capacity and will own the hygiene long-term.

The protocol on this page is not secret. If you have people who will build it, run it on a schedule, keep the sampling honest when results are inconvenient, and maintain it for a year rather than a quarter, you should build it. The failure mode is not capability — it is the fourth month, when the person who built it is on another project and nobody re-runs the baseline.

Your site fundamentals are still broken.

If key pages are not indexed, load slowly, or say nothing a machine can extract, AI visibility work is premature. Engines cannot cite what they cannot retrieve. Spend the budget on the foundation, then measure. We will tell you this during a snapshot rather than after you have signed something.

The honest summary: an agency is worth paying for when the measurement itself becomes the hard part — many intents, several engines, more than one market, and a need for the numbers to survive a sceptical question from someone who was not in the room. Below that threshold, a disciplined monthly habit run by your own team beats an expensive one run by ours.

03 — PROBE HYGIENE

The Discipline That Makes AI Measurement Mean Something

If you decide to build this in-house, take this section as the specification. None of it is proprietary — it is simply what has to be true before a number is worth reporting to a board. It is also, in our experience, the part that quietly decays first when nobody owns it.

01
Two samples per intent — because we measured how many it takes.

One run is an anecdote, so we ran the experiment on our own archive: across 503 stored responses, the mention rate at one, two, three and eight rounds was identical. Mention rate settles almost immediately, so we sample twice — the second run is a failure buffer, not theatre. Source sets are the opposite: the eighth round was still turning up sources we had not seen, which is why we do not treat API-side source lists as settled. Below twenty-five samples we report ranges, never a point estimate.

02
Lock the intent, not the sentence.

Each monitored intent carries five to eight semantically equivalent phrasings — word order, synonyms, question versus instruction. Runs draw from that pool and the pool is refreshed monthly, so results describe the intent rather than one wording that happened to flatter us.

03
At least six hours between repeat samples of the same intent.

Shorter gaps risk reading a cached response rather than a fresh one. Scheduling also carries roughly twenty percent random jitter, so probes never fire on the hour in a recognisable pattern.

04
Stay under half of every published rate limit.

Throttling is the most direct source of silent data corruption: a rate-limited call is not a null result, it is a wrong one. We plan volume against the provider's own limits rather than against how much we would like to collect.

05
Treat one query per prompt, per model, per day as the ceiling.

That is the line the serious tools in this category converge on, and going above it buys nothing. Higher frequency does not reduce the randomness — that is what repeat sampling across days is for — it only makes the traffic pattern more distinctive.

06
A CAPTCHA stops the channel automatically.

Not an alert. A stop: the channel halts, the remaining runs in that period are cancelled, and a human reviews before anything resumes. No auto-retry, and never any attempt to bypass the challenge. A tripwire that only notifies is not a control.

One consequence worth stating plainly: this protocol reports frequencies and relative share, never a single-run ranking. Position is the most volatile thing an AI answer produces, so a report built on it looks precise and means nothing. The same reasoning shapes our full GEO methodology.

04 — RED LINES

Three Ways In-House Probing Goes Wrong

These are the failures we see most often, and the first two are the reason some teams stop measuring altogether after a bad week. None of them are subtle once named — but all three are the natural thing to do when a marketing lead is asked for AI-visibility numbers by Friday.

Firing the same question a hundred times in one burst.

This is the pattern that most closely resembles the behaviour platforms publicly act against — and it does not even produce better data, because samples clustered in one minute share the same cache state and the same conditions. Spacing is what buys independence.

Probing from your company's main account.

The most expensive mistake available to an in-house team. Automated querying against a logged-in surface can take the whole account down with it, and account loss is not reversible on your schedule. Probing must run on dedicated identities that exist for nothing else, so the blast radius of a bad day is one throwaway account.

Pointing a cloud headless browser at an AI surface.

A datacentre IP plus a headless fingerprint is identifiable on the first request. Whatever comes back is not what a buyer in your market would see, so the measurement is wrong before any compliance question arises. This is why we run the probe layer on real machines with real browsers.

The pattern behind all three: low volume alone does not keep you safe. What gets a channel cut off is a recognisable behavioural fingerprint — bursts, headless signatures, datacentre origins — not the raw number of queries. An in-house setup that only budgets for frequency, and not for how the traffic looks, is exposed regardless of how modest the volume is.

05 — HOW WE RUN IT

Real Machines, Real Browsers, Physically Separated

Our probe layer does not run in a container. It runs on ordinary computers with ordinary browsers, because that is the only configuration that returns what a buyer would actually see. Each machine is set up the same way: one browser for the people who use the machine, and a completely separate browser instance for probing, signed in to a dedicated identity that exists for nothing else. The two never share a profile, a session, or a window list.

That separation is not a nicety. Automation that reaches across a profile boundary reads whatever account happens to be in front — which contaminates the measurement with someone else's personalised answers and exposes data that was never ours to collect. Keeping the boundary physical rather than procedural is what makes unattended operation acceptable at all.

Running a distributed set of machines rather than one busy server has a second effect that matters more than it sounds: the more machines carry the work, the lower each individual machine's request rate becomes, and the further every one of them sits from any behavioural threshold. Capacity here is bought by adding separation, not by turning up frequency — which is the opposite of how an in-house script under deadline tends to evolve.

The output of all this is deliberately unglamorous: a raw evidence package. Prompt text, engine and version, region, sampling timestamp, and the full response, for every probe, handed to you. You can re-run the same intents and check whether our ranges hold. That reproducibility is the actual product — see how it reads in practice in our engagement write-ups.

06 — SIDE BY SIDE

In-House vs Agency, Dimension by Dimension

Read this as a specification you could hand to your own team, not as a scorecard. Every row on the right is something an in-house effort can do — the question is whether it will still be doing all of them in month nine.

Dimension Typical in-house Working with us
Sampling One ask, one answer Five or more runs per intent, per snapshot
Wording Whatever you typed that day 5–8 equivalent phrasings, rotated per run
Account risk Usually a company account Dedicated probe identities, isolated
Evidence Screenshots in a chat thread Prompt, engine, region, timestamp, full response
Engine coverage Whichever tab is open Multiple engines, one protocol
Reproducibility Not repeatable next month You can re-run the pool yourself
Reporting unit “We showed up” / “We did not” Frequency ranges against baseline
Execution Findings become a backlog Preview page, you click confirm
Cost shape Staff time, indefinitely Fixed monthly scope

Note what is absent from the right-hand column: any promise about where you will rank. We commit to deliverables, cadence and a protocol that does not change between reports — not to outcomes nobody controls. How that translates into engagement shapes is set out on our pricing page, and the questions worth asking any vendor in this category are in our guide to choosing a GEO agency in Singapore.

07 — WHAT CHANGES HANDS

Measurement Is Half the Job. Someone Still Has to Ship the Changes.

This is the part in-house efforts underestimate most. Measurement produces a list of things to fix, and that list then sits in a backlog behind the quarter's roadmap. The measurement was never the bottleneck — getting the pages written, structured, reviewed and published was.

So our delivery is built to reduce your side of that work to two actions: look at a preview, and click confirm. When we produce or optimise a page, you review the real rendered page at a live preview link — your layout, your site, scrollable and clickable — not a document describing what we plan to publish. Nothing goes live until you approve that exact version, and approval is bound to that version, so any later change comes back to you before it ships.

We also run this programme on ourselves and publish the result. Our own July 2026 self-audit baseline showed canlah.ai cited in zero of our own commercial-intent queries — a blunt starting point we chose to publish rather than bury, because it is the same “before” picture every engagement begins with. Publishing a number like that is uncomfortable for the obvious reason, and it is also the cheapest way to demonstrate that the pipeline is real: it takes a working probe stack, a willingness to show an unflattering figure, and confidence that it will move.

If you take one thing from this page and never contact us, take the protocol in section 03. A team measuring honestly with repeat samples and rotated phrasings is in better shape than a team paying an agency for a dashboard with no archive behind it.

FAQ

Frequently Asked Questions

Can we just check ChatGPT ourselves every month?

For your own brand name, yes — and we recommend it. Type your company name into an AI assistant, read what comes back, and note it. That habit costs nothing and it catches obvious problems: a wrong description, a stale address, a competitor named where you should be. What a manual check cannot do is tell you whether a change you observed is real. The same question asked twice returns an identical answer less than 1% of the time, and across repeated runs of the same prompt the sets of brands named overlap only about 45–59%. So a month-over-month difference from one ask per month is inside the noise. Once your question widens past your own name — categories, comparisons, buying-intent phrasings — the manual check stops being informative.

What does disciplined sampling actually mean in practice?

Five things, all boring, all load-bearing. First, never a single run — we sample each intent twice per snapshot, a number we arrived at by testing rather than by picking: on 503 archived responses the mention rate was identical at one, two, three and eight rounds, so extra rounds buy nothing there. The same test showed source sets never settle — the eighth round still found new ones — which is why source claims come from a real browser instead of more API rounds. Second, we lock the intent rather than the sentence: each monitored intent carries five to eight semantically equivalent phrasings, rotated between runs and refreshed monthly, so a result reflects the intent rather than one lucky wording. Third, repeat samples of the same intent are spaced at least six hours apart, with roughly twenty percent random jitter on scheduling, so nothing fires on the hour. Fourth, request volume stays well under half of each provider's published rate limit. Fifth, results are reported as ranges — appearance frequency, citation frequency, relative share — never as a single-run ranking.

Why not just run this on a cloud server with a headless browser?

Because it is the fastest way to get identified and cut off. A datacentre IP combined with a headless browser fingerprint is recognisable on the first request, which means the answers you collect are not the answers a real buyer would see — and that is a data quality problem before it is a compliance problem. The related mistake is running probes from a company account. Automated querying against a logged-in surface can take the whole account with it, and there is no way to un-ring that bell. Probing should never touch an account your business depends on.

What happens when a probe hits a CAPTCHA?

The channel stops immediately and automatically, and the remaining scheduled runs for that period are cancelled pending human review. We do not auto-retry, and we never attempt to bypass a CAPTCHA under any circumstances. This is a hard stop written into the pipeline rather than an alert someone might notice — an alert-only tripwire is not a control, because an unattended process will simply keep digging. If your in-house setup does not have an equivalent hard stop, it will eventually cross a line nobody decided to cross.

If we hire you, can we verify the numbers ourselves?

That is the design goal. You receive the raw evidence for every probe: the exact prompt text, the engine and version, the region, the sampling timestamp, and the full response. With that package, you or a rival agency can re-run the same intents and check whether our reported ranges hold. We would rather hand you the means to falsify our reporting than ask you to trust a dashboard. It also means the relationship survives a change of marketing lead — the evidence belongs to you, not to us.

Can you guarantee our brand will appear in AI answers?

No, and any agency that guarantees an AI ranking is a red flag you should act on. These systems are non-deterministic and their retrieval behaviour changes without notice, so nobody controls placement. What can be committed to is different and checkable: named deliverables, a fixed measurement cadence, a protocol that does not change between reports, and an evidence archive you own. If a vendor will not put the protocol in writing but will guarantee an outcome, they have the incentives exactly backwards.

How much work does this create for our team?

Two actions, repeatedly: look at a preview, and click confirm. When we produce or optimise a page, you review the real rendered page in your browser — the actual layout, on your actual site, at a live preview link — rather than a document describing what we intend to publish. Nothing goes live until you approve that specific version. Approval is tied to the exact version reviewed, so if anything changes afterwards it goes back to you. Beyond that, we ask for a quarterly session to review whether the intents we are monitoring still match how your buyers actually ask.

Find Out Whether You Have a Measurement Problem or a Visibility Problem

Request a free 48-hour AI-visibility snapshot. We probe a starter set of your real buyer intents under the protocol on this page and send you the findings with the raw records attached — so you can check every claim, including the ones that do not favour us.