How do I check my brand's visibility in AI search results?
CANLAH AI checks brand visibility in AI search by asking a buying question that never contains the brand name, logging every brand the answer names, then repeating that identical prompt in fresh logged-out sessions. Instability across repeated runs is the measurement, not a fault.
What does AI visibility actually mean?
The phrase covers two different things that are measured in different ways, and separating them is the first move of any honest self-check. Most DIY checks go wrong right here, before a single prompt has been typed.
Being named is measurable; being known is not
Visibility in AI search means one observable thing: whether a brand's name appears in the answer a stranger receives after asking a buying question. It is not whether the model has heard of the brand. Any large model will produce a plausible paragraph about almost any registered company when asked directly, because describing a name is a different task from selecting a recommendation. The check that matters therefore starts from a question a buyer would actually type, with the brand name absent from the prompt. If the name appears anyway, that is visibility. If it does not, the brand is invisible on that question, however fluently the model can talk about it when prompted by name.
The term is a description, not a demand signal
Worth saying plainly on a page that carries the phrase in its title: AI visibility is language the industry uses about itself. In the sample of query rewrites CANLAH AI observes, this exact phrasing did not appear at all — buyers describe the same problem in their own words, usually some version of whether ChatGPT mentions them. That is not an argument against the concept, which is real and measurable. It is an argument against building a content plan around the label, and the reason this page is written as a method anyone can run rather than as a pitch.
What exactly am I counting when I check?
Before running anything, fix the units. Named, cited and linked are different events inside the same answer and they move independently. Collapsing them into one visible-or-not flag is the most common way a self-check produces a result nobody can act on.
Keep named, cited and linked in separate columns
A brand can be named in prose with no link, cited as a source with a link but never recommended, or linked in a sources panel the answer body never mentions. Names in prose are what a reader remembers and repeats. Citations are what sends a click. Source links without a mention are the weakest of the set, and they usually come from a directory or a listicle rather than the brand's own domain. Each column implies different work afterwards, so recording them together destroys the only part of the check that leads anywhere.
One answer is a sample, not a position
Search rankings are stable enough that a single look is a fair reading. AI answers are not. Treat every answer as one draw from a distribution: the unit of measurement is the share of runs in which a brand is named across a fixed prompt set, and that share only exists once the same prompt has been run more than once. Anyone reporting AI visibility as a single position number is reporting something the surface does not produce. It is also why screenshots of AI answers circulate so easily and settle so little.
How unstable are AI answers between identical runs?
This is the property that governs the whole method, so it is stated here with its evidence rather than offered as a caution. CANLAH AI ran the audit against itself and kept every run.
Same prompt, same conditions, different answer
In our own self-audit, three rounds returned three different vendor sets. Nothing changed between them: identical wording, identical logged-out condition, same day — only the run. The competitor list a buyer sees was therefore not a property of the question but a property of the draw. Read that the way it deserves to be read. A single screenshot of an AI answer cannot establish that a brand is absent, and it cannot establish that a rival is present. It establishes what happened once, and treating it as a state of the market is the error the rest of this method exists to prevent.
Read the intersection, not the latest run
Once several runs exist, the useful reading is the set of names appearing in all of them, then the names appearing in some. Brands in the intersection are the engine's stable answer for that question. Brands appearing once are unproven, including yours — a one-run appearance is exactly the evidence a single check would have reported as fact. The variation itself is not a bug to fix or a reason to distrust the method. It is the property being measured, which is why the repeat run is not optional and why a report without run numbers cannot be audited.
How do I ask the first question properly?
The opening check is a single prompt and a note. It needs no tool, no upgraded account and no crawler. Run it in a fresh logged-out session so that what gets recorded is the answer a stranger receives.
Type what a buyer types, and leave your name out of it
Open a new logged-out session in ChatGPT, Gemini or Perplexity and type the question a buyer would type: the category, the constraint and the market — a service, a budget shape, and 'in Singapore'. Never include the brand name. Then list every brand the answer names, in order, spelled as the answer spells them. If the brand is on the list, note where. If it is not, note who is. Copy the prompt string into the log at this point, verbatim, rather than reconstructing it later from memory.
The competitive set is the real output here
Most brands open this check expecting a verdict about themselves and leave with something more useful: the list of companies the engine currently treats as the answer to the question. That set is frequently not the set on the search results page, and frequently not the set the sales team would have named. Before drawing any conclusion about your own placement, read the list as market intelligence. It names the competitors already installed in the answer, and those are the ones that have to be displaced for anything else to change.
Why run the identical prompt more than once?
The repeat is what turns a screenshot into a measurement. It is also the step most people skip, because the first answer already feels like a result and reading it again feels like wasted time.
Fresh session, identical wording, no rephrasing
Close the session and run the same prompt again in two further fresh sessions. Do not rephrase it. A rephrased prompt tests a different question and quietly merges two measurements into one unreadable result. Do not run the repeats inside the same conversation either: earlier turns stay in context and pull the later answers towards whatever has already been established. Number each run in the log. An unnumbered run cannot be checked for stability, and stability is the only reason the repeat exists.
Compare the lists, not the prose
The wording will differ every time and that difference carries no information. Compare only the brand lists and the source domains. Names present in every run are the stable answer; names present once are noise. If the runs agree exactly, record that too — some questions genuinely are settled, and knowing which ones are settled tells you where the expensive fights are. If they disagree, add runs rather than conclusions. The honest report of a disagreement between runs is 'not yet measured', which is an uncomfortable line to write and a cheap one to fix.
What do the sources under an answer tell me?
Where the answer shows sources, open them. This step costs a few minutes and reframes the work more often than the other two, because it moves the question away from ranking and towards trust.
Record the domain of every source shown
For each source, note what kind of page it is: the brand's own domain, a directory, an agency listicle, a marketplace, a forum thread, a news article. Do this across several answers rather than one, because the pattern is the finding and no single answer carries it. Keep the raw URLs. A source written into a note as 'some blog' is not recoverable later, and the pages carrying a category into AI answers are exactly the ones worth revisiting when the work starts.
Most brands are absent from the pages carrying their own category
The common result is that the sources feeding the answer are third-party lists nobody at the brand controls, and the brand's own domain appears nowhere among them. That single observation changes the question being asked. It stops being how to rank a page and becomes which pages this engine already trusts for the category, and whether the brand appears on them. It is inconvenient precisely because it cannot be fixed by editing a homepage, which is why the trace is worth running before any content work is commissioned.
What should I write down while checking?
A check that cannot be repeated next month is entertainment. The log is the deliverable: a small set of columns kept in one file rather than scattered through a chat history.
The minimum columns that survive a month
Date, engine, exact prompt string, run number, brands named in order, source domains shown, and whether the session was logged out. That is enough to re-run the same measurement later and compare like with like. The fields people skip are the exact prompt string and the run number, and those are the two that make a log worthless when missing: a prompt remembered approximately is a different prompt. Keep it in a spreadsheet. A chat history is a transcript, not a record, and it will read differently after the next model update.
Define 'named' before you look at results
Write the rule first, or it will bend towards good news. A workable definition: the brand name appears in the answer body as one of the options offered, spelled correctly, without the user having asked for it. A passing mention inside a quoted source title does not count. A name that appears only after a follow-up question does not count, because the follow-up came from you. A brand listed in a sources panel with no mention in the body goes in the sources column. Ambiguity resolved afterwards is how a self-check becomes a self-portrait.
How do I keep the result defensible?
Findings get disputed — internally, by a competitor, or by an agency defending a report of its own. What survives that is the raw capture, not the summary written beside it.
Screenshot the answer, not your reading of it
Models are updated without notice and answers change between runs, so a paraphrase in a note cannot be re-examined later. Capture the full answer including the sources panel, store it beside the log row, and name the file with the run number. This feels excessive during the first check and stops feeling excessive the first time somebody asks whether the answer really said that. A check that produces only conclusions is indistinguishable from an opinion, and it will be treated as one.
Have someone try to break the finding
Before a result travels further than the person who ran it, hand the log to somebody whose job is to argue with it: wrong prompt, contaminated session, too few runs, market word missing, definition of 'named' quietly widened. Most claims survive that treatment, and the ones that do not were going to fail in front of a client instead. CANLAH AI puts its published citation evidence through the same adversarial pass, and the /our AEO versus GEO guide reports what survived it. A finding nobody has attacked has not been tested.
Why doesn't asking an AI 'do you know my brand?' work?
The self-check has a small number of ways to produce a confident wrong answer, and the most common one is the prompt people reach for first.
Asking about your own brand is a recall test
Type a brand name into any assistant and it will produce a paragraph about it. That paragraph proves the model can generate text about a name it has encountered. It proves nothing about whether the brand is ever offered to somebody who did not type that name. The two tasks are unrelated inside the system, and the prompt that feels like the obvious test is the one certain to return a reassuring answer. The rule is blunt: if the brand name is in the prompt, the test is void. Delete the result and start from the buyer's question instead.
A shared chat window contaminates everything after it
Running later checks inside one conversation carries earlier turns into the input, so a brand mentioned once keeps reappearing and the answer set narrows towards whatever has already been established. The same applies to checking a competitor immediately after checking yourself: the competitor names you typed are now part of the context for everything that follows. When results look suspiciously stable across runs, contaminated context is the first thing to rule out, before concluding that the engine holds a firm view of the category.
Does it matter which account and wording I use?
Two settings decide whether the check measures a market or measures you. Both are free to get right, and both are silently wrong by default.
Your signed-in account is not a stranger
Assistants adapt to account history, saved memories and location. Checking from your own signed-in account measures what the engine shows to somebody who has spent months asking about your industry and your company — the least representative reader available. Use a logged-out or private session and record the setting in the log. If the brand appears signed in and disappears signed out, that gap is a finding about the sample rather than about the market, and reporting the signed-in version as a market reading is the quiet form of marking your own homework.
Say the market, and pitch the category at the right level
Answers to 'best supplier' and 'best supplier in Singapore' are frequently unrelated brand sets. Locale words do heavy work in these systems, and dropping them is how a check ends up measuring a market the brand does not sell into. The category term behaves the same way: one level too broad returns generalists and pulls the whole reading out of alignment with the buyers you want. Record the prompt verbatim, market word included. Precision in the prompt is not pedantry — it is the only variable fully under your control.
Do I need to check every AI engine?
Coverage gets over-thought. In most categories the assistants agree more than they differ, and the split that actually matters is not the split between one assistant and another.
Assistants and AI Overviews are genuinely different surfaces
A conversational assistant answers a typed question inside a chat window. An AI Overview sits above the Google results page and is triggered by a search query. They draw on different retrieval paths and frequently name different companies for the same intent, so a check covering only one leaves the other entirely unmeasured. Given the budget for two surfaces, make them one assistant and Google's overview rather than two assistants. That pairing covers the two behaviours buyers actually have — asking and searching — instead of covering one behaviour twice.
Adding assistants buys less than adding runs
Running the same category question through several assistants usually returns overlapping brand sets with different ordering and different prose. Most of the time there is no difference worth reporting, and that overlap is a finding rather than a flaw in the method. Spend the extra effort on repeating one prompt instead: repetition exposes instability, while breadth mostly exposes rewording. Add an engine when that engine matters to your buyers — because their procurement team lives inside it, for instance — not for completeness on a slide.
When does a manual check stop being enough?
The self-check is honest and cheap, and it has a ceiling. Saying where that ceiling sits is what makes a DIY result trustworthy rather than overclaimed, and it is the part most tool dashboards leave out.
A handful of prompts cannot carry a decision
Repeating one question tells you whether that question is stable. It cannot tell you whether a brand is visible across a category, because a category is many questions with different intents behind them. A CANLAH AI audit is sized for exactly that: 36 questions × 2 engines × 2 rounds. The questions carry the intent range, the engines carry the split between surfaces, and the rounds carry the instability — that shape exists because single-question checks kept producing contradictory conclusions from one week to the next. Nobody needs that scale to begin. What everybody needs is to stop short of announcing 'we are invisible in AI' on one prompt and one afternoon.
Some of what decides visibility never appears in the answer
Part of the machine-readable layer around a brand is issued by its platform rather than written by the brand, and nothing in a chat answer reveals which is which; the /our agentic SEO guide carries the store-level evidence for that. A self-check tells you the outcome, not the layer that produced it. Reading a cause into a result is where DIY analysis usually goes wrong, and the discipline is to keep the check descriptive and hold the diagnosis until there is evidence to support it.
What do I do with the result once I have it?
The check returns a state, not a plan. What happens next depends on how the result is read, and the most common reading is the one brands least expect when they open the first answer.
Read an absence as information, not as failure
Absence on a question is normal and worth circulating internally exactly as recorded. CANLAH AI was named 0/5 on a pure SEO query in its own audit, and the /our AI SEO guide sets out what that result means and why it was published anyway. A zero says the engine holds an established set for that phrase, and that displacing it is a different project from being named on questions where a brand already appears. Run your own zeros before deciding which questions are worth contesting, because the cheapest win is usually a question nobody has claimed.
Separate a page problem from a category problem
If the citation trace shows third-party lists carrying your category while your own domain never appears among the sources, the gap sits upstream of the website. If the domain does appear as a source but the answer never names you in the body, the gap is in how that page states what the business does and for whom. Those findings lead to entirely different work, and the source column is what tells them apart. If the result convinces you outside help is warranted, read the published Singapore retainer bands on the /our AEO versus GEO guide first, then judge providers on method: how many questions, how many engines and how many repeat runs sit behind the claim.