AI Visibility

How to Audit Brand Mentions in AI Engines

A brand audit for LLMs is four passes over one question list: build the list, run it on every engine, check the crawlers can reach you, then read the pages that got cited instead of you. With the numbers from our own audits, including the ones that embarrassed us.

Telman GadimovFounder, CueScout8 min read

An AI visibility audit is four passes over the same list of questions: build the list, run it on every engine that matters to you, check that the crawlers can read your site at all, then read the pages that got cited in your place. Only the fourth pass produces work. The first three tell you where you stand and how wrong your assumptions were.

I have run this on cuescout.com and on customer domains often enough to know where it goes sideways, and it is almost always the first pass.

What the audit is actually counting

Two numbers come out of the exercise, and they answer different questions.

The first is how often you get named. Across fifty questions, in how many answers does your brand appear at all. That is your share of voice, and it is the number you put in a slide.

The second is which pages the engine read to write those answers. Almost nobody collects it, because the answer text is what fills the screen and the source panel takes another click. It is the more useful of the two. A grounded assistant does not hold opinions about vendors; it runs a search, pulls back a handful of pages, and summarises what those pages say between them. The names in the answer are the names in the retrieved pages. So the citation list is the audit, and your share of voice is a symptom of it.

Collect both. If you only have time for one, collect the URLs.

Step 1: build the question list, then clean it

Search Console is the right starting point, because it tells you which phrasings your buyers already use. Filter to queries that start with best, how to, alternative, vs, or for, sort by impressions, and export the top fifty. Those are your prompts.

Two things to do before you trust a single row of that export.

Filter out -site: first. Over three months, 87 of 777 impressions on cuescout.com came from queries stuffed with a dozen exclusion operators, which are brand-monitoring tools scraping Google rather than people searching. Zero clicks, average position 6.9, while the rest of the site sat at 53. I wrote that up in 11% of my Search Console impressions are competitor-monitoring bots, including how to isolate them in about a minute. Build a prompt list off unfiltered data and you will spend a week auditing queries no buyer has ever typed.

Then check who the clicks belong to. When I read our own window on 16 September, 20 of the 28 clicks came from Azerbaijan, which is where I sit. So the site had 28 clicks and eight of them were strangers. That is a small-site problem, but under a few hundred clicks a month you have it too, and it changes which queries look like they matter.

The bigger caveat is structural. Your Search Console queries are your demand, not your citation surface. In one customer audit we pulled every page the engines had cited across their question set and checked each one against Google's top 100 for the query that surfaced it. 58 of 96 cited pages were not in the top 100 at all. Not ranking badly, absent. Search Console tells you which questions are worth money. It has very little to say about which pages win them, which is the whole reason this audit exists as a separate job from an SEO audit.

If you have no Search Console data worth exporting, write twenty questions by hand the way a buyer would type them, with the qualifier attached: "best X for Y team size", "X vs Y", "X alternative for Z". Vague category questions return vague category answers and both are useless.

Step 2: run every question on every engine you sell into

The manual version is exactly as dull as it sounds. Paste the question, record whether you were named, record every cited URL, move on. Roughly two minutes per question per engine once you have a rhythm, so fifty questions across four engines is a day you will not enjoy.

Do not shortcut it down to one engine. On 20 August we put 36 questions to Grok, ChatGPT and Perplexity and kept all 740 citations. Grok and Perplexity shared 74 cited domains. Grok and ChatGPT shared eight. Six domains were cited by all three. Perplexity cited Reddit 38 times out of its 346 citations, and ChatGPT cited it zero times out of 139. The full breakdown is in what Grok actually cites, but the short version is that an audit of one engine is an audit of one engine.

For B2B software the set worth checking is ChatGPT, Perplexity, Gemini, and Google's two answer surfaces, AI Overviews and AI Mode. Add Grok only if your category turns on current sentiment, since on that 36-question run it cited X posts on sentiment questions and never once on a buying question.

One discipline that decides whether your data is real: record whether the answer was grounded. A model answering from memory will happily produce a list of sources, and they will be URLs it half-remembers rather than pages it read. The tell is bare homepages with no paths. If there was no visible search step and no source panel, log the brand names and throw the URLs away.

Step 3: check that the crawlers can reach you

Low visibility sometimes has a boring cause. Open your robots.txt and look for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended, because plenty of sites are still running a template that blocked them in 2024 and nobody has read the file since. OpenAI documents its crawlers by user agent, and the directives are ordinary robots.txt syntax.

Google-Extended is the one people get wrong. It governs Gemini, not AI Overviews. AI Overviews are assembled from Google's normal index, so blocking Google-Extended will not remove you from them and will remove you from Gemini.

After the crawl directives, three things decide whether a page that is reachable is also usable: whether your content renders without JavaScript, whether you publish an llms.txt, and whether your answers carry FAQPage markup that can be extracted as question and answer pairs rather than parsed out of prose. Our AI crawler checker reads the robots side and the readiness checker runs the rest, both without a signup.

Worth admitting: the check our own site failed for weeks was the sameAs list on our Organization schema, which is about the least glamorous fix available. Run the audit on yourself before you run it for anyone else.

Step 4: read the cited pages, not the answer

This is the pass that turns an audit into a plan.

For every question where you were missing, open the pages the engine did cite and sort each one into one of three buckets. It is a third-party page you could plausibly be added to. It is a competitor's own page, which you cannot join and have to displace with an equivalent of your own. Or it is one of your pages, which got read and still did not produce your name, which usually means the answer to the question is buried three paragraphs into something else.

The good news is how short the resulting list is. We classified 533 cited URLs across 360 domains from 50 buying questions in our own category, and 288 of those domains appeared exactly once and never again. Thirty pages out of 533 carried about a fifth of all citation weight. The full corpus study has the dataset attached. A year of blogging on your own domain competes with the other 483 pages; getting named on twenty specific third-party pages competes for the fifth.

If the audit ends with a ranked list of pages and an owner for each, it worked. If it ends with a percentage, it did not.

When a spreadsheet is enough

SpreadsheetSoftware
Good forUnder ~20 questions, checked quarterly50+ questions, several engines, weekly
First passAbout a day, and you learn how the engines answerMinutes
Every pass afterThe same day, againNothing
CostFreePaid

The first audit is genuinely worth doing by hand, because you cannot interpret a visibility chart until you have watched a few answers get assembled. The problem is never the first pass. It is the fourth, when you are re-running fifty questions across five surfaces for the third month and have 250 answers to log by hand, and the run stops happening.

What we run instead

CueScout does the four passes on a schedule. It pulls your question set from Search Console demand rather than from a keyword tool, checks ChatGPT, Perplexity and Gemini through their search-grounded APIs, and reads Google AI Overviews and AI Mode off Google's own results page. Every prompt is checked at least once a week, and one whose answers keep moving gets checked every three days. Claude is a paid add-on and Grok is sold as a pack of twenty prompts, for the sentiment reason above.

Each run stores the cited URLs, not just whether you were named, which is what makes the fourth pass possible without a spreadsheet. The site readiness side covers robots.txt, JavaScript rendering, llms.txt and FAQPage markup, and the output is a queue of specific pages to write or pitch rather than a score.

The first check is free and needs no account: put one real buying question through the AI visibility checker, or go straight to the sources with the citation source radar. If you want the mechanics behind why any of this moves an answer, how AI engines choose brands to recommend is the longer read, and how to get your brand into AI recommendations is what to do with the list once you have it.

Frequently asked questions

How often should I check my AI citations?

Monthly is enough for a manual audit of brand terms, and weekly is the point where you would want something running it for you. Grounded answers are re-retrieved on every run, so two identical questions asked three days apart can return different sources. That means a single audit is a snapshot with error bars, not a baseline, and the thing you are actually watching is whether the same hosts keep reappearing.

Does blocking AI crawlers improve search rankings?

No. Blocking GPTBot or PerplexityBot has no effect on how Googlebot ranks you, and it removes you from the retrieval pool the answer engines draw on. The one that gets confused is Google-Extended, which governs Gemini, not AI Overviews. AI Overviews are built on Google's ordinary index, so you cannot opt out of them with a crawler directive while staying in normal search.

Which AI engine drives the most B2B referral traffic?

Perplexity and ChatGPT are the two that show up in referrer logs, because both put visible source links next to the answer. But traffic is the wrong thing to audit for. Most of an AI answer's influence is the recommendation itself, which a buyer acts on without ever clicking, so measure whether you are named and cited rather than whether anyone arrived.

Can I pay to appear in AI model answers?

Not in the synthesized answer itself. Some engines run separately labelled ad slots, but the brand names inside the text come from whatever the model retrieved or memorised. You get in by being described on pages the engine reads, which is why an audit that records cited URLs is more useful than one that only records your share of voice.

Related cluster

Keep reading

Generative Engine Optimization guide

Find the questions worth writing about

CueScout scans Reddit, Hacker News, and Quora for the buyer questions AI answers are built from, explains why each one matched, and turns the ones that keep repeating into pages to publish on your own site. Nothing gets posted anywhere else.

Start your first scan