AI Visibility

The 533 URLs AI Reads About One Category: A Citation Corpus Study

We pulled every source three AI engines cited across 50 buying questions in the AI-visibility category, classified all 533 URLs, then fetched the top 30 pages to see who was named on them. Listing presence tracked visibility almost one to one. Full dataset included.

Telman GadimovFounder, CueScout8 min read

Between 9 and 15 August 2026 I paid $125 for a month of Searchable, a competitor's AI-visibility tool, mostly to see what its product did. The useful part turned out to be an export I could have got nowhere else: every source URL that three AI engines cited while answering 50 buying questions about AI-visibility tools. 533 URLs, 360 domains, each with a usage frequency.

That is a map of what the models in one category are actually reading. I have not seen anyone publish one, so here is ours, along with the file itself.

Download the dataset (CSV, 533 rows), with rank, URL, source type, content format and usage share. Nothing removed.

How the data was collected, and by whom

The collection is not mine. Searchable ran 50 prompts against its base engine set for seven days, US region only, three engine responses per prompt, and recorded the sources behind each answer. I exported the Sources tab on 15 August and classified nothing myself except the arithmetic below, which anybody can redo from the CSV.

Two things about that provenance matter more than they might seem to.

The first is that the prompt set includes prompts naming CueScout, which is why our own domain shows up 15 times in the corpus. Any per-brand number that includes branded prompts is inflated, ours worst of all. Where I quote visibility figures further down they are the unbranded ones.

The second is the engine question. The interface reported ChatGPT, Claude and Google; the vendor's own pricing page named ChatGPT, Google AI Overviews and Perplexity. Those disagree and I could not resolve it from the outside, so every number here should be read as a three-engine blend of unverified composition. For what it is worth, CueScout reads ChatGPT and Perplexity and nothing else, and I have never claimed a Google AI Overviews reading in a report.

The shape of the corpus

533 URLs across 360 domains, with 640.3 total usage points distributed across them (a page cited in 9.3% of answers scores 9.3).

The first surprise is how thin the long tail is. 288 of the 360 domains, exactly 80%, appear once and never again. The category's answers are not being assembled from a broad reading of the web; they are assembled from a few dozen pages plus a scattering of stragglers.

Slice of the corpusShare of all citation weight
Top 10 URLs10.2%
Top 20 URLs16.2%
Top 30 URLs21.0%
Top 50 URLs28.8%

Thirty pages out of 533 carry a fifth of everything. That number reframes the work. A year of blogging on your own domain competes with the other 483; getting your name onto twenty specific third-party pages competes for the fifth.

Who owns the sources

Source typeURLsShare of URLsShare of citation weight
Editorial23644.3%40.5%
Competitor-owned11822.1%31.5%
Corporate6311.8%8.3%
Institutional499.2%7.2%
Other275.1%3.6%
Our own domain152.8%2.3%
UGC (mostly Reddit)132.4%3.1%
Social112.1%3.3%
Review sites10.2%0.2%

Editorial is the biggest block, and it is the reachable one: agency blogs, publications, independent writers, none of whom sell a competing product and most of whom take pitches. Competitor-owned pages are a fifth of the URLs but nearly a third of the weight, which is the part of the table I found least comfortable. In this category the machine that supplies AI answers is largely competitors writing ranked lists about each other, and those lists get read harder than anything else per page.

Review sites, one URL and 0.2%, are the other end of it. If you have been buying G2 placement to influence AI answers in a category like this one, the corpus says you bought something else.

What formats get cited

The export labelled a content format for 299 of the 533 URLs and left the rest unlabelled, so this table is computed over those 299 and I have not tried to fill the gaps by guessing.

FormatURLsShare of labelled URLsShare of labelled weight
Ranked list11638.8%43.6%
Topic guide9632.1%24.4%
How-to guide4615.4%14.5%
Comparison155.0%7.3%
Reddit thread93.0%4.0%
LinkedIn article41.3%3.3%
Social, other51.7%0.8%
YouTube video31.0%0.9%
Everything else51.7%1.0%

Three formats, 86% of the labelled URLs and 82% of their weight. Searchable's own chart, drawn across all citations including the ones my export left unlabelled, put the same three at 73.6%, so the finding survives both denominators even though the exact figure moves.

The four LinkedIn articles are worth a second look. Four URLs, 3.3% of the labelled weight, and two of them sit in the six most-cited pages in the entire category. Both were written by individuals, on LinkedIn Pulse, in the last few months. Not a publication with a masthead. A person with an opinion and a headline containing a number.

The part that actually explains visibility

The corpus tells you what gets read. It does not tell you why one brand appears in answers and another does not. So on 15 August I fetched all 30 of the top-cited URLs with a browser user-agent, stripped them to text, and searched each page for every brand name in the category. That check is reproducible in an afternoon and I would rather you rerun it than take my word.

BrandOn how many of the top 30 pagesUnbranded visibility
Profound2130%
Otterly1921%
Peec AI1821%
Ahrefs189%
AthenaHQ106%
Conductor64%
ZipTie55%
Trakkr12%
CueScout00.7%

Two readings come out of this table.

The obvious one is the slope. Across nine brands, how many of thirty pages name you predicts how often an engine mentions you, and it does so about as cleanly as anything in marketing ever predicts anything. There is no mystery layer where a model forms a view of your product. It reads a handful of pages and repeats the names on them.

The less obvious one is Ahrefs. It sits on 18 pages, the same count as Peec AI, and gets under half the visibility. Ahrefs has a domain profile Peec cannot approach and a twenty-year brand, and in this dataset none of that converted. My guess, and it is a guess, is that Ahrefs appears in these lists as an SEO suite mentioned in passing while Peec appears as one of the products the page is about, and engines summarise what a page is about. I cannot prove that from the data I have. It is the next thing I want to measure.

Then there is our row. CueScout is on zero of the thirty most-cited pages in its own category and has 0.7% unbranded visibility, and I would rather publish that than a study in which we come out well. It also settles an internal argument: our problem was never the site, the product or the domain authority. We were not in the corpus, and no amount of writing on cuescout.com was going to change a number that is decided on somebody else's pages.

What I would do with this if it were your category

Run the same procedure. It took an afternoon once the export existed, and none of it needs our product.

  1. Write down 30 to 50 questions a buyer would actually type. Not keywords.
  2. Run them through a grounded engine and record every cited URL. Repeat a few days later, because retrieval moves around more than you expect.
  3. Count by host, then by URL. Look at where the top 20 cut falls.
  4. Fetch those top pages and search each for your own name and each competitor's.
  5. Sort the ones you are missing from by who owns them. Editorial and agency pages you can pitch. Competitor pages you cannot, and displacing those is a different project.

The output is a ranked list of specific URLs, which is a thing you can act on this week, rather than a content strategy, which is a thing you can talk about for a quarter.

What this dataset cannot tell you

  • One category, one 50-prompt set, seven days, US only, collected by a third party's crawler. Any of those could move the numbers.
  • The engine mix is unverified for the reasons above. Do not quote these as ChatGPT figures.
  • 234 of the 533 URLs carry no format label in the export, so the format table describes the labelled 299 and nothing more.
  • Usage percentages are the vendor's, computed by a method I cannot inspect. The share-of-total arithmetic in this post is mine and reruns from the CSV.
  • Correlation across nine brands is nine data points. It is a strong slope on a small sample, and the causal direction is at least arguable: being visible probably helps you get listed, too.
  • The visibility figures come from the same vendor's dashboard, unbranded filter applied. They are not independent of the corpus they explain.

If you rerun any of this and get a different answer, I would like to see it. The CSV is right there, and a second category would be worth more than anything else I could add to this one.


Related reading: AI citation strategy covers what to do once you have the list, and how to get added to a best-tools listicle is the pitch mechanics for the editorial half of the corpus. Our own numbers on what AI cites and how much of it ranks live on AI search citation statistics.

Frequently asked questions

Where did this citation data come from?

From a paid trial of Searchable, a competitor's AI-visibility product, which we bought for $125 and ran between 9 and 15 August 2026. The 533 URLs are its Sources export for our 50-prompt set. We did not collect the citations ourselves, and we say so on the page because a reader who finds that out later has every reason to stop trusting the rest.

Which engines produced these citations?

The tool ran three engines per prompt and its interface named ChatGPT, Claude and Google. Its own marketing page named a different trio, so we cannot state the engine mix with certainty. Treat the corpus as a three-engine blend and not as a claim about any single one. CueScout itself reads two engines, ChatGPT and Perplexity, and we make no claim about Google AI Overviews anywhere.

Does this generalise to other categories?

The method does. The numbers almost certainly do not. This is one category, one 50-prompt set, seven days, US only. What is worth copying is the procedure: export every cited URL, count by host, then fetch the top pages and search them for your own name.

What is the single most useful thing in the dataset?

The list of the 30 highest-usage URLs, with who owns each one. About 21% of all citation weight sits in those 30 pages, and roughly half are owned by editorial sites that accept pitches.

Find the questions worth writing about

CueScout scans Reddit, Hacker News, and Quora for the buyer questions AI answers are built from, explains why each one matched, and turns the ones that keep repeating into pages to publish on your own site. Nothing gets posted anywhere else.

Start your first scan