Citation corpus

The pages AI answers cite, and what to do about each one

6

verdicts, one per cited page, from "write this" to "you already hold it"

own-it, earn-it, get-listed, competitor-owned, blocked, held — geo/sources.go

Every tool in this category answers the same question: did the engine say your name. Citation Sources answers the question after it. When Perplexity recommends something for a buying question in your category, the answer is assembled out of pages it retrieved, and those pages are listed. We collect them, read them, and say what each one is, whether you are in it, and what it would take to be.

The list is usually uncomfortable reading. Most of it is other people's writing: a Reddit thread from two years ago, somebody's "best tools for X" roundup, a G2 category page. Those pages are doing the recommending, and not one of them shows up in an audit of your own site.

Where it lives: The Citations tab in the dashboard, per product.

What it does

Classifies every cited page

Eight types: community thread, listicle, comparison page, review site, documentation, a competitor's own domain, your own domain, or a plain article. The classification runs off the URL's host and path shape, which is deterministic, free, and explainable — three properties that matter more here than catching every edge case.

Says whether you are in it

We fetch the page and look. Present, absent, or unfavourable when the engine's own answer about you came out negative. Pages we could not read stay marked unknown and are never counted as absence, because "we have not looked" and "you are not there" are different facts and only one is worth acting on.

Turns each page into a verdict

A Reddit thread is not the same job as a G2 listing. Threads get "earn it" (write the page that answers the question properly), third-party lists get "get listed" (pitch the author), directories get "blocked" (reviews move those, writing does not), and reference material gets "own it". The list sorts by verdict, so the pages you can win sit at the top.

How it works

  • The corpus comes from citation rows the visibility checks already produced, so nothing here costs an extra fan-out of AI calls.
  • Pages are fetched on a budget per press, with failures backed off for 24 hours so a dead URL cannot crowd out pages that have never been read once.
  • A cited page on your own domain gets its own verdict, "held", and is reported as a win rather than a task.
  • Nothing on this surface suggests replying in a thread, and there is no Reddit account to connect. Four permanent bans on our own accounts taught us what the other CTA costs a customer.

What it does not do

  • We can tell you a page was cited. We cannot tell you which sentence in it earned the citation.
  • Classification reads URL shape, so an engineering blog hosted on a subdomain of a review site can land in the wrong bucket.
  • Pages behind a login or a hard bot wall stay unknown until they are readable, which on some review sites is never.
  • A citation is not a click. This measures whether the answer was built from you, not traffic.

Which plans have it

GEO score + cited-thread opportunities. On every plan. What changes by plan is how many questions feed the corpus.

GEO score + cited-thread opportunities by plan
BasicGrowthAgency
IncludedIncludedIncluded

Read the full plan comparison.

Try it free

The free version: which sites supply the AI answers in your category, ranked by share of citations, and whether your domain appears in that corpus at all.

AI Citation Source Radar

Live demo

See Citation Sources on a real account, with no signup and nothing to configure.

Open the demo

Frequently asked questions

Where do the cited URLs come from?

From the answers themselves. Search-grounded engines return the sources they retrieved alongside the text, and we store that list per question, per engine, per run. We read the structured annotation the API returns rather than scraping links out of the prose, which is what we used to do and got wrong.

Why is most of my corpus Reddit and forum threads?

Because that is what the engines retrieve for buying questions. In our own July 2026 audit of our category, 10 of the 22 URLs Perplexity cited were Reddit threads. The action is still a page on your own site, written well enough to be cited next to the thread.

Can I see this without signing up?

Partly. The public citation scan runs three questions on one engine and shows you the first five cited pages and the top five domains, no account needed.

Related features

Every feature we publish

Run it on your own product

Start with the free visibility check to see whether the engines name you today, then run the scan that shows which pages they used instead.