Citation corpus

One corpus, built from every check, instead of a snapshot

50

buyer questions per brand on Starter and Growth, each checked on a schedule

fixedPromptsPerBrand in billing/fixed_plans.go. Custom quotes for more than three brands run from 25 to 300.

One visibility check is a snapshot, and snapshots in this category are noisy. Ask ChatGPT the same buying question twice in a week and you can get two different source lists. So the useful object is not a single answer, it is the corpus: every page the engines have retrieved for your questions, accumulated over runs, with the date each one was last cited.

That corpus is the thing the rest of the product reads from. The domain leaderboard, the rank overlap check, the To-dos and the Opportunity Report are all second readings of the same rows. No extra crawl, no second bill.

Where it lives: Sources in the dashboard, and underneath Citations, Analytics, and the To-dos.

What it looks like

Illustrative
Six weekly runs48 URLs kept
Questionw1w2w3w4w5w6URLs
best product analytics for small teams14
alternatives to mixpanel11
cheapest event analytics tool9
posthog vs amplitude8
analytics that does not need an engineer6

A filled square is a run that returned sources for that question. The URL count only ever goes up, which is the difference between a corpus and a screenshot.

Six weeks of runs against one library. Illustrative data, drawn to show the accumulation rather than a single snapshot.

What it does

01

Accumulates instead of replacing

Each check adds its retrieved URLs to the corpus rather than overwriting it. A page cited in three different weeks for two different questions is one row with a citation count, not three unrelated results.

02

Keeps the question attached

Every URL carries the question it was cited for and the engine that cited it. The same page can be the top source for one buying question and absent from the next, which is the difference between a category authority and a lucky match.

03

Resets when you reposition

Change your product's description, keywords, or competitor set and the baseline is invalidated: the prompt library is regenerated and the mirror ignores checks from before the change. A repositioned product showing the names its old positioning collected is worse than showing nothing.

How it works

  1. 01Questions are seeded once from your product context, then edited and extended by hand. Generation returns about a dozen; the library you buy starts at 25.
  2. 02The library rotates by staleness, so the question checked longest ago goes next rather than the first one in the list.
  3. 03Ungrounded answers are stored but excluded from the citation rate. An engine answering from its weights has no way to cite anyone, and counting that as a miss would score you down for a question we never really asked.
  4. 04Each stored row records whether the engine searched the web, what it said about you, its sentiment, and the URLs it returned.

What it does not do

  • The corpus only covers questions in your library. A buying question nobody thought to add is invisible to it.
  • Engines differ. A page in your Perplexity corpus may never appear in ChatGPT's, so the selected engine set matters.
  • It is not a web crawl. We see what the engines chose to retrieve, which is a much smaller and more opinionated set than "pages about your category".

What you get

Every plan builds a corpus. The check budget decides how fast it fills.

Read the full plan comparison.

Live demo

See Source corpus on a real account, with no signup and nothing to configure.

Open the demo

Frequently asked questions

How many questions should I track?

Starter and Growth include 50 per brand, and very few products need more. Questions with real buying intent ("best X for Y", "alternatives to Z") produce a corpus you can act on; brand-name questions mostly confirm what you already know.

How often is it refreshed?

Every prompt is checked at least once a week, and a prompt whose answers are changing is checked every 3 days. A manual run that hits the rolling guard asks the stalest prompts first.

What happens to the old corpus when I change my positioning?

It stays stored but the baseline is reset, so the first-run mirror and the freshness reads ignore anything from before the change. New checks rebuild against the new positioning.

Related features

Every feature we publish

Run it on your own product

Start with the free visibility check to see whether the engines name you today, then run the scan that shows which pages they used instead.