Skip to content

AI Visibility

Which Sources ChatGPT, Gemini, and Perplexity Cite When They Recommend Products (And How to Track Them)

We asked five AI engines the same 50 B2B buying questions on two days and kept every link they cited: 4,709 of them. ChatGPT leaned on vendors' own pages, Perplexity on 'best tools' lists, Google's AI Overviews on YouTube and Reddit. Here is how each engine picks its sources, what that means for getting recommended, and how to track it.

When ChatGPT recommends three tools, it is usually repeating something it just read. So the useful question is not only "does ChatGPT mention us?" but "which pages did it read before it answered, and are we on them?" Each engine answers that differently, and the difference is bigger than most GEO advice assumes.

To see how big, we asked five engines the same 50 buying questions on 1 and 2 October 2026 and kept every link they cited. That is 500 answers and 4,709 links across ChatGPT, Perplexity, Gemini, Google AI Overviews and Google AI Mode. The questions came from five B2B software categories (sales tools, CRM, workflow automation, ERP and marketing automation), written the way a buyer types them, with no brand names. The full method and the headline findings are on our State of AI Search research page. This post is the part about sources: which pages each engine picks, how it shows them, and how to keep track.

The run behind the numbers
500
answers
5
engines
4,709
cited links
~$3.15
total spend

50 buying questions in 5 B2B software categories, each asked once a day on 1 and 2 October 2026. Spend covers the engines plus the model that read each answer.

Why the cited sources decide the recommendation

Every engine in this post answers a buying question roughly the same way. It turns your question into one or more web searches, reads some of the pages that come back, and writes an answer from them. Google says AI Overviews and AI Mode may use a "query fan-out" technique, and OpenAI's help page says ChatGPT rewrites a prompt into "one or more targeted queries" before it searches. If your product is not on the pages the engine reads, the answer has nothing to name you from.

Here is one question from the run, asked on all five engines in the same run on 2 October:

The same question on five engines: 44 cited pages from 27 sites, and no site cited by all five

"What is the best marketing automation platform for B2B?" got 44 cited pages from 27 different sites. Three sites (gartner.com, ventureharbour.com and likehoney.nl) were cited by three engines each. None was cited by all five. ChatGPT's four sources did not show up in any other engine's answer at all. The recommendations were similar (all five put HubSpot first, then some order of Marketo, Salesforce and ActiveCampaign), but if you were a vendor trying to get into that answer, you would have had five different reading lists to get onto.

Independent research points the same way. The original GEO paper from Princeton and colleagues found that adding citations, quotes and statistics to a page lifted its visibility in generated answers by up to 40% on their benchmark, and by 22% for quotations on the live Perplexity site. Ahrefs found that branded web mentions correlate with AI Overview visibility at 0.664, against 0.218 for backlinks. Taken together, what the pages say, and what other sites say about you, seems to count for more than raw link authority.

There is a second reason to care. Pew Research tracked 900 US adults' browsing and found they clicked a result 8% of the time when an AI summary appeared, against 15% without one, and clicked a link inside the summary only 1% of the time. Fewer people reach your page to make up their own mind, so the cited page's version of you is often the only one they see.

Which engines you can track, per plan

CueScout checks every prompt you track once a day on each engine your plan includes:

  • Starter (20 prompts) checks ChatGPT, Perplexity and Google AI Overviews.
  • Growth (100 prompts) checks the same three plus Gemini and Google AI Mode.
  • Claude is an add-on on either plan, checked daily on every prompt.
  • Grok comes as a pack of 20 prompts you pick, checked weekly. We sell it that way because in our own Grok run it cited X only on "what are people saying" questions and never on buying questions.

Prices, add-on costs and limits are on the pricing page. If you only want to look at one question on one engine, the free ChatGPT visibility checker runs a live, search-grounded question and lists the sources the answer drew on. There are free Perplexity and Grok versions too.

What gets captured from each answer

For every prompt, engine and day, CueScout stores:

  • the full answer text, as the engine wrote it
  • every URL the answer cited, in order
  • every brand the answer named, yours and your competitors'
  • whether your domain was among the cited pages, and where you sat if the answer was a ranked list
  • the date and time of the check
  • for ChatGPT, the searches it ran before answering
  • for Google AI Overviews and AI Mode, the country and language the search ran in, whether Google showed an overview at all, and where each cited page ranked in the normal results on the same page

We do not store screenshots. A screenshot is nice evidence for one day, but you cannot search it, count it or compare it with last week. The answer text and the link list can be lined up day against day, which is how the overlap numbers further down were measured.

How each engine picks and shows its sources

The screenshots below are each engine's real answer to the marketing automation question, captured on 2 October 2026 and laid out with the cited pages on the right. The text is cut for length and nothing else is edited. The numbers under each one come from all 100 of that engine's answers in the run.

Distinct pages cited per answer, by engine
Perplexity19.7
Google AI Overviews6.5
ChatGPT4.7
Gemini3.5
Google AI Mode2.9

Mean distinct pages per answer, 100 answers per engine. State of AI Search run: 50 B2B buying questions asked on 1 and 2 October 2026, US English. A page is a URL with its query string and #:~:text= fragment removed.

ChatGPT: few sources, mostly the vendors' own pages

ChatGPT ran five searches, four asking for 'official' pages, then cited three vendor pages and one comparison article

ChatGPT shows its searches if you capture them, and they explain most of what it cites. For this question it ran five: one comparison search, then one per vendor with the word "official" in it. Three of its four citations were Adobe's, Salesforce's and ActiveCampaign's own documentation.

That was the pattern across the whole run, not a one-off. ChatGPT averaged 3.7 searches per answer. 31% of those searches contained "official" and 35% asked about pricing; none mentioned Reddit. In the end, 63% of the pages it cited were on the website of a brand the answer named (hubspot.com in 25% of answers, salesforce.com in 18%, microsoft.com in 17%). It cited 4.7 distinct pages per answer and gave no sources at all in 7 of 100 answers.

It also cited Reddit, YouTube and Wikipedia zero times in those 100 answers. I expected Reddit at least, because nearly every GEO deck I have read says ChatGPT loves it, so I went back through the raw captures to make sure the parser had not dropped them. It had not. Semrush tracked ChatGPT's Reddit citations falling from close to 60% of responses to about 10% over the summer of 2025, and our sample (software buying questions, October 2026) found none. Sources shift within weeks, which is the main argument for checking every day.

Two practical notes. First, 398 of the 477 links ChatGPT gave us carried utm_source=chatgpt.com, so that traffic is visible in your analytics if anyone clicks. Second, OpenAI says a site has to allow its OAI-SearchBot crawler: sites that block it "will not be shown in ChatGPT search answers". GPTBot is the separate training crawler. Blocking one does not block the other. In ChatGPT itself, the citations are inline links and a Sources button lists them, as described in OpenAI's help article.

If ChatGPT is the engine you care about, start with your own site. Your product, pricing and docs pages are the ones it goes looking for, so they should say plainly what the product does, who it is for and what it costs, in words a search for "[your brand] official pricing" would match.

Perplexity: twenty sources, and the most stable

Perplexity cited 20 pages, nearly all 'best tools' lists, with ZoomInfo's blog four times

Perplexity is the opposite. It cited a median of 20 pages per answer, and all 100 answers had sources. For our example question, nearly every one of the 20 was a "best marketing automation tools" list: GetResponse, Drip, Brevo and Klaviyo's own roundups, ZoomInfo's blog four times, G2 three times. Across the run, Zapier's blog was cited in 59% of Perplexity's answers, G2 in 56% and HubSpot's blog in 45%.

Why so many? Perplexity searches the web for every question and, as its help center puts it, every answer "includes numbered citations linking to the original sources". The list you get is closer to the search results it read than to the handful of pages the answer actually leans on. The numbered markers in the text tell you which of them carried a given claim.

Perplexity was also the steadiest engine from one day to the next, by a wide margin:

Same question, next day: how much of the cited-site list repeated
Perplexity73
Gemini43
Google AI Overviews42
Google AI Mode35
ChatGPT23

Mean Jaccard overlap of cited sites, day 1 vs day 2, %. 50 questions per engine, each asked on 1 and 2 October 2026. Overlap is shared sites divided by all sites cited on either day.

Ask Perplexity the same question two days running and its cited sites overlap by 73%. For ChatGPT the overlap is 23%. That matches what SparkToro found with a much larger panel in January 2026: there is a less than 1 in 100 chance that ChatGPT or Google's AI gives the same list of brands twice. A single reading of ChatGPT tells you very little, and it takes a week or two of daily readings before a pattern in its sources means much.

One caveat. We captured Perplexity through its Sonar API rather than the perplexity.ai app, and the app can return a different source list. Perplexity's crawler for search is PerplexityBot, which it says is not used to train models.

For Perplexity, most of the work is getting onto the category's "best tools" lists and comparison pages. Its sources are stable, so a placement on one of them tends to keep showing up. (Many of those lists are paid placements, so budget for that rather than expecting a free pitch to land.)

Gemini: the same page, linked again and again

Gemini gave 11 links but only 6 pages, linking drip.com four times, each link to a quoted sentence

The first time I counted Gemini's links I thought it cited more than ChatGPT. It returned 12.6 links per answer on average, second only to Perplexity. Strip the duplicates and it cited 3.5 distinct pages, fewer than ChatGPT's 4.7. In our example, drip.com's roundup was linked four times and Gartner's twice.

The reason is how Gemini links. 96% of its links ended in #:~:text=, a text fragment that jumps to the exact sentence it used. So each link is a claim pinned to a passage, and one good page can carry several of them. Google's Gemini Apps help page says sources appear "at the bottom of the response or in-line" when they are available, and that not every response includes them; 5 of our 100 answers had none. The Gemini API's grounding documentation describes the same loop: the model decides whether a Google search would help, runs one or more queries, and ties each part of the answer to a source.

Gemini also picked a winner far less often than the others. In the research run it named no single top choice in 97% of its answers, and gave a list sorted by use case instead. In six answers it linked Google Shopping product listings rather than web pages.

Because Gemini quotes passages, the sentence on a page that describes your product has to make sense on its own when lifted out of the paragraph around it.

Google AI Overviews: YouTube and Reddit

AI Overviews cited nine pages including a YouTube video, Gartner Peer Insights and G2

AI Overviews cited 6.5 pages per answer. Its most-cited site was YouTube, in 51% of answers, with Reddit in 48% and Quora in 15%. No other engine came close on those three. In our example the overview cited a YouTube video, Gartner Peer Insights and G2, next to six agency and vendor roundups.

Google is clear about eligibility. Its page on AI features and your website says a page only has to be indexed and eligible for a snippet, and that there are "no additional requirements" or special optimizations. Being eligible is a long way from being cited, though. Ahrefs' study of 4 million AI Overview citations in March 2026 found only 37.9% of cited URLs also ranked in the top 10 for the query. The rest ranked lower, or not at all. A good ranking helps, but a video or a thread about your category can get you into the overview without one.

Google also counts an overview as a single position in Search Console, with every link inside sharing it. We covered how to read that next to your citation data in our post on Search Console and AI citations.

Google AI Mode: the shortest list

AI Mode cited five pages, one of them a Reddit thread

AI Mode cited the fewest pages, 2.9 per answer, and gave none in 8 of 100. Like AI Overviews it leaned on forums and video: Reddit in 38% of answers, YouTube in 24%. It shared more with AI Overviews than any other pair of engines did (at least one common site on 71% of question-days), but Google says the two may use different models and techniques, and the overlap was still only 20% by site.

An earlier Semrush comparison of 5,000 keywords measured how far each engine's cited domains overlap with Google's own top 10: over 91% for Perplexity, around 86% for AI Overviews, about 54% for AI Mode, and lowest of all for ChatGPT. That fits what we saw. The engines that cite the most mainstream "best of" pages agree most with Google's rankings.

Claude and Grok

Neither was in the five-engine run, so I only have smaller numbers. Claude, with web search on, returned 4.9 citations per answer with no uncited answers in our 36-question test in September. Anthropic's help page says that when Claude searches, "every response includes citations". Grok uses xAI's web search tool and, as mentioned above, only reached for X posts on questions about what people were saying.

How the engines compare with each other

Put the engines side by side and they barely agree on sources. For the same question on the same day:

  • ChatGPT and Gemini cited at least one common site in only 18% of cases.
  • ChatGPT and Perplexity did in 77% of cases, but the overlap by site was 7%.
  • Perplexity and AI Overviews shared a site in 84% of cases, the most of any pair.
  • All five engines cited one common site on only 4 of 100 question-days.

So a single "AI visibility" number hides most of what is going on. You can be on Perplexity's reading list and absent from ChatGPT's, and the fix for each is a different page.

A last caution on accuracy. In March 2025 the Tow Center at Columbia tested eight AI search engines on news articles and found they got the source wrong in more than 60% of queries. Our run did not check whether each cited page actually said what the answer claimed. A cited page is the page the engine pointed to, which is still worth knowing, but open the page before you build a plan on it.

How often to check, and how long the data is kept

Every engine on your plan runs every prompt once a day. A new prompt is checked straight away instead of waiting for the next run, up to 100 new prompts per brand in any 30 days. The Grok pack runs weekly.

Daily matters most for ChatGPT, whose cited sites overlapped by only 23% from one day to the next. Perplexity moves slowly enough that weekly would mostly do, but you would not know that without the daily data.

Cited sources and answers stay available for 3 months on Starter and 12 months on Growth. That is long enough to compare the month before you published a page with the month after.

Engine-by-plan coverage table

The plan columns are what CueScout checks. The citation columns are real captures from the October run described above, 100 answers per engine.

EngineStarterGrowthPages cited per answerMost-cited sites in our runEngine's own docs
ChatGPTDailyDaily4.7hubspot.com 25%, salesforce.com 18%, microsoft.com 17%Search help, crawlers
PerplexityDailyDaily19.7zapier.com 59%, g2.com 56%, hubspot.com 45%How it works, crawlers
Google AI OverviewsDailyDaily6.5youtube.com 51%, reddit.com 48%, zapier.com 15%AI features
GeminiNoDaily3.5 (12.6 links)zapier.com 15%, technologyadvice.com 8%, monday.com 7%Sources help, grounding
Google AI ModeNoDaily2.9reddit.com 38%, youtube.com 24%, zapier.com 8%AI features
ClaudeAdd-on, dailyAdd-on, daily4.9 (36-question test)Not in this runWeb search help
GrokPack, weeklyPack, weeklyNot in this runx.com on sentiment questions onlyWeb search tool

The percentages are the share of that engine's 100 answers that cited the site at least once. The site lists are from software buying questions, so a question about running shoes or tax law would produce a different list. The pattern by engine is the part I would expect to carry over.

How the run was done

ChatGPT and Gemini answers were captured from their web apps, with ChatGPT's search switched on. AI Overviews and AI Mode came from Google's results pages, and Perplexity from its Sonar API. Everything ran in US English from a US location. We counted a page as a URL with its query string and text fragment removed, and a site as the registered domain, so blog.hubspot.com and hubspot.com count as one site. "On the website of a brand the answer named" is a name match between the cited domain and the brands in the answer, which we spot-checked by hand. The two days are a small window. Treat the percentages as a snapshot from October 2026, not a constant.

If you want to see where you stand before tracking anything, run your main buying question through the free ChatGPT visibility checker and look at the source list it returns. Those pages are the ones to get onto. For the wider playbook on what to change once you know, the GEO guide covers the rest.

Frequently asked questions

Which sources does ChatGPT cite most?

In our 100 ChatGPT answers to B2B software questions (October 2026), the most-cited sites were vendors' own: hubspot.com in 25% of answers, salesforce.com in 18%, microsoft.com in 17%. Overall, 63% of the pages ChatGPT cited were on the website of a brand the answer named. It did not cite Reddit, YouTube or Wikipedia once in that sample.

Why does Perplexity cite so many sources?

Perplexity runs a web search for every question and returns the result set with the answer. Through its Sonar API we got a median of 20 cited pages per answer, and every one of 100 answers had sources. The answer itself leans on only a few of them; the numbered markers in the text show which.

Do Google AI Overviews only cite pages that rank in the top 10?

No. Ahrefs' March 2026 study of 4 million AI Overview citations found 37.9% of cited URLs also ranked in the top 10 for the query. The rest ranked lower or not at all. In our run, YouTube videos and Reddit threads were the most-cited sources in AI Overviews.

Do ChatGPT and Gemini cite the same pages?

Rarely. On the same question and day, ChatGPT and Gemini shared at least one cited site in 18% of cases. Across all five engines, a single site was cited by every engine on only 4 of 100 question-days.

Can I see which pages ChatGPT cited for my brand without paying?

Yes, one question at a time. The free ChatGPT visibility checker runs one live, search-grounded question and lists the sources the answer drew on. Tracking the same questions every day, on every engine, is what the paid plans do.

Does CueScout save screenshots of the answers?

No. It stores the answer text, every cited URL, the brands named, whether your domain was cited, and when the check ran. Text and links can be searched and compared over time; a screenshot cannot.

See whether AI answers name you

CueScout checks your buying questions on ChatGPT, Perplexity, Gemini and Google’s AI answers every day, shows which pages they cite, and turns the gaps into To-dos.