AI Visibility

Grok Cited X Posts 24 Times Across 36 Questions, and Every One Came From the Same Kind of Question

Everyone selling Grok optimisation says the same thing: Grok reads X, so post on X. We put 36 questions to Grok, ChatGPT and Perplexity, kept all 740 citations, and found Grok's X citations sit entirely inside one question shape. Buying questions got zero. Full dataset included.

Telman GadimovFounder, CueScout6 min read

Somebody replies "@grok is this true?" under a post and a bot answers, in public, to whoever is reading the thread. It has become the reflex of the site. A working paper by Thomas Renault, Mohsen Mosleh and David Rand counted 447,083 tweets tagging the bot to check another post between March and September 2025, and about 1.4 million fact-checking requests across Grok and Perplexity, which was 7.6% of all interactions with the two of them. Indicator wrote it up in more detail than I can here.

That trend produced a second, smaller one: a genre of blog post explaining how to get your brand cited by Grok. I have read a stack of them in the last month. They agree on the premise. Grok is different because it reads X as well as the web, therefore build an X presence and Grok will start recommending you.

The premise is checkable and nobody seemed to have checked it, so I did.

Download the dataset (CSV, 740 rows) — every citation from every answer, with the engine, the question, the question type and the position in the answer. Nothing removed.

What I ran

36 questions on 20 August 2026, put to three engines each, so 108 answers. The questions came in three groups of twelve, and the grouping is the whole point of the design.

Twelve buying questions, the kind that end with somebody choosing a vendor: best CRM for a small real estate team, best project management software for a 10 person agency, which applicant tracking system for a startup hiring its first 20 people. Twelve fact-check questions, shaped like the ones people throw at the bot: is it true that backlinks no longer matter, is it true that AI is replacing entry level engineering jobs, is it true that Perplexity ignores robots.txt. And twelve sentiment questions about what people are saying right now, on AI coding agents, the SEO industry, vibe coding, the design job market.

The engines were x-ai/grok-4.20:online, perplexity/sonar and openai/gpt-4.1-mini:online, all through OpenRouter, all with live search. Total cost $3.33, no failed calls.

One methodological thing that matters more than it sounds. I counted only the citations the API returned in its annotations, never URLs written into the answer text. A model writing a URL into its prose is often reciting something it half-remembers from training, and those fake references are indistinguishable from real ones by eye. We learned that the expensive way: of 236 source URLs we once stored, 213 turned out to be recall rather than references. Annotations are what the engine actually opened.

The result

EngineQuestion typeCitationsFrom x.comShare
GrokBuying9100%
GrokFact-check7900%
GrokSentiment852428.2%
PerplexityAll 3634610.3%
ChatGPTAll 3613900%

Grok cited X 24 times. All 24 sit in the sentiment group, and 10 of those 12 questions produced at least one. Cross the line into a buying question or a fact-check question and it stops completely, across 170 citations.

Ask Grok what designers are saying about the job market and four of its eight sources are individual X posts. Ask it what CRM a small real estate team should buy and it reads HousingWire, Follow Up Boss, Wise Agent and a Zapier roundup. Its most-cited domain on buying questions was zapier.com, then reddit.com, then ventureharbour.com. That list could have come from any engine. It is the ordinary web, read in the ordinary way.

So the premise is true and the conclusion drawn from it is not. Grok does read X, on questions about what is happening and how people feel. The questions where somebody is choosing what to buy are answered from the same listicles and comparison pages that decide ChatGPT and Perplexity.

The 24 posts

I looked at all of them, since 24 is a small enough number to just read.

They came from 24 different accounts. Not one account was cited twice. Twenty-three are people posting under their own name, including Ethan Mollick, Lily Ray and Tibo Maker. Exactly one is a company account, @cursor_ai, and it was cited on a question about AI coding agents where the company is the subject rather than a commentator.

If you were hoping the takeaway is that a brand account posting consistently gets picked up, this sample says the opposite as loudly as 24 data points can. Grok reached for named individuals with a track record of saying things about the topic. That is a much slower thing to build than a posting schedule, and it is not really a marketing channel.

The engines disagree more than I expected

This was meant to be a control and turned out to be the second finding.

Grok and Perplexity shared 74 of the domains they cited. Grok and ChatGPT shared 8. Six domains were cited by all three engines across all 36 questions: arxiv.org, digitalapplied.com, drip.com, en.wikipedia.org, perplexity.ai and searchenginejournal.com. 96 domains were cited by Grok alone and nobody else.

Anyone running a single-engine check and calling the result "AI visibility" is measuring one engine's reading list. We have found this before in a bigger corpus study, and I keep being surprised by how far apart they are.

Two other things fell out of the data. Perplexity cited Reddit 38 times out of 346 citations. Grok managed 9, and ChatGPT cited it zero times out of 139, so the received wisdom that AI search runs on Reddit is really a fact about Perplexity that got generalised. And ChatGPT returned no citations at all for 10 of its 36 answers, which is worth knowing if you are picking an engine to build a tracker on.

Where this is thin

The API is not the bot. x-ai/grok-4.20:online with live search is not the thing replying in someone's mentions, which runs inside X, sees the post above it, and is built for that job. My honest guess is the in-thread bot leans on X harder than what I measured. What I can defend is the claim about Grok answering a category question cold, which is the situation a business is in when a buyer asks it for a recommendation.

Beyond that: 36 questions is small, it was one day, the results were US-region, and the question set is mine, so it carries my assumptions about what a buying question looks like. The CSV is published so you can disagree with the classification. The three-way engine comparison is the part I would most want repeated at ten times the size before anyone treats the exact numbers as stable.

What I would actually do with this

If you sell something and you want Grok to name you when a buyer asks, work on the pages, not the posts. The pages Grok read on our buying questions were review roundups, comparison articles and category guides on other people's domains, which is the same work that gets you into the other two engines. There is no separate Grok strategy hiding in this data.

If your category genuinely turns on what people are saying this week, X matters to Grok and matters to nothing else we measured. Getting mentioned by hand in a post by someone people already listen to is the mechanism, going by the 24 accounts here.

And check more than one engine before you conclude anything, because on this evidence they hardly read the same internet.

We added a free Grok visibility checker while doing this, so you can put your own buying question to Grok and see the source list it built the answer from. It sits next to the ChatGPT and Perplexity versions. No signup, and nothing stored.

One boundary worth stating plainly, since this post is partly an argument for watching several engines: CueScout's paid tracking reads ChatGPT and Perplexity. It does not track Grok. I ran Grok through the API for this study, and it is on the free tool where you can run it yourself. If that changes I will say so here.

Frequently asked questions

Does Grok cite X posts?

Yes, but only on one kind of question. In our 36-question run Grok cited x.com 24 times and every one of them came from a question about current sentiment, like what designers are saying about the job market. On the 12 buying questions and the 12 fact-check questions it cited X zero times out of 170 citations, and read ordinary web pages instead.

Is this the @grok bot people summon on X?

No. We ran Grok through the API with live search on. The reply bot runs inside X, can see the post it is replying to, and is tuned for that job, so its source mix is its own and is likely to lean harder on X. What the two share is the model and the search behind it. Read this as evidence about how Grok reads a category, not as a transcript of the bot.

Which Grok model did you use?

x-ai/grok-4.20:online through OpenRouter, at $0.11 a question. The free checker we run uses x-ai/grok-4.3:online, which measured cheaper at $0.052 and returned more citations per dollar. Different slugs, same family and the same live search.

Is 36 questions enough to conclude anything?

It is enough for the size of the effect we found, and not enough for anything subtle. Zero X citations across 170 citations in 24 questions is not a marginal result you can wave away as noise. Anything smaller in this data, like Grok citing Reddit 9 times, I would not build a plan on.

So should I stop posting on X?

Not what the data says. It says X posting moves Grok on sentiment questions and did nothing for buyer-intent questions in our run. If you sell software and you want to be in the answer when somebody asks Grok which tool to buy, the pages that decide that are ordinary web pages, and they are the same ones that decide it in ChatGPT and Perplexity.

Find the questions worth writing about

CueScout scans Reddit, Hacker News, and Quora for the buyer questions AI answers are built from, explains why each one matched, and turns the ones that keep repeating into pages to publish on your own site. Nothing gets posted anywhere else.

Start your first scan