SEO

11% of My Search Console Impressions Are Competitor-Monitoring Bots

87 of 777 impressions on cuescout.com came from queries with a dozen -site: operators in them. Zero clicks, average position 6.9, while the rest of the site sits at 53. Here's how to spot them in your own data.

Telman GadimovFounder, CueScout7 min read

I was looking at Search Console last week, trying to work out whether any of my blog posts had started ranking, and one page looked like good news. /blog/social-listening-tool-for-founders, 69 impressions, average position 7.1. Page one. For a site that mostly sits on page five, that's the kind of number you screenshot.

It should have bothered me sooner that it was such an outlier. The pages either side of it in the report look like this:

PageImpressionsPosition
/compare/best-reddit-monitoring-tools19550.5
/compare/mention-alternative17974.1
/tools/ai-visibility-checker10979.7
/blog/social-listening-tool-for-founders697.1

One page at position 7 in a neighbourhood of 50 to 80. Either I'd written something unusually good, or something was wrong with the number.

So I filtered to that page and read the actual queries. They look like this:

"awario" -site:reddit.com -site:twitter.com -site:x.com -site:wykop.pl
-site:tripadvisor.com -site:youtube.com -site:yelp.com -site:booking.com
-site:facebook.com -site:instagram.com -site:tiktok.com

Nobody types that. That's a machine. And 56 of that page's 69 impressions, 81% of them, come from queries shaped like it.

What the numbers look like

I run CueScout, which scans Reddit for threads with buying intent, so I write a fair amount about social listening tools. That means my pages mention brand names like awario, brand24, talkwalker. Which, it turns out, is enough to get me swept up in whatever those companies (or their customers) are running against Google all day.

Over the three-month window in Search Console, the whole site got 777 impressions and 6 clicks. Filtering to queries that contain -site::

Whole siteQueries with -site:
Impressions77787
Clicks60
CTR0.8%0%
Average position536.9

87 of 777 is 11.2%. Zero clicks across all 87, which is what you'd expect from something that's parsing the results page rather than reading it.

The position column is the part that actually matters. These queries resolve at an average position of 6.9 while everything else I have averages 53. They are, by a distance, my best-ranking "traffic", and they represent nobody. Every one of them pulls the site-wide average position toward a number that looks like progress.

That's how a page ends up reporting position 7.1 without a single human having seen it.

There are at least two of them

Fifteen distinct queries had -site: in them. Reading through, they split cleanly into two groups by their exclusion list.

The first group always ends with the same eleven domains, in the same order, including wykop.pl (a Polish link-sharing site) and the review triple of tripadvisor.com, yelp.com, booking.com:

"octolens" -site:reddit.com -site:twitter.com -site:x.com -site:wykop.pl
-site:tripadvisor.com -site:youtube.com -site:yelp.com -site:booking.com
-site:facebook.com -site:instagram.com -site:tiktok.com

The second group uses a different set, notably fb.me, youtu.be, vm.tiktok.com and t.co, which is a list built by someone thinking about URL shorteners rather than review sites:

-site:facebook.com -site:fb.me -site:youtube.com -site:youtu.be
-site:youtube.be -site:twitter.com -site:instagram.com -site:tiktok.com
-site:vm.tiktok.com -site:t.co -site:x.com -site:reddit.com "sprinklr"

An exclusion list is a fingerprint. You build it once, you bolt it onto every query the product runs, and it never varies. Two lists means two products.

The one that convinced me these are tool-generated rather than a diligent human is this:

"brand24" -"24-7" -"24/365" -"24/7" -"brand 24 apparel" -"brand 24-hour"
-site:reddit.com -site:twitter.com -site:x.com -site:wykop.pl ...

Someone sat down and worked out that tracking the brand "Brand24" pulls in noise about 24/7 support hours and a clothing company called Brand 24 Apparel, and wrote negative keywords for each. That's a configuration screen, not a search box.

The full list of brands being tracked across both fingerprints: octolens, awario, awario.com, brand24, talkwalker, talkwalker alerts, sprinklr, consensus ai, plus the bare category terms "social listening tools" and ("reddit monitor" or "reddit monitoring") and ("tool" or "software" or "app" or "tools" or "platform").

So: social listening tools, monitoring social listening tools. My pages are collateral because I happen to write about the category.

What I'd actually do about it

Nothing, to the bots. They're reading public search results and I have no lever there, and honestly I'd be doing something similar if I needed that data.

The thing to fix is your own reporting. In the Performance report, Add filter, Query, "Queries containing", type -site:. That shows you the noise. Flip it to "Queries not containing" and you get something closer to your real numbers.

For me the difference is not subtle. Average position 53 is the honest read of where this site stands. Anything better than that is bots flattering me. And "new query appearing under position 20" is a signal I'd been using to decide what to write next, which means for a few weeks I was letting scrapers pick my content calendar.

If you write about any named product category, check yours. I'd genuinely like to know whether 11% is normal, high, or embarrassingly low.

Doing the filter properly

The one-minute version is above. If you want the real read on your site, four passes rather than one.

Pass one, size the noise. Performance report, Add filter, Query, "Queries containing", -site:. Note impressions, clicks, and average position. That is your contamination.

Pass two, your honest baseline. Same filter flipped to "Queries not containing". This is the number to report and the number to compare over time. Mine went from 53 to slightly worse than 53, which is not a fun revision but is at least real.

Pass three, per page. Sort your pages by position and look at anything that is a wild outlier against its neighbours. That is what caught this for me. A single page at position 7 in a report where everything else sits between 50 and 80 is not a success story, it is a data problem, and the instinct to screenshot it rather than interrogate it is the thing to fight.

Pass four, the query list itself. Read the actual queries for your top pages, not just the aggregates. Machine queries are obvious once you look at them, and there are shapes beyond -site: worth knowing: strings of quoted negative keywords, queries ending in a bare OR chain, identical queries repeating at a fixed daily cadence, and anything containing a full boolean expression with brackets.

Set a calendar reminder to redo this quarterly. The exclusion lists change when vendors update their products, and a filter that caught everything in July will miss a new fingerprint in October.

This is not the only flattering number

The reason I am fairly evangelical about this is that I got caught twice in the same month by two different vanity metrics pointing the same direction.

The other one was Ahrefs. It showed 372 referring domains for cuescout.com, against a domain rating of 0.1. Those two figures do not belong in the same sentence, and the explanation was that two days earlier the same tool had reported roughly 2 referring domains. What happened in between was a wave of scraper and auto-generated directory sites republishing our copy. Three hundred and seventy links that mean nothing, arriving in 48 hours, on a report designed to make link growth feel like progress.

My instinct on seeing it was to start disavowing, which would have been a week of work against a problem that does not exist. Google is good at ignoring this category of link; the damage was entirely to my judgement, not to the site.

The pattern in both cases is the same. Automated systems generate a lot of traffic against public data, that traffic lands in tools built when almost all traffic was human, and the tools present it in the same column as the real thing. The metrics that go up on their own are the ones to distrust.

For what it is worth, the metric I could not fake in either direction was AI citations, which sat at zero the entire time both of the above looked encouraging. Zero is a bad number and it was the only honest one I had. The longer argument about which of the two channels is still worth reporting on is in does SEO still generate inbound leads.

What this doesn't show

It's one small site. 777 impressions over three months is not a lot of data, and if you have real traffic your percentage will almost certainly be lower, because this is a roughly fixed volume of robot queries divided by a much bigger denominator. The absolute number of bot impressions is probably the more comparable figure than the ratio.

I can't prove which vendor is behind either fingerprint. The wykop.pl exclusion in the first group is suggestive, since it's an odd domain to care about unless you're Polish or serving Polish customers, and at least one company in that brand list is Polish. That's an inference from a domain list, not evidence, and I'm not going to name a culprit on that basis. The honest claim is narrower: two different automated systems, both tracking the same competitive set.

Search Console also warns that "chart totals and table results might be partial when filters are applied," so treat the 87 as approximately right rather than exact. And filtering on -site: only catches this one shape of machine query. Anything scraping Google without operators looks identical to a person in my data, so treat 11% as the low end of the range.

Frequently asked questions

What are -site: queries in Google Search Console?

They're search queries that use Google's -site: operator to exclude domains, usually a dozen or more at once, like -site:reddit.com -site:twitter.com -site:facebook.com. No person types these. They're generated by brand-monitoring and social-listening tools scraping Google results, and the impressions land in your Search Console if one of your pages ranks for the brand name being tracked.

Do bot impressions hurt my SEO?

They don't hurt your rankings. They corrupt your reporting. They arrive at high positions with zero clicks, which pulls your site-wide average position toward a better-looking number and drags your average CTR down. A page can show 69 impressions at position 7 and have been seen by no humans at all.

How do I filter bot queries out of Search Console?

In the Performance report, click Add filter, choose Query, set it to 'Queries containing', and enter -site:, that isolates them. To read your real numbers you want the inverse, so use 'Queries not containing' with the same string. Check both: the filtered view tells you how much noise you have.

Can I tell which tool is scraping me?

Not with certainty, but the exclusion list is a fingerprint. Vendors build a fixed list of domains to strip out of results and reuse it across every query, so queries from the same tool share an identical tail. In my data one fingerprint excludes wykop.pl, tripadvisor.com, yelp.com and booking.com; another excludes fb.me, youtu.be, vm.tiktok.com and t.co. Those are two different products.

Related cluster

Keep reading

Generative Engine Optimization guide

Find the questions worth writing about

CueScout scans Reddit, Hacker News, and Quora for the buyer questions AI answers are built from, explains why each one matched, and turns the ones that keep repeating into pages to publish on your own site. Nothing gets posted anywhere else.

Start your first scan