AI crawlers allowed in robots.txt
Not possible hereSubstack controls robots.txt for the whole platform and you cannot edit it. Nothing here is yours to fix.
Site platforms
This page is shorter than the others because there is less to say. Substack controls robots.txt for the whole platform, serves no root-level files, and offers no custom head code. Four of the eight readiness checks cannot be passed on a Substack domain no matter what you do, and pretending otherwise would waste your time.
That is not the same as being invisible. Public Substack posts are crawled and quoted routinely — the platform is well-indexed and its posts are exactly the shape an answer engine likes: one question, answered at length, by a named person. What you cannot do is tune it. So the work moves to what you write and where else it appears.
How we recognise it: assets on substackcdn.com, or Substack post media URLs. Run the free readiness check and the fixes it gives you are the Substack ones below.
A crawler sees the paywall, not the post. Whatever is behind it cannot be cited, quoted, or used to decide that you are the authority on something. The archive you leave public is the entire surface an engine has.
Substack gives you no schema controls, so the title and the first paragraph are doing all the work of telling a retrieval system what this post answers. "What we learned about X" tells it nothing; the question itself tells it everything.
A one-page site on a domain you control, with Organization schema and sameAs links pointing at your Substack, gets you the checks Substack cannot. It is a weekend of work and it is the only route to them.
CueScout runs eight checks on whether an answer engine can read a site at all. These are the ones whose answer is different on Substack — the rest are the same everywhere.
Substack controls robots.txt for the whole platform and you cannot edit it. Nothing here is yours to fix.
Not possible on Substack — there is no root-level file access and no custom code. Skip this one.
Substack emits its own schema and gives you no custom head code, so this check cannot be passed on a Substack domain. If entity linking matters to you, it is an argument for a real site with the newsletter alongside it.
Substack publishes your subscription price on the /subscribe page, which is a real URL an assistant can cite. If this is failing, paid subscriptions are switched off.
These describe Substack’s own settings, which move without warning us. Last checked against the product on 2026-08-14.
Not on its own. Substack posts get cited on their merits and the platform is well-crawled. The argument for moving is control over your own domain and entity, which matters more the more the newsletter is a company rather than a person.
It changes what the URL says. Substack still serves the pages, still controls robots.txt, and still gives you no head code, so every constraint on this page still applies.
Because assistants asked whether something is worth paying for look for a price and skip what they cannot find. Substack's /subscribe page is a real URL with the number on it, which is why this one usually passes.
Ghost
robots.txt lives in the theme, llms.txt goes in /public, and Ghost already writes half the schema.
Site platformsWordPress
The virtual robots.txt, the Yoast and Rank Math settings that produce sameAs, and where llms.txt goes.
Site platformsWebflow
The robots.txt setting that only works on a custom domain, where JSON-LD goes, and why /llms.txt needs a proxy.
The readiness check is free, needs no signup, and takes about half a minute. It works out what your site is built on and gives you the fixes written for it.