AI crawler checker
Check what a site tells GPTBot, ClaudeBot, PerplexityBot and five more, and whether the edge agrees with the file.
For example: · No account, no email, no daily limit.
Also free, no account:llms.txt checkerDomain Rating checkerCertificate expiry checkerAll tools
Which crawlers this checks
Eight agents, and the list is the one Domduck reads every day across its whole corpus rather than a list assembled for this page.
- GPTBot OpenAIBlocking it means: no ChatGPT browsing or training.
- OAI-SearchBot OpenAIBlocking it means: no ChatGPT search surface.
- ClaudeBot AnthropicBlocking it means: no Claude retrieval.
- PerplexityBot PerplexityBlocking it means: no Perplexity citation.
- Google-Extended GoogleBlocking it means: out of Gemini and AI Overviews grounding, Search ranking unaffected.
- Applebot-Extended AppleBlocking it means: out of Apple Intelligence training.
- CCBot Common CrawlBlocking it means: out of most open training corpora.
- Bytespider ByteDanceBlocking it means: out of TikTok and Doubao.
Why the answer has three states and not two
Every other free checker reports allowed or blocked. That is wrong whenever a robots.txt carries a catch-all rule and never names the agent, because the agent is then blocked by a rule that does not mention it. Domduck reports three states: allowed by name, blocked, and not named.
Not named is not the same as allowed. It means the file says nothing about that crawler, so whatever the catch-all rule says is what applies, and the catch-all can change tomorrow without the agent ever being mentioned.
We measured how often this matters on 3 August 2026. Across 87 well-known sites, a boolean answer was wrong for 10.3% of them, including reddit.com, facebook.com, x.com and netflix.com. Across 69 domains sampled from the Tranco list between ranks 1,700 and 200,000, it was wrong for 1.4%. So it matters most for exactly the sites people type first when they are testing a new tool.
Declared is not the same as enforced
robots.txt is a request, not a lock. Roughly 39.5% of GPTBot bans written in robots.txt are not enforced, and blocking can also happen at the edge with nothing written in the file at all.
So this checker sends one request to the domain as each crawler and compares what comes back against what our own user agent gets. Four readings come out of the two columns, and two of them are invisible from robots.txt alone: blocked in the file and served anyway, and blocked at the edge with nothing in the file.
Read that column carefully. Our requests come from our address, not from OpenAI’s or Anthropic’s. Those companies publish their crawler address ranges, so a site that allows the real GPTBot by address will still refuse us, and will show here as blocking. It is evidence, not proof, and no page here will tell you otherwise.
What happens on 15 September 2026
Cloudflare changes its AI crawler default that day. It applies to new customers, to new sites of existing customers, and to existing free-tier customers. A large number of sites will change what AI crawlers can reach without anybody editing a robots.txt.
That change happens at the firewall layer, so a checker that reads robots.txt will not see it move. Ours reads both, and it has been recording the robots.txt side across a quarter of a million domains since before the date, which is the part that cannot be bought back afterwards. Every policy flip we have observed is dated.
Should you block AI crawlers?
We do not have an opinion to sell you, but two facts are worth knowing before you decide. Blocking Google-Extended does not affect normal Google Search ranking, and many people block it believing it does. And blocking retrieval agents such as OAI-SearchBot or PerplexityBot removes you from answers that would have cited you, which is a different trade from blocking training crawlers.
How to block one
Add a group to robots.txt naming the agent, then a rule. Name the agents you care about rather than relying on a catch-all, because a catch-all tells a reader nothing about your intent and tools like this one cannot report it as a decision you made.
Our own robots.txt names all eight for that reason. Read it.
Every free tool here
- Domain Rating checker
Type a domain. See its Domain Rating, its popularity rank, how old it is, when its certificate expires and which AI crawlers it blocks.
- Bulk Domain Rating checker
Paste your list, one domain per line. No account, no captcha, no daily limit.
- llms.txt checker
Check whether a site publishes an llms.txt file, and whether the file is actually valid.
- Certificate expiry checker
Check when a certificate expires, who issued it, and how many names it covers.
- Domain badges
Five embeddable badges for any domain, each linking to that domain’s own record.