OAI-SearchBot
The crawler behind ChatGPT’s search surface: it fetches pages so they can be cited in answers.
6.7% of the 2,228 domains Domduck read on disallow it, against 5.1% on .
What blocking OAI-SearchBot costs
Your pages stop appearing as citations in ChatGPT search results. For most publishers this is the expensive one to block, because a citation is a link with a reader attached.
What it does not cost
It is not a training crawler. Blocking it does not keep your content out of a model, and allowing it does not put your content into one.
This is the agent to think hardest about. The traffic case for allowing it is the strongest of the eight, and the training objection does not apply to it.
How to block it
or allow it explicitlyUser-agent: OAI-SearchBot
Disallow: /To allow it and say so, use a bare Disallow: with nothing after it, which RFC 9309 defines as “nothing is disallowed”. That matters more than it looks: a file that never names an agent leaves it in a third state, neither allowed nor blocked, and a catch-all rule added later will block it without anyone deciding to.
Block rate over time
2 daysThe percentage of domains in the corpus whose robots.txt disallows OAI-SearchBot, measured once per domain per pass. The denominator is how many domains were read that day and is stored with each point, so the line is not restated when the corpus grows.
The other seven
- OpenAI
- GPTBot
- Anthropic
- ClaudeBot
- Perplexity
- PerplexityBot
- Google-Extended
- Apple
- Applebot-Extended
- Common Crawl
- CCBot
- ByteDance
- Bytespider
How this page is made
The explanation above is written by a person. The numbers are computed from Domduck’s own daily reads of each domain’s robots.txt, parsed to RFC 9309 rather than searched for a string: a rule binds to the nearest preceding group of user-agent lines, the longest matching pattern wins, and Allow breaks a tie. No prefix matching is done, so a group naming Google is not treated as a policy about Google-Extended. What the collectors have read.