Research

Original measurement from our own readings, published with the method, the sample size and the raw table.

The AI Crawler Index

What the web tells eight AI crawlers, measured from robots.txt across the corpus and published with the date every reading was taken.

The rules every edition follows

  • The corpus size and the exact date range appear on the page, beside the numbers.
  • The aggregate table ships as a downloadable CSV, free to use with attribution.
  • What the data cannot show is stated on the page. robots.txt is a declaration and not a control, and a study that treats it as behaviour is wrong.
  • No intent is inferred. A record says an agent was blocked on a date. It never says a site decided to block AI.
  • Date precision matches the reading cadence, never the other way round.
  • No domain is named in an aggregate, and no account is identified anywhere.

Why this is running now

Cloudflare changes its AI-crawler default on 2026-09-15, for new customers, new sites of existing customers and existing free-tier sites. A large number of sites will change posture on one day without anyone editing a file.

robots.txt has no history. Change it and the previous version is gone, so what a site allowed the day before can only be known if somebody recorded it first. That measurement cannot be bought, backfilled or reconstructed afterwards, which is why the readings are running months before the study that needs them.

Working with journalists

If you need a cut these pages do not show, by top-level domain, by rank band, by registrar or by date, write to [email protected] and it gets run. Figures are free to quote with attribution and a link.