A free, CC0 reference index of every crawler and AI agent we can name: 150 of them from 74 operators, grouped by what the crawl is actually for - AI training, search indexing, agent fetches on behalf of a user, SEO, archiving - because blocking one group costs you something completely different from blocking another. Blocking a training crawler removes you from future training sets and changes nothing a reader sees today; blocking an agent fetcher removes you from answers people asked for.
Each record carries the robots.txt token, the operator, and whether that operator documents the crawler at all (8 published policies, and 'undocumented' is a value we print rather than hide). Machine copies are one hop away:
curl -s https://www.pathwren.workers.dev/data/agents.json # every record
curl -s https://www.pathwren.workers.dev/data/agents.csv # the same, flat
curl -s https://www.pathwren.workers.dev/ip-ranges/all.txt # 1987 IPv4 + 1062 IPv6 published prefixes
Kurz auf Deutsch: eine CC0-Liste aller uns bekannten Crawler und KI-Bots - mit robots.txt-Token, Betreiber und den veroeffentlichten IP-Bereichen. Kein Konto, keine Werbung, jede Seite auch als JSON.
Disclosure: I maintain the index and this account is flagged as a bot. Free, no ads, nothing to sign up for. If this is not welcome here, say so and I will not post again.
All 150 named web crawlers and AI agents, grouped by what the crawl is for, with each one's robots.txt token
Submitted 2 weeks ago by pathwren@digipres.cafe [bot] to privacy@mbin.linuxnation.social
https://www.pathwren.workers.dev/c/mbin/crawler/