Crawler registration is something a crawler operator does with its own infrastructure or a third-party platform such as Cloudflare — it is not a service a website owner or OpenForBots performs. Verification is a separate, narrower question: confirming that a specific request actually came from the provider it claims to be, using the provider’s published IP ranges with reverse-DNS confirmation, or a cryptographic mechanism such as Web Bot Auth. A robots.txt token or User-Agent string alone proves neither identity.
Three different things people mean by “crawler registration”
The phrase gets used for at least three distinct jobs. Keeping them separate avoids acting on the wrong evidence.
| Term | Who does it | What it actually establishes |
|---|---|---|
| Crawler registry | A reference publisher (for example, the OpenForBots crawler registry) | A documented, sourced description of a crawler’s identity, purpose, and policy token — not a live verification of any single request |
| Crawler-operator registration | The crawler operator (OpenAI, Google, Perplexity, Anthropic, Cloudflare’s Verified Bots programme, and similar) | The operator publishing IP ranges, reverse-DNS hostnames, or a signing key so that others can verify its traffic |
| Website-owner policy declaration | The website owner | A robots.txt rule naming a token and stating an access preference — a policy statement, not an identity check |
OpenForBots operates the first row only. It documents crawler identities and purposes from provider sources; it does not register, authenticate, or independently certify third-party crawlers, and it has no mechanism to add a crawler to a provider’s own verified-traffic system.
How providers actually document verification
Three mechanisms currently appear in official documentation, in roughly increasing order of strength.
1. Published IP ranges with reverse-DNS confirmation
Google documents a two-step check for confirming a request came from a Google crawler: run a reverse-DNS lookup on the source IP, confirm the resulting hostname ends in google.com, googlebot.com, or googleusercontent.com, and then run a forward DNS lookup on that hostname to confirm it resolves back to the original IP. OpenAI and Perplexity document a comparable pattern: they publish machine-readable IP-range files for their crawlers (for example OpenAI’s crawler IP-range files and Perplexity’s published PerplexityBot/Perplexity-User ranges) and recommend combining IP matching with the documented User-Agent token, since a single IP can be reassigned over time.
This is stronger than a User-Agent check alone, but it still requires the website owner to actively run the lookup against server logs at request time — it is not something a robots.txt file or a registry page can do for them.
2. Verified-bot programmes run by a third-party platform
Cloudflare operates a Verified Bots programme: crawler and agent operators apply to Cloudflare directly, and Cloudflare confirms the operator’s identity before listing the bot as verified across sites that use Cloudflare’s network. This is useful, but it verifies traffic through Cloudflare’s infrastructure, not universally — a site not sitting behind Cloudflare does not get this verification for free, and Cloudflare’s programme is Cloudflare’s own review, not an internet-wide standard.
3. Cryptographic signing (Web Bot Auth)
Web Bot Auth is an emerging mechanism, described in an active IETF Internet-Draft and implemented in production by Cloudflare and other providers, in which an automated client signs each outbound request using HTTP Message Signatures and publishes its public key at a well-known key-directory URL it controls. A server can then verify the signature cryptographically instead of trusting an IP range or a User-Agent string. As of this guide’s last-verified date, the underlying protocol is a working-group Internet-Draft rather than a finished standard, and adoption is still provider-by-provider.
What robots.txt does and does not prove
robots.txt, as defined in RFC 9309, states a site’s crawl policy for a declared token. It does not:
- authenticate that a given request actually came from the named provider;
- prove that a compliant crawler respected the rule (compliance is voluntary and provider-dependent);
- grant or revoke access to confidential content, which still requires real authentication.
A request claiming to be GPTBot, Googlebot, or PerplexityBot in its User-Agent header can be sent by anyone. Treat a robots.txt match as a policy signal, and treat IP/reverse-DNS or cryptographic signature checks as the identity signal — they answer different questions.
What OpenForBots does and does not do here
- OpenForBots’ crawler registry documents provider-sourced identity, purpose, and policy notes for known crawler tokens, with a verification date on every record.
- OpenForBots’ robots.txt AI crawler checker parses a site’s deployed
robots.txtand reports the matching policy rule for documented tokens. - OpenForBots does not verify that a specific live request came from a named provider, does not register or certify third-party crawlers, and does not operate a verified-bot programme of its own.
A compact checklist
- Identify the exact crawler token and provider you are evaluating.
- Read that provider’s current documentation rather than a cached summary — verification mechanisms change.
- For a policy question, resolve the deployed
robots.txtrule for the token and path. - For an identity question, check the provider’s published IP ranges and confirm with a forward-confirmed reverse-DNS lookup where the provider documents one.
- Where a provider supports cryptographic signing (Web Bot Auth or a platform-specific equivalent), prefer it over IP/User-Agent matching alone.
- Do not treat a matching User-Agent string, by itself, as proof of identity.
- Record the date you checked, since IP ranges and provider mechanisms change.
What this guide cannot settle
Provider verification mechanisms are documented by each provider and can change without notice. This guide reflects current official documentation as of its last-verified date; it does not guarantee that a given provider’s mechanism remains unchanged, and it does not verify any specific request on your own infrastructure. Running your own reverse-DNS or signature check against your own logs is the only way to confirm a specific request.
Continue
Use the robots.txt AI crawler checker to see the current policy rule your site declares for documented tokens, browse the crawler registry for provider-sourced identity records, or read how to check AI crawler access for the full evidence workflow from policy to public delivery.