Assistant Crawlers: Who to Allow and How to Verify
Learn which assistant crawlers to allow for AI citations, how to verify Googlebot and GPTBot by IP, and how to block unwanted bots with robots.txt.
ADS Beast editorial teamPublished Updated 11 min read
Assistant crawlers are automated bots that fetch your pages on behalf of AI systems such as ChatGPT, Claude, Perplexity and Google Gemini. You decide who gets in through robots.txt, then verify identity at the server level. Allowing the right ones keeps you citable in AI answers; blocking the wrong ones removes you from them.
Short version
- Assistant crawlers fall into two groups: retrieval bots that fetch pages to answer a live question, and training bots that collect text for model development. Blocking a retrieval bot removes you from that assistant's answers.
- GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot and Google-Extended are the tokens most sites need to make a decision about.
- Google-Extended is separate from Googlebot. Blocking it does not remove you from Google Search.
- User-agent strings prove nothing. Verify by IP range and reverse DNS, then forward-confirm.
- robots.txt is a request, not enforcement. Hard blocking happens at the server, WAF or CDN.
Which assistant crawlers matter for your site?
The four tokens that decide whether AI assistants can read and cite you are GPTBot, ClaudeBot, PerplexityBot and Google-Extended. Each belongs to a different vendor and controls a different thing: some handle live retrieval, some collect training data, and Google-Extended does both training and AI Overviews grounding for Google.
The distinction between retrieval and training is the one that costs people traffic. A retrieval bot fetches your page at the moment a user asks a question, so the assistant can quote or link you in the answer. A training bot collects your text into a dataset for later model versions. If you block a retrieval bot, you disappear from that assistant's answers today. If you block a training bot, you lose influence over future model behavior but keep your current citations.
Vendors do not always separate the two cleanly under one token, so check each vendor's documentation before you decide. The token name tells you the vendor, not the purpose.
| Crawler token | Vendor | What it typically does | Blocking it affects |
|---|---|---|---|
| GPTBot | OpenAI | Fetches pages for OpenAI systems | ChatGPT answers and training, depending on current vendor policy |
| ClaudeBot | Anthropic | Fetches pages for Anthropic systems | Claude answers and training, depending on current vendor policy |
| PerplexityBot | Perplexity | Live retrieval for answer generation | Your presence in Perplexity answers |
| Google-Extended | Controls Gemini training and AI Overviews grounding | Gemini and AI Overviews only, not Search | |
| Googlebot | Search indexing and ranking | Your Google Search rankings |
The table is a starting map, not a contract. Vendor behavior changes, and the only reliable source is the vendor's own crawler documentation, which you should re-read whenever you make a policy decision.
Googlebot and Google-Extended are not the same thing
Googlebot fetches pages so Google can index and rank them in Search. Google-Extended is a robots.txt token that controls whether your content trains Gemini models and grounds AI Overviews. Blocking Google-Extended leaves your Search rankings untouched.
This trips up more site owners than any other point on the list. People see "Google" in both names, assume one switch controls everything, and either block Googlebot by accident or leave Google-Extended open when they meant to close it.
Googlebot obeys robots.txt for crawling. It does not use robots.txt to decide indexing, which you control with meta tags and HTTP headers instead. So a Disallow rule for Googlebot limits what Google can fetch, while a noindex directive limits what Google can show. They do different jobs.
Google-Extended works the other way around. It is a permission token with no crawling role of its own. You either grant or withhold permission for Gemini training and AI Overviews grounding, and nothing about your ordinary Search presence changes either way.
If your goal is Search visibility plus AI answer visibility, you leave both open. If your goal is Search visibility without feeding Gemini, you keep Googlebot open and block Google-Extended. Those are two separate decisions and you should make them separately.
How to verify a crawler is really from Google
Verify Googlebot by matching the source IP against Google's published IP ranges, then running a reverse DNS lookup that resolves to a hostname ending in googlebot.com or google.com, then a forward lookup on that hostname that returns the same IP. Any step that fails means the request is spoofed.
The three-step check exists because each step alone is forgeable or incomplete:
- Get the source IP of the request from your server logs.
- Match it against the official IP ranges Google publishes for its crawlers.
- Run a reverse DNS lookup on the IP. The hostname must end in googlebot.com or google.com.
- Run a forward DNS lookup on that hostname. It must resolve back to the same IP you started with.
- If all three match, treat the request as genuine Googlebot. If any step fails, treat it as a fake.
The forward-confirm step is the one people skip. Reverse DNS alone can be faked by anyone who controls the PTR record for their own IP block, which is why the round trip back to the original IP is the part that actually proves anything.
User-agent strings prove nothing on their own. Anyone can send a request with "Googlebot" in the header. Log analysis tools that filter by user-agent alone will show you fake Googlebot traffic as if it were real, and you will make policy decisions on bad data.
The same method applies to other vendors. Most publish their crawler IP ranges or hostname suffixes in their documentation. If a vendor does not publish ranges, you have less to verify against, and you should treat unverifiable traffic with more caution rather than less.
How to block an unwanted AI crawler with robots.txt
Add a user-agent line with the bot's exact token, then a Disallow rule under it. Place the block before any wildcard rules, because the most specific matching user-agent group wins.
A working block looks like this:
User-agent: GPTBot
Disallow: /
Two things go wrong here often. First, people use the wrong token. "GPT" is not a token. "OpenAI" is not a token. The token must match exactly what the vendor publishes, or the rule is ignored. Second, people put the specific block after a wildcard group and expect it to override. Robots.txt resolves by specificity, not by order of appearance in the way most people assume, so structure your file with specific groups clearly separated from wildcards.
Robots.txt is voluntary. It is a convention that well-behaved crawlers follow and bad actors ignore. Treat every rule in it as a request.
If you also want AI assistants to understand your site structure and content priorities, publishing an llms.txt file for assistant crawlers gives cooperative systems a map they can actually use. It is not a substitute for robots.txt, and it is just as voluntary.
Why some crawlers ignore your robots.txt
Robots.txt is a convention, not a security control. Scrapers, malicious bots and some unlisted crawlers skip it entirely, and no text file can stop them.
For those, enforcement moves to the server layer:
- Rate limiting per IP and per user-agent, so a scraper cannot pull your whole site in minutes.
- WAF rules that return 403 to suspicious user-agent strings and known-bad IP ranges.
- IP allowlisting for known good bots, which inverts the model: everything is blocked unless it proves it belongs.
- CDN-level bot management, which filters before requests reach your origin.
Allowlisting is the strongest option and the most maintenance. You get hard enforcement, but you take on the job of keeping the list current as vendors add and retire IP ranges. For most sites, a combination of rate limiting and WAF rules catches the bulk of abuse without the upkeep.
The tradeoff is real. Aggressive blocking catches legitimate crawlers too, and a false positive on a retrieval bot quietly removes you from an assistant's answers without any error message anywhere. If AI citations are part of your traffic strategy, review your block rules after every change and check whether the assistants still mention you. An audit of what AI answers say about your company shows you the effect of a block within days, which is faster than waiting for a ranking report to notice.
Verification at scale: logs, sampling and false positives
At small volume you can verify crawlers by hand. Past a few thousand requests a day, you need a process, because manual checks stop happening and stale assumptions take over.
A workable process has three parts. First, log the source IP and user-agent for every request that claims to be a known crawler. Second, run the reverse-forward DNS check automatically on a sample, not on every request, and alert when the failure rate moves. Third, keep a short list of confirmed vendor IP ranges and refresh it on a schedule, since vendors add ranges without announcing it to you.
The failure mode to watch for is a false positive that nobody notices. You block an IP range because it looked suspicious, a retrieval bot was inside it, and your citations in that assistant drop to zero. Nothing in your analytics flags this, because AI answer traffic often arrives without a referrer you can attribute. This is the same attribution problem that shows up in ad measurement when the same period shows different numbers: the signal exists, but it is not where you expect to find it.
If AI-driven traffic feeds into your conversion tracking, make sure server-side events are set up so you can enable server-side conversion tracking without double-counting. Otherwise you will see the traffic in one system and the conversions in another, and you will not be able to tell whether a crawler policy change helped or hurt.
What to allow, in practice
Allow retrieval bots you want citations from. Allow training bots if you want influence over future model behavior and have no licensing objection. Block everything you cannot verify. That order of operations keeps you citable without opening the door to scrapers.
A concrete starting policy for most content sites:
- Allow GPTBot, ClaudeBot and PerplexityBot if AI citations are part of your strategy.
- Decide on Google-Extended separately from Googlebot, based on whether you want Gemini training and AI Overviews grounding.
- Block unverifiable traffic at the server, not in robots.txt.
- Re-check your rules quarterly against vendor documentation, because tokens and behaviors change.
The right answer depends on your business model. A publisher with licensing revenue has a different calculus than a SaaS company that wants to be mentioned in every relevant AI answer. Neither is wrong. What is wrong is leaving the decision unmade and letting defaults decide for you.
When AI-sourced traffic becomes a meaningful share of conversions, the crawler policy stops being a technical detail and becomes a channel decision. That is the point where measurement discipline matters, and where advertising operations for agencies that treat AI answers as a real acquisition channel start to separate from those that treat it as noise.
Next step
Open your server logs and pull every request from the past seven days that claims to be GPTBot, ClaudeBot, PerplexityBot or Googlebot. Run the reverse-forward DNS check on a sample of each. Whatever fails verification, block at the server. Whatever passes, confirm you actually want it, then leave it alone.
FAQ
How do I know if a crawler is really from Google?
Googlebot requests come from IP ranges published at developers.google.com/search/apis/ipranges. Match the source IP against that list. Reverse DNS must resolve to a hostname ending in googlebot.com or google.com, then a forward lookup on that hostname must return the same IP. User-agent strings alone prove nothing, since anyone can send a fake Googlebot header.
Which assistant crawlers should I allow if I want my content cited in AI answers?
The main ones are GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot and Google-Extended, which controls Gemini and AI Overviews training and grounding. If citations matter more than training, check each vendor's docs, because some bots handle live retrieval while others only collect training data. Blocking a retrieval bot usually removes you from that assistant's answers, not just its training set.
What is the difference between Googlebot and Google-Extended?
Googlebot fetches pages for Search indexing and ranking. Google-Extended is a robots.txt token that controls whether your content trains Gemini models and grounds AI Overviews. Blocking Google-Extended does not remove you from regular Google Search results, and blocking Googlebot does not affect Gemini training.
How do I block an unwanted AI crawler with robots.txt?
Add a user-agent line with the bot's exact token, then a Disallow: / rule under it, for example User-agent: GPTBot followed by Disallow: /. Place the block before any wildcard rules, since the most specific matching user-agent group wins. Robots.txt is voluntary, so treat it as a request, not enforcement.
Why do some crawlers ignore my robots.txt rules?
Robots.txt is a convention, not a security control. Bad actors, scrapers and some unlisted bots skip it. For those, use server-side blocks: rate limiting, WAF rules, or returning 403 to suspicious user agents and IP ranges. If you need hard enforcement, IP allowlisting for known good bots works better than any text file.
Does allowing assistant crawlers hurt my Google rankings?
No. Googlebot and the AI crawler tokens are separate. Allowing GPTBot, ClaudeBot or PerplexityBot has no effect on how Googlebot crawls or ranks your pages. The only token with a Google-side effect is Google-Extended, and it controls Gemini and AI Overviews, not Search.
See how this works in ADS Beast: assistant crawlers.