On this page8 sections
- 01What are the three kinds of AI crawler?
- 02AI search crawlers: the ones that get you cited
- 03User-triggered fetchers: when someone asks about your page
- 04Training crawlers and opt-out tokens
- 05Which AI crawlers should I allow?
- 06Is Cloudflare blocking AI crawlers on my site?
- 07How do I see AI crawlers in my logs?
- 08FAQ
AI crawlers come in three kinds. Search crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot index your pages so an assistant can cite them. User-triggered fetchers such as ChatGPT-User read a page when someone asks about it. Training crawlers and tokens such as GPTBot, ClaudeBot and Google-Extended decide whether your content trains future models.
The rule that follows: block a search crawler and you disappear from that assistant’s answers, while blocking a training crawler only opts you out of training. Below is every major token from the vendors’ own pages, a robots.txt we tested, the Cloudflare setting that since September 15, 2026 also blocks Googlebot when you block AI training, and how to spot fake bots in your logs.
- To be cited, allow the search crawlers: OAI-SearchBot, Claude-SearchBot, PerplexityBot, plus Googlebot and Bingbot, which feed Google’s and Microsoft’s AI answers.
- Blocking GPTBot, ClaudeBot, Applebot-Extended or meta-externalagent opts you out of training without removing you from search answers.
- Google-Extended is the exception to watch: it also controls grounding in the Gemini app, so blocking it can cost you Gemini citations.
- Since September 15, 2026, Cloudflare settings that block AI training also block Googlebot, Bingbot and Applebot. Check yours.
- Anyone can claim to be GPTBot. Verify with the vendor’s published IP list or a reverse DNS lookup.
What are the three kinds of AI crawler?
AI crawlers differ by what they do with your page, and that decides what blocking them costs you.
| Kind | What it does | Examples | Blocking it means |
|---|---|---|---|
| Search crawler | Indexes pages so an assistant can find and cite them | OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, bingbot | Fewer or no citations in that assistant |
| User-triggered fetcher | Reads a page because a person asked about it | ChatGPT-User, Claude-User, Perplexity-User | The assistant cannot read the page live; several ignore robots.txt anyway |
| Training crawler or token | Collects pages to train future models | GPTBot, ClaudeBot, Google-Extended, Applebot-Extended | Your content is left out of training; search answers are unaffected, with one Google exception |
AI search crawlers: the ones that get you cited
These crawlers build the indexes that AI answers draw on. If you want to appear in an assistant’s answers, allow its search crawler.
Checked on September 28, 2026, on each vendor’s crawler page:
| Token | Company | What it does | Blocking it | Source |
|---|---|---|---|---|
| OAI-SearchBot | OpenAI | Surfaces sites in ChatGPT search; not used for training | Pages “will not be shown in ChatGPT search answers”, though they can still appear as navigational links. Changes take about 24 hours | OpenAI |
| Claude-SearchBot | Anthropic | Indexes content to improve Claude’s search results | “May reduce your site’s visibility and accuracy in user search results” | Anthropic |
| PerplexityBot | Perplexity | Surfaces and links sites in Perplexity’s results; not used for training | Perplexity cannot index your pages. Changes take up to 24 hours | Perplexity |
| Googlebot | Crawls for Google Search, including all Search features such as AI Overviews and AI Mode | You leave Google Search entirely. To leave only the AI features, use Search Console instead | ||
| bingbot | Microsoft | Crawls for Bing, which shares its index with Copilot | You leave Bing and Copilot. To limit Copilot only, use the NOARCHIVE or NOCACHE meta tags | Bing |
| Applebot | Apple | Powers search in Spotlight, Siri and Safari; data may also train Apple’s models | You leave Apple’s search features. With no Applebot rules, it follows your Googlebot rules | Apple |
| meta-webindexer | Meta | Improves Meta AI’s search results | Not described by Meta | Meta |
| MistralAI-Index | Mistral | Indexes pages for search in Mistral’s assistant, Vibe; not used for training | Not described by Mistral | Mistral |
| Amzn-SearchBot | Amazon | Indexes pages for search in Amazon products such as Alexa; not used for training | Not described. With no rule of its own, it follows your rules for other search bots | Amazon |
| DuckAssistBot | DuckDuckGo | Crawls pages in real time for DuckDuckGo’s AI-assisted answers; not used for training | Out of those answers about 72 hours after you disallow it; organic rankings are unaffected | DuckDuckGo |
Two assistants also lean on other search engines. Microsoft’s guidelines say Bing and Copilot “rely on the same core crawling, indexing, and ranking foundation”, so bingbot is the only crawler Copilot needs. Claude’s web search draws on Brave Search, which Anthropic lists as a subprocessor, as Simon Willison reported in March 2025. Brave’s crawler “does not advertise a differentiated user agent” and skips anything Googlebot may not crawl.
User-triggered fetchers: when someone asks about your page
These fetchers visit a page because a person asked an assistant about it, for example by pasting your link. Several of them ignore robots.txt by design, on the reasoning that a person, not a bot, made the request.
Checked on September 28, 2026:
| Token | Company | What it does | Follows robots.txt? | Source |
|---|---|---|---|---|
| ChatGPT-User | OpenAI | Visits pages for user actions in ChatGPT and custom GPTs | “robots.txt rules may not apply” | OpenAI |
| Claude-User | Anthropic | Fetches pages when a Claude user asks | Yes. Blocking it “may reduce your site’s visibility for user-directed web search” | Anthropic |
| Perplexity-User | Perplexity | Visits a page to answer a user’s question; not used for training | “Generally ignores robots.txt rules” | Perplexity |
| Google-Agent, Google-GeminiNotebook, Google-Read-Aloud and others | Fetch pages for Google’s agents, Gemini Notebook sources and read-aloud, at a user’s request | “Generally ignore robots.txt rules” | ||
| meta-externalfetcher | Meta | Fetches links at a user’s request, including for agentic AI features | “May bypass robots.txt rules” | Meta |
| MistralAI-User | Mistral | Visits a page when a Vibe user asks, and links the source | Yes | Mistral |
| Amzn-User | Amazon | Fetches live information, for example for Alexa | “May not follow all robots.txt directives” | Amazon |
Visits from these fetchers are useful signals: each one means an assistant read your page to answer someone. Our guide to tracking AI traffic shows how to count them per page.
Training crawlers and opt-out tokens
Blocking these keeps your content out of future model training. Except for Google-Extended, none of them affects whether assistants cite you today.
Checked on September 28, 2026:
| Token | Company | What blocking it does | Effect on search answers | Source |
|---|---|---|---|---|
| GPTBot | OpenAI | Signals that crawled content should not train OpenAI’s foundation models | None. OpenAI says each setting “is independent of the others” | OpenAI |
| ClaudeBot | Anthropic | Excludes the site’s future content from Anthropic’s training data | None; search runs on Claude-SearchBot | Anthropic |
| Google-Extended (token only) | Stops use for Gemini training and for grounding in Gemini Apps and Vertex AI | “Does not impact a site’s inclusion in Google Search”, but can remove you from Gemini app answers | ||
| Applebot-Extended (token only) | Apple | Opts out of training Apple’s foundation models | Pages “can still be included in search results” | Apple |
| meta-externalagent | Meta | Crawls for “training foundation AI models or improving products” | None stated | Meta |
| MistralAI-Training | Mistral | Crawls to build training datasets | None; it is “not used for search indexing” or live answers | Mistral |
| Amazonbot | Amazon | Crawls to improve Amazon’s products; data “may be used to train Amazon AI models” | None; Amazon search uses Amzn-SearchBot | Amazon |
| CCBot | Common Crawl | Builds Common Crawl’s open archive of the web | None; it is not a search engine. Cloudflare’s managed robots.txt treats it as an AI crawler | Common Crawl |
Microsoft and Perplexity list no training crawler. Microsoft controls training with page tags instead: in 2023 it said content tagged NOARCHIVE is not used to train its generative AI models. Perplexity says PerplexityBot “is not used to crawl content for AI foundation models.”
Two names you may see in logs have no vendor page we could find: ByteDance’s Bytespider and xAI’s Grok crawlers. Cloudflare’s managed robots.txt blocks Bytespider as an AI crawler. For Grok, third-party lists name several user agents, and xAI has confirmed none of them.
Which AI crawlers should I allow?
Allow every search crawler and user fetcher if you want to be cited, and decide on training separately. For most sites that want visitors from AI answers, the simplest choice is to allow everything, which is what a robots.txt with no AI rules already does.
If you want to stay in AI answers but opt out of training, block only the training tokens. We ran this file through Python’s robots.txt parser (Python 3.14) to confirm who gets through:
# Search crawlers and user fetchers fall under *
User-agent: *
Disallow: /account/
# Training crawlers and tokens
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: MistralAI-Training
User-agent: Amazonbot
User-agent: CCBot
Disallow: /Allowed: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User,
PerplexityBot, Googlebot, bingbot, Applebot, meta-webindexer,
MistralAI-Index, Amzn-SearchBot, DuckAssistBot
Blocked: GPTBot, ClaudeBot, Applebot-Extended, meta-externalagent,
MistralAI-Training, Amazonbot, CCBotWe left out Google-Extended on purpose. Add it only if you accept losing grounding in the Gemini app.
For what else helps a page get cited once crawlers can reach it, see our guide on getting cited by AI search. How AI answers changed search traffic is covered in SEO in the age of AI.
Is Cloudflare blocking AI crawlers on my site?
It might be, and since September 15, 2026 it might be blocking Googlebot too. Cloudflare now treats crawlers that serve both search and training, such as Googlebot, Bingbot and Applebot, as training crawlers. Any setting that blocks AI training, including the old “Block AI bots” toggle, now blocks them as well.
Cloudflare’s blog puts it plainly: “Multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training.” A site with that setting on can quietly drop out of Google and Bing. That also takes it out of AI Overviews and Copilot.
Checked on September 28, 2026, in Cloudflare’s docs:
| Cloudflare feature | What it does to crawlers | Where to check |
|---|---|---|
| AI bot policies: Search, Agent, Training | Each can be set to block on all pages, block on pages with ads, or allow. Training now includes mixed-purpose crawlers | Security Settings, then Configure AI bot policies |
| Block AI bots (legacy toggle) | Deprecated on September 15, 2026; while on, it also blocks mixed-purpose crawlers | Security Settings |
| AI Crawl Control | Allow or block each crawler, or charge it (a private beta); a block becomes a WAF rule. The Free plan identifies crawlers by user agent only | AI Crawl Control, in your domain’s menu |
| Managed robots.txt | Adds Disallow rules for Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent | Security Settings |
| Defaults for new domains | From September 15, 2026, Training and Agent bots are blocked on pages that show ads; Search stays allowed | Applies at signup |
Open your AI bot policies
In the Cloudflare dashboard, pick your domain, open Security Settings, then Configure AI bot policies.
Set Search and Agent to Allow
Cloudflare’s Search covers crawlers that index content to answer questions later. Agent covers bots acting for a person in real time, such as chat fetch bots. Blocking Search keeps AI search crawlers out, and blocking Agent stops assistants reading your pages when someone asks.
Decide on Training, knowing the new cost
Blocking Training now also blocks Googlebot, Bingbot and Applebot. To opt out of training without that, set Training to Allow and use the robots.txt above.
Check each crawler in AI Crawl Control
Make sure no search crawler shows Block or Charge. Remember that other WAF rules can still block a crawler you allowed.
The managed robots.txt is not a neutral default: it also blocks Google-Extended, which costs you grounding in the Gemini app.
How do I see AI crawlers in my logs?
Search your server’s access log for each token. This loop counts log lines per crawler and works in both bash and zsh:
for bot in GPTBot OAI-SearchBot ChatGPT-User ClaudeBot Claude-SearchBot Claude-User \
PerplexityBot Perplexity-User Googlebot bingbot Applebot meta-externalagent \
meta-externalfetcher CCBot Amazonbot Bytespider DuckAssistBot MistralAI-User; do
n=$(grep -ci "$bot" access.log)
[ "$n" -gt 0 ] && echo "$n $bot"
done | sort -rn2 Googlebot
2 GPTBot
1 bingbot
1 PerplexityBot
1 OAI-SearchBot
1 ClaudeBot
1 Claude-SearchBot
1 CCBot
1 ApplebotYou will never see Google-Extended or Applebot-Extended in a log, because no crawler uses those names. Google’s existing crawlers do the fetching, and the token only controls how the content is used. Apple says Applebot-Extended “does not crawl webpages” at all.
A user agent is only a claim. In our sample we planted a fake GPTBot line from an address outside OpenAI’s range, and the checks below caught it. There are two ways to verify a bot.
Published IP lists. OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple, Common Crawl and Mistral publish their crawler addresses as JSON files. Every file in the table below loaded when we downloaded it on September 28, 2026, and this short script read all of them. It checks an address against one or more lists:
import ipaddress, json, sys
ip = ipaddress.ip_address(sys.argv[1])
for path in sys.argv[2:]:
data = json.load(open(path))
nets = [ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix")) for p in data["prefixes"]]
hit = next((n for n in nets if ip.version == n.version and ip in n), None)
print(f"{sys.argv[1]} in {path}: {'YES, ' + str(hit) if hit else 'no'}")curl -s https://openai.com/gptbot.json -o gptbot.json
grep GPTBot access.log | awk '{print $1}' | sort -u | while read ip; do python3 is_bot_ip.py "$ip" gptbot.json; done132.196.86.10 in gptbot.json: YES, 132.196.86.0/24
203.0.113.7 in gptbot.json: noReverse DNS. Google, Bing, Apple and Common Crawl crawlers resolve to their own domains. Look up the address, then look up the name to confirm it points back:
host 66.249.66.1
host crawl-66-249-66-1.googlebot.comIn our test the four vendors’ addresses resolved to googlebot.com, search.msn.com, applebot.apple.com and crawl.commoncrawl.org, and each name pointed back to the same address. OpenAI and Anthropic crawler addresses had no reverse DNS, so use their IP lists instead.
Checked on September 28, 2026:
| Vendor | Published IP list | Reverse DNS works |
|---|---|---|
| OpenAI | gptbot.json, searchbot.json, chatgpt-user.json | No, in our test |
| Anthropic | bots.json | No, in our test |
| Perplexity | perplexitybot.json, perplexity-user.json | Not tested |
| googlebot.json | Yes, googlebot.com | |
| Microsoft | bingbot.json | Yes, search.msn.com |
| Apple | applebot.json | Yes, applebot.apple.com |
| Common Crawl | ccbot.json | Yes, crawl.commoncrawl.org |
| Mistral | User and Index lists | Not tested |
FAQ
Does blocking GPTBot remove my site from ChatGPT?
No. GPTBot only controls training. ChatGPT search uses OAI-SearchBot, and OpenAI says each setting “is independent of the others”, so you can allow OAI-SearchBot and block GPTBot.
How do I stop Google’s AI Overviews from using my site?
Use Search Console’s “Search generative AI” setting and choose Exclude. It removes your site from AI Overviews, AI Mode and AI features in Discover, not from regular results. Blocking Google-Extended does not do this, because Google says it does not affect Search.
What is the difference between ClaudeBot, Claude-SearchBot and Claude-User?
ClaudeBot collects training data, Claude-SearchBot indexes pages for Claude’s search results, and Claude-User fetches a page when a user asks. Anthropic says all three follow robots.txt, so you can allow the last two and block the first.
Do AI crawlers run JavaScript?
Not in Vercel’s December 2024 measurement, which found that none of the major AI crawlers rendered JavaScript. Googlebot and Applebot do. The data is old, so the safe choice is the same: put your main content in the HTML the server sends, not only in scripts.
Why do I still see GPTBot after blocking it?
Check the address first, because anyone can send a GPTBot user agent. Then allow for delays: OpenAI says robots.txt changes take about 24 hours for its search crawler, and Perplexity says up to 24 hours.
- Search crawlers get you cited; training crawlers only decide whether you train models.
- Allow OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and bingbot if you want AI citations.
- Block training with robots.txt tokens, and treat Google-Extended as a Gemini citation trade-off.
- On Cloudflare, a Training block now also blocks Googlebot, Bingbot and Applebot.
- Verify bots by IP list or reverse DNS before you trust a user agent.
Next, learn how to track the traffic AI sends you, or see what the numbers say in AI search statistics.
- Overview of OpenAI crawlers, OpenAI
- Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic, April 2026
- Perplexity crawlers, Perplexity
- Google’s common crawlers, Google, July 2026
- Google’s user-triggered fetchers, Google, August 2026
- Search generative AI control, Google Search Console Help
- Which crawlers does Bing use?, Microsoft Bing
- Bing Webmaster Guidelines, Microsoft Bing
- New options to control usage of content in Bing Chat, Microsoft Bing, September 2023
- About Applebot, Apple, September 2026
- Meta web crawlers, Meta for Developers
- Mistral AI crawlers, Mistral AI
- DuckAssistBot, DuckDuckGo
- Amazonbot and other Amazon agents, Amazon
- CCBot, Common Crawl
- Brave Search crawler, Brave
- Anthropic use Brave for Claude’s web search, Simon Willison, March 2025
- Block AI bots, Cloudflare Docs
- Your site, your rules: new AI traffic options for all customers, Cloudflare, July 2026
- New AI traffic options, Cloudflare Changelog, July 2026
- Manage AI crawlers, Cloudflare Docs
- Managed robots.txt, Cloudflare Docs
- RFC 9309: Robots Exclusion Protocol, IETF, September 2022
- The rise of the AI crawler, Vercel, December 2024




