Every AI crawler, and what it wants.

Search bots, fetchers that act for a user, and training crawlers, each from the vendor’s own docs, plus a tested robots.txt and the Cloudflare switch to check.

Two ants walking along the edge of a curved surface
Photo by Maksim Shutov on Unsplashdithered by Cyborb

AI crawlers come in three kinds. Search crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot index your pages so an assistant can cite them. User-triggered fetchers such as ChatGPT-User read a page when someone asks about it. Training crawlers and tokens such as GPTBot, ClaudeBot and Google-Extended decide whether your content trains future models.

The rule that follows: block a search crawler and you disappear from that assistant’s answers, while blocking a training crawler only opts you out of training. Below is every major token from the vendors’ own pages, a robots.txt we tested, the Cloudflare setting that since September 15, 2026 also blocks Googlebot when you block AI training, and how to spot fake bots in your logs.

The short version
  • To be cited, allow the search crawlers: OAI-SearchBot, Claude-SearchBot, PerplexityBot, plus Googlebot and Bingbot, which feed Google’s and Microsoft’s AI answers.
  • Blocking GPTBot, ClaudeBot, Applebot-Extended or meta-externalagent opts you out of training without removing you from search answers.
  • Google-Extended is the exception to watch: it also controls grounding in the Gemini app, so blocking it can cost you Gemini citations.
  • Since September 15, 2026, Cloudflare settings that block AI training also block Googlebot, Bingbot and Applebot. Check yours.
  • Anyone can claim to be GPTBot. Verify with the vendor’s published IP list or a reverse DNS lookup.

What are the three kinds of AI crawler?

AI crawlers differ by what they do with your page, and that decides what blocking them costs you.

KindWhat it doesExamplesBlocking it means
Search crawlerIndexes pages so an assistant can find and cite themOAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, bingbotFewer or no citations in that assistant
User-triggered fetcherReads a page because a person asked about itChatGPT-User, Claude-User, Perplexity-UserThe assistant cannot read the page live; several ignore robots.txt anyway
Training crawler or tokenCollects pages to train future modelsGPTBot, ClaudeBot, Google-Extended, Applebot-ExtendedYour content is left out of training; search answers are unaffected, with one Google exception

AI search crawlers: the ones that get you cited

These crawlers build the indexes that AI answers draw on. If you want to appear in an assistant’s answers, allow its search crawler.

Checked on September 28, 2026, on each vendor’s crawler page:

TokenCompanyWhat it doesBlocking itSource
OAI-SearchBotOpenAISurfaces sites in ChatGPT search; not used for trainingPages “will not be shown in ChatGPT search answers”, though they can still appear as navigational links. Changes take about 24 hoursOpenAI
Claude-SearchBotAnthropicIndexes content to improve Claude’s search results“May reduce your site’s visibility and accuracy in user search results”Anthropic
PerplexityBotPerplexitySurfaces and links sites in Perplexity’s results; not used for trainingPerplexity cannot index your pages. Changes take up to 24 hoursPerplexity
GooglebotGoogleCrawls for Google Search, including all Search features such as AI Overviews and AI ModeYou leave Google Search entirely. To leave only the AI features, use Search Console insteadGoogle
bingbotMicrosoftCrawls for Bing, which shares its index with CopilotYou leave Bing and Copilot. To limit Copilot only, use the NOARCHIVE or NOCACHE meta tagsBing
ApplebotApplePowers search in Spotlight, Siri and Safari; data may also train Apple’s modelsYou leave Apple’s search features. With no Applebot rules, it follows your Googlebot rulesApple
meta-webindexerMetaImproves Meta AI’s search resultsNot described by MetaMeta
MistralAI-IndexMistralIndexes pages for search in Mistral’s assistant, Vibe; not used for trainingNot described by MistralMistral
Amzn-SearchBotAmazonIndexes pages for search in Amazon products such as Alexa; not used for trainingNot described. With no rule of its own, it follows your rules for other search botsAmazon
DuckAssistBotDuckDuckGoCrawls pages in real time for DuckDuckGo’s AI-assisted answers; not used for trainingOut of those answers about 72 hours after you disallow it; organic rankings are unaffectedDuckDuckGo

Two assistants also lean on other search engines. Microsoft’s guidelines say Bing and Copilot “rely on the same core crawling, indexing, and ranking foundation”, so bingbot is the only crawler Copilot needs. Claude’s web search draws on Brave Search, which Anthropic lists as a subprocessor, as Simon Willison reported in March 2025. Brave’s crawler “does not advertise a differentiated user agent” and skips anything Googlebot may not crawl.

User-triggered fetchers: when someone asks about your page

These fetchers visit a page because a person asked an assistant about it, for example by pasting your link. Several of them ignore robots.txt by design, on the reasoning that a person, not a bot, made the request.

Checked on September 28, 2026:

TokenCompanyWhat it doesFollows robots.txt?Source
ChatGPT-UserOpenAIVisits pages for user actions in ChatGPT and custom GPTs“robots.txt rules may not apply”OpenAI
Claude-UserAnthropicFetches pages when a Claude user asksYes. Blocking it “may reduce your site’s visibility for user-directed web search”Anthropic
Perplexity-UserPerplexityVisits a page to answer a user’s question; not used for training“Generally ignores robots.txt rules”Perplexity
Google-Agent, Google-GeminiNotebook, Google-Read-Aloud and othersGoogleFetch pages for Google’s agents, Gemini Notebook sources and read-aloud, at a user’s request“Generally ignore robots.txt rules”Google
meta-externalfetcherMetaFetches links at a user’s request, including for agentic AI features“May bypass robots.txt rules”Meta
MistralAI-UserMistralVisits a page when a Vibe user asks, and links the sourceYesMistral
Amzn-UserAmazonFetches live information, for example for Alexa“May not follow all robots.txt directives”Amazon

Visits from these fetchers are useful signals: each one means an assistant read your page to answer someone. Our guide to tracking AI traffic shows how to count them per page.

Training crawlers and opt-out tokens

Blocking these keeps your content out of future model training. Except for Google-Extended, none of them affects whether assistants cite you today.

Checked on September 28, 2026:

TokenCompanyWhat blocking it doesEffect on search answersSource
GPTBotOpenAISignals that crawled content should not train OpenAI’s foundation modelsNone. OpenAI says each setting “is independent of the others”OpenAI
ClaudeBotAnthropicExcludes the site’s future content from Anthropic’s training dataNone; search runs on Claude-SearchBotAnthropic
Google-Extended (token only)GoogleStops use for Gemini training and for grounding in Gemini Apps and Vertex AI“Does not impact a site’s inclusion in Google Search”, but can remove you from Gemini app answersGoogle
Applebot-Extended (token only)AppleOpts out of training Apple’s foundation modelsPages “can still be included in search results”Apple
meta-externalagentMetaCrawls for “training foundation AI models or improving products”None statedMeta
MistralAI-TrainingMistralCrawls to build training datasetsNone; it is “not used for search indexing” or live answersMistral
AmazonbotAmazonCrawls to improve Amazon’s products; data “may be used to train Amazon AI models”None; Amazon search uses Amzn-SearchBotAmazon
CCBotCommon CrawlBuilds Common Crawl’s open archive of the webNone; it is not a search engine. Cloudflare’s managed robots.txt treats it as an AI crawlerCommon Crawl

Microsoft and Perplexity list no training crawler. Microsoft controls training with page tags instead: in 2023 it said content tagged NOARCHIVE is not used to train its generative AI models. Perplexity says PerplexityBot “is not used to crawl content for AI foundation models.”

Two names you may see in logs have no vendor page we could find: ByteDance’s Bytespider and xAI’s Grok crawlers. Cloudflare’s managed robots.txt blocks Bytespider as an AI crawler. For Grok, third-party lists name several user agents, and xAI has confirmed none of them.

Which AI crawlers should I allow?

Allow every search crawler and user fetcher if you want to be cited, and decide on training separately. For most sites that want visitors from AI answers, the simplest choice is to allow everything, which is what a robots.txt with no AI rules already does.

If you want to stay in AI answers but opt out of training, block only the training tokens. We ran this file through Python’s robots.txt parser (Python 3.14) to confirm who gets through:

robots.txt
# Search crawlers and user fetchers fall under *
User-agent: *
Disallow: /account/

# Training crawlers and tokens
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: MistralAI-Training
User-agent: Amazonbot
User-agent: CCBot
Disallow: /
What our parser test allowed for /blog/post
Allowed: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User,
         PerplexityBot, Googlebot, bingbot, Applebot, meta-webindexer,
         MistralAI-Index, Amzn-SearchBot, DuckAssistBot
Blocked: GPTBot, ClaudeBot, Applebot-Extended, meta-externalagent,
         MistralAI-Training, Amazonbot, CCBot

We left out Google-Extended on purpose. Add it only if you accept losing grounding in the Gemini app.

For what else helps a page get cited once crawlers can reach it, see our guide on getting cited by AI search. How AI answers changed search traffic is covered in SEO in the age of AI.

Is Cloudflare blocking AI crawlers on my site?

It might be, and since September 15, 2026 it might be blocking Googlebot too. Cloudflare now treats crawlers that serve both search and training, such as Googlebot, Bingbot and Applebot, as training crawlers. Any setting that blocks AI training, including the old “Block AI bots” toggle, now blocks them as well.

Cloudflare’s blog puts it plainly: “Multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training.” A site with that setting on can quietly drop out of Google and Bing. That also takes it out of AI Overviews and Copilot.

Checked on September 28, 2026, in Cloudflare’s docs:

Cloudflare featureWhat it does to crawlersWhere to check
AI bot policies: Search, Agent, TrainingEach can be set to block on all pages, block on pages with ads, or allow. Training now includes mixed-purpose crawlersSecurity Settings, then Configure AI bot policies
Block AI bots (legacy toggle)Deprecated on September 15, 2026; while on, it also blocks mixed-purpose crawlersSecurity Settings
AI Crawl ControlAllow or block each crawler, or charge it (a private beta); a block becomes a WAF rule. The Free plan identifies crawlers by user agent onlyAI Crawl Control, in your domain’s menu
Managed robots.txtAdds Disallow rules for Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagentSecurity Settings
Defaults for new domainsFrom September 15, 2026, Training and Agent bots are blocked on pages that show ads; Search stays allowedApplies at signup
  1. Open your AI bot policies

    In the Cloudflare dashboard, pick your domain, open Security Settings, then Configure AI bot policies.

  2. Set Search and Agent to Allow

    Cloudflare’s Search covers crawlers that index content to answer questions later. Agent covers bots acting for a person in real time, such as chat fetch bots. Blocking Search keeps AI search crawlers out, and blocking Agent stops assistants reading your pages when someone asks.

  3. Decide on Training, knowing the new cost

    Blocking Training now also blocks Googlebot, Bingbot and Applebot. To opt out of training without that, set Training to Allow and use the robots.txt above.

  4. Check each crawler in AI Crawl Control

    Make sure no search crawler shows Block or Charge. Remember that other WAF rules can still block a crawler you allowed.

The managed robots.txt is not a neutral default: it also blocks Google-Extended, which costs you grounding in the Gemini app.

How do I see AI crawlers in my logs?

Search your server’s access log for each token. This loop counts log lines per crawler and works in both bash and zsh:

count-ai-bots.sh
for bot in GPTBot OAI-SearchBot ChatGPT-User ClaudeBot Claude-SearchBot Claude-User \
  PerplexityBot Perplexity-User Googlebot bingbot Applebot meta-externalagent \
  meta-externalfetcher CCBot Amazonbot Bytespider DuckAssistBot MistralAI-User; do
  n=$(grep -ci "$bot" access.log)
  [ "$n" -gt 0 ] && echo "$n $bot"
done | sort -rn
Output on our 12-line sample log
2 Googlebot
2 GPTBot
1 bingbot
1 PerplexityBot
1 OAI-SearchBot
1 ClaudeBot
1 Claude-SearchBot
1 CCBot
1 Applebot

You will never see Google-Extended or Applebot-Extended in a log, because no crawler uses those names. Google’s existing crawlers do the fetching, and the token only controls how the content is used. Apple says Applebot-Extended “does not crawl webpages” at all.

A user agent is only a claim. In our sample we planted a fake GPTBot line from an address outside OpenAI’s range, and the checks below caught it. There are two ways to verify a bot.

Published IP lists. OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple, Common Crawl and Mistral publish their crawler addresses as JSON files. Every file in the table below loaded when we downloaded it on September 28, 2026, and this short script read all of them. It checks an address against one or more lists:

is_bot_ip.py
import ipaddress, json, sys

ip = ipaddress.ip_address(sys.argv[1])
for path in sys.argv[2:]:
    data = json.load(open(path))
    nets = [ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix")) for p in data["prefixes"]]
    hit = next((n for n in nets if ip.version == n.version and ip in n), None)
    print(f"{sys.argv[1]} in {path}: {'YES, ' + str(hit) if hit else 'no'}")
Terminal
curl -s https://openai.com/gptbot.json -o gptbot.json
grep GPTBot access.log | awk '{print $1}' | sort -u | while read ip; do python3 is_bot_ip.py "$ip" gptbot.json; done
Output
132.196.86.10 in gptbot.json: YES, 132.196.86.0/24
203.0.113.7 in gptbot.json: no

Reverse DNS. Google, Bing, Apple and Common Crawl crawlers resolve to their own domains. Look up the address, then look up the name to confirm it points back:

Terminal
host 66.249.66.1
host crawl-66-249-66-1.googlebot.com

In our test the four vendors’ addresses resolved to googlebot.com, search.msn.com, applebot.apple.com and crawl.commoncrawl.org, and each name pointed back to the same address. OpenAI and Anthropic crawler addresses had no reverse DNS, so use their IP lists instead.

Checked on September 28, 2026:

VendorPublished IP listReverse DNS works
OpenAIgptbot.json, searchbot.json, chatgpt-user.jsonNo, in our test
Anthropicbots.jsonNo, in our test
Perplexityperplexitybot.json, perplexity-user.jsonNot tested
Googlegooglebot.jsonYes, googlebot.com
Microsoftbingbot.jsonYes, search.msn.com
Appleapplebot.jsonYes, applebot.apple.com
Common Crawlccbot.jsonYes, crawl.commoncrawl.org
MistralUser and Index listsNot tested

FAQ

Does blocking GPTBot remove my site from ChatGPT?

No. GPTBot only controls training. ChatGPT search uses OAI-SearchBot, and OpenAI says each setting “is independent of the others”, so you can allow OAI-SearchBot and block GPTBot.

How do I stop Google’s AI Overviews from using my site?

Use Search Console’s “Search generative AI” setting and choose Exclude. It removes your site from AI Overviews, AI Mode and AI features in Discover, not from regular results. Blocking Google-Extended does not do this, because Google says it does not affect Search.

What is the difference between ClaudeBot, Claude-SearchBot and Claude-User?

ClaudeBot collects training data, Claude-SearchBot indexes pages for Claude’s search results, and Claude-User fetches a page when a user asks. Anthropic says all three follow robots.txt, so you can allow the last two and block the first.

Do AI crawlers run JavaScript?

Not in Vercel’s December 2024 measurement, which found that none of the major AI crawlers rendered JavaScript. Googlebot and Applebot do. The data is old, so the safe choice is the same: put your main content in the HTML the server sends, not only in scripts.

Why do I still see GPTBot after blocking it?

Check the address first, because anyone can send a GPTBot user agent. Then allow for delays: OpenAI says robots.txt changes take about 24 hours for its search crawler, and Perplexity says up to 24 hours.

Key takeaways
  • Search crawlers get you cited; training crawlers only decide whether you train models.
  • Allow OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and bingbot if you want AI citations.
  • Block training with robots.txt tokens, and treat Google-Extended as a Gemini citation trade-off.
  • On Cloudflare, a Training block now also blocks Googlebot, Bingbot and Applebot.
  • Verify bots by IP list or reverse DNS before you trust a user agent.

Next, learn how to track the traffic AI sends you, or see what the numbers say in AI search statistics.

Sources
  1. Overview of OpenAI crawlers, OpenAI
  2. Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic, April 2026
  3. Perplexity crawlers, Perplexity
  4. Google’s common crawlers, Google, July 2026
  5. Google’s user-triggered fetchers, Google, August 2026
  6. Search generative AI control, Google Search Console Help
  7. Which crawlers does Bing use?, Microsoft Bing
  8. Bing Webmaster Guidelines, Microsoft Bing
  9. New options to control usage of content in Bing Chat, Microsoft Bing, September 2023
  10. About Applebot, Apple, September 2026
  11. Meta web crawlers, Meta for Developers
  12. Mistral AI crawlers, Mistral AI
  13. DuckAssistBot, DuckDuckGo
  14. Amazonbot and other Amazon agents, Amazon
  15. CCBot, Common Crawl
  16. Brave Search crawler, Brave
  17. Anthropic use Brave for Claude’s web search, Simon Willison, March 2025
  18. Block AI bots, Cloudflare Docs
  19. Your site, your rules: new AI traffic options for all customers, Cloudflare, July 2026
  20. New AI traffic options, Cloudflare Changelog, July 2026
  21. Manage AI crawlers, Cloudflare Docs
  22. Managed robots.txt, Cloudflare Docs
  23. RFC 9309: Robots Exclusion Protocol, IETF, September 2022
  24. The rise of the AI crawler, Vercel, December 2024
cyborb.ai

Stop reading about it. Build it.

Describe what you want in plain words. Cyborb plans the work, writes and runs the code, makes the assets, and puts the result online.

Download Cyborb

Free to start. No card required.