AI Ranked
Reference

AI crawlers list: every user agent for robots.txt

The user agent tokens OpenAI, Anthropic, Perplexity, Google, Apple, Microsoft, Meta and others publish, what eight major sites actually do with them, and how to block training without disappearing from AI search.

Short answerAI crawlers are the bots AI companies use to fetch web pages, and each has a user agent name you can allow or block in robots.txt. There are three kinds: training crawlers (GPTBot, ClaudeBot), search crawlers that decide whether you can be cited (OAI-SearchBot, Claude-SearchBot, PerplexityBot), and user-triggered fetchers (ChatGPT-User, Perplexity-User), which may ignore robots.txt.
Key takeaways
  • Of the 59 major sites that block GPTBot in our 261-site study, 31 still allow OAI-SearchBot: they refuse OpenAI training and keep ChatGPT search citations.
  • Blocking OAI-SearchBot keeps a site out of ChatGPT search answers; blocking GPTBot only opts it out of training1.
  • Of eight major sites whose robots.txt we fetched for this guide, the New York Times, CNN and the BBC block OAI-SearchBot too, while Medium blocks training bots and leaves AI search crawlers alone.
  • Google-Extended "does not impact a site's inclusion in Google Search," so it does not remove you from AI Overviews or AI Mode4.
  • OpenAI, Anthropic and Perplexity publish JSON lists of their crawler IPs, so a request claiming to be GPTBot can be checked in seconds123.

What are the three kinds of AI crawler?

Training crawlers, search crawlers and user-triggered fetchers. They do different jobs, so blocking each one costs you something different, and a robots.txt that lumps them together as "AI bots" usually throws away search visibility to stop training.

Training crawlers such as GPTBot, ClaudeBot, meta-externalagent and MistralAI-Training collect public pages that may train future models. Blocking them changes what future models learn and, per the operators' own docs, has no effect on whether you're cited in search answers today.

Search crawlers build the index an AI engine retrieves from: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and bingbot. Block one of these and you drop out of that engine's answers.

User-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User) load a specific page because someone asked about it or pasted the link. Several operators say robots.txt may not apply to them, since a person started the request.

Two entries aren't crawlers at all. Google-Extended and Applebot-Extended never fetch pages; they tell Google and Apple how pages already crawled by Googlebot or Applebot may be used45.

Which AI crawler user agents should I know?

The table lists every AI-related user agent token we could verify in the operator's own documentation. "Respects robots.txt" is what each operator states, which isn't independent proof. Use the token exactly as written in the first column of your robots.txt.

User agent tokenOperatorPurposeRespects robots.txt (per docs)Official docs
GPTBotOpenAITrainingYes. Disallow signals content should not be used for trainingOpenAI crawlers
OAI-SearchBotOpenAISearch index (ChatGPT search)Yes. Opted-out sites are not shown in ChatGPT search answersOpenAI crawlers
ChatGPT-UserOpenAIUser-triggered fetchNot always. OpenAI says robots.txt rules may not applyOpenAI crawlers
ClaudeBotAnthropicTrainingYes, including Crawl-delayAnthropic help center
Claude-SearchBotAnthropicSearch index (Claude search)YesAnthropic help center
Claude-UserAnthropicUser-triggered fetchYesAnthropic help center
PerplexityBotPerplexitySearch index (not training)Docs recommend allowing it in robots.txt; compliance disputed (see below)Perplexity crawlers
Perplexity-UserPerplexityUser-triggered fetchGenerally ignores robots.txtPerplexity crawlers
GooglebotGoogleSearch index, including AI Overviews and AI ModeYesGoogle crawlers
Google-ExtendedGoogleControl token for Gemini training and grounding (does not crawl)Yes. Token only, crawling is done by other Google agentsGoogle crawlers
Google-CloudVertexBotGoogleCrawls requested by site owners building Vertex AI AgentsYes. No effect on Google SearchGoogle crawlers
ApplebotAppleSearch (Spotlight, Siri, Safari); may also train Apple modelsYes. Falls back to Googlebot rules if Applebot is not namedAbout Applebot
Applebot-ExtendedAppleControl token for training on Applebot data (does not crawl)Yes. Token onlyAbout Applebot
bingbotMicrosoftBing index, which grounds CopilotYesBing crawlers
meta-externalagentMetaTraining and product indexingYes, block with a disallow ruleMeta web crawlers
meta-externalfetcherMetaUser-triggered fetchMay bypass robots.txtMeta web crawlers
CCBotCommon CrawlOpen repository of web crawl dataYesCCBot
AmazonbotAmazonProduct improvement; may train Amazon AI modelsYes. No Crawl-delay supportAmazonbot
Amzn-SearchBotAmazonSearch (for example Alexa); not trainingYesAmazonbot
Amzn-UserAmazonUser-triggered fetch (live Alexa answers)May not follow all directivesAmazonbot
DuckAssistBotDuckDuckGoReal-time AI-assisted answers; not trainingYes, applied within 72 hoursDuckAssistBot
MistralAI-UserMistral AIUser-triggered fetch; not trainingNot detailed in docsMistral robots
MistralAI-IndexMistral AISearch index; not trainingNot detailed in docsMistral robots
MistralAI-TrainingMistral AITrainingYesMistral robots

Under RFC 9309, the Robots Exclusion Protocol standard, crawlers match user agent tokens case-insensitively, and a crawler that finds a group naming it obeys that group and ignores User-agent: *12. That trips people up constantly: if you give GPTBot its own group, every shared disallow line from your * group has to be repeated inside it. The reverse trap is just as common. In our 261-site study, 11 of the 38 sites blocking Claude-SearchBot never named it; their * group said Disallow: / and the newer bot fell through to it.

Meta, Mistral AI and Amazon split their bots the same way OpenAI does. Meta separates its training and indexing crawler from meta-externalfetcher, which may bypass robots.txt for user requests15. Mistral AI documents separate training, index and user agents16. Microsoft lists bingbot as its standard crawler17.

What does blocking each AI crawler cost you?

Blocking a training crawler costs little visibility today. Blocking a search crawler takes you out of that engine's answers. Blocking a user fetcher stops the engine reading a page a user explicitly asked about. The table uses each operator's own wording where it has one.

If you blockYou loseYou keep
OAI-SearchBotAppearing in ChatGPT search answers (OpenAI: such sites "will not be shown")1Navigational links can still appear
GPTBotFuture OpenAI training on your new pagesChatGPT search eligibility; settings are independent1
Claude-SearchBotIndexing "for search optimization"; possibly visibility in Claude's results2Claude-User fetches if allowed
ClaudeBotFuture Anthropic training on your materials2Claude search eligibility
PerplexityBotBeing surfaced and linked in Perplexity results3Nothing in Perplexity search; it is not a training crawler
GooglebotGoogle Search, including AI Overviews and AI Mode, and likely Brave Search tooNothing in Google
Google-ExtendedGemini training and grounding in the Gemini app and Vertex AI4Google Search, AI Overviews and AI Mode
Applebot-ExtendedUse in Apple foundation model training5Apple search features
bingbotBing and Copilot citations6Nothing in Bing; use NOCACHE or NOARCHIVE for finer control7
DuckAssistBotDuckDuckGo's AI-assisted answersOrganic DuckDuckGo rankings, per DuckDuckGo10

The Googlebot row has a side effect most people miss. Brave Search, which Claude appears to use for web search, says its crawler won't fetch anything Googlebot isn't allowed to18. For the full engine-by-engine picture, read how AI engines choose which sources to cite.

Which AI crawlers do the New York Times, BBC, Reddit and GitHub block?

The big news publishers block AI search crawlers along with training bots, while GitHub and Wikipedia leave AI search open. That split holds at scale. In our study of 261 major websites' robots.txt files, 88% of news sites block at least one training crawler and 69% block an AI search crawler, against 12% and 1% of B2B SaaS sites.

We fetched the live robots.txt of eight well-known sites with curl while researching this guide and pulled out the AI-related lines. The files are public and change often, so open them yourself before copying anything.

The New York Times blocks training and AI search, with one exception

The file opens with a comment listing prohibited uses, including the development of "machine learning, artificial intelligence (AI), and/or large language models." Then it gives most AI agents their own Disallow: / group:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: PerplexityBot
Disallow: /

ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User and Perplexity-User get the same treatment. Amazonbot is the odd one out: its group only says Disallow: /wirecutter/. That matches the multiyear AI licensing deal the Times signed with Amazon in May 2025, which reportedly excluded Wirecutter19. Robots.txt here reads like a contract summary.

CNN puts 77 user agents in one block

CNN uses a single group: 77 User-agent lines, from AI2Bot to YouBot, sharing one Disallow: /. It includes GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Google-Extended, PerplexityBot, Perplexity-User, Amzn-SearchBot and DuckAssistBot. Googlebot isn't on the list, so CNN stays in Google Search and AI Overviews while opting out of nearly every other AI engine.

The BBC bans AI search in plain words

The BBC's header is the most explicit we found. Among its listed prohibitions:

# - No retrieval-augmented generation (RAG), AI-powered search, agentic AI or grounding using BBC content

The rules follow through, with a carve-out for one path:

User-agent: OAI-SearchBot
Disallow: /
Allow: /storyworks
Allow: /storyworks/

GPTBot, ChatGPT-User and Google-Extended get the same /storyworks exception; PerplexityBot and Perplexity-User are blocked outright. StoryWorks is the BBC's branded-content studio, so the one section made for advertisers who want reach stays open to AI engines.

The Guardian blocks Claude and Perplexity but not OpenAI

The Guardian's big disallow group names ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, CCBot, Applebot-Extended, meta-externalagent, Amazonbot, DuckAssistBot and Google-CloudVertexBot, among others. GPTBot and OAI-SearchBot aren't in it. In February 2025 the Guardian and OpenAI announced a content partnership that brings Guardian reporting into ChatGPT with attribution and links20. We can't see the contract, but the robots.txt lines up with it.

Reddit blocks everyone

The file Reddit served to our request, minus comments, is two lines:

User-agent: *
Disallow: /

Reddit made this change in mid-2024, and Microsoft confirmed Bing stopped crawling it21. Reddit threads still show up heavily in AI answers because Reddit licenses its data instead, including a partnership with Google announced in February 202422. For everyone else, robots.txt is the only lever they've got.

Medium blocks training and leaves AI search open

Medium's AI group is close to what we'd recommend for most publishers:

User-Agent: Amazonbot
User-Agent: Applebot-Extended
User-Agent: Bytespider
User-Agent: ClaudeBot
User-Agent: FacebookBot
User-Agent: GoogleOther
User-Agent: GPTBot
User-Agent: meta-externalagent
Disallow: /
Allow: /about
Allow: /membership

(The real group allows a few more marketing paths.) OAI-SearchBot, Claude-SearchBot and PerplexityBot aren't named, so they fall through to Medium's normal * rules and can index articles for AI search.

GitHub rate-limits AI bots and fences off repo internals

GitHub groups GPTBot, OAI-SearchBot, ClaudeBot, anthropic-ai and PerplexityBot together, sets Crawl-delay: 1, explicitly allows marketing pages such as /about, /pricing and /features, and disallows a long list of repository views like /*/blame/, /*/raw/ and /*/commits/*?author. It's the clearest example of path-level control on the list: let AI bots read what you want repeated, keep them off the expensive or noisy URLs.

Wikipedia has no AI rules at all

Wikipedia's robots.txt is mostly about misbehaving crawlers that go "way too fast" and internal paths like /w/ and /wiki/Special:. It doesn't name a single AI agent. That openness pays off in citations. In Profound's analysis of ChatGPT citations from August 2024 to June 2025, Wikipedia was the single most cited source, with 7.8% of all citations26.

Publishers that sell or license content (the Times, CNN, the BBC) block AI search crawlers too, accepting absence from ChatGPT search as the price of their licensing position. Old tokens also linger: anthropic-ai still appears in the Times, CNN, BBC, Guardian and GitHub files even though Anthropic no longer documents it, and Claude-Web shows up in three of them. Robots.txt files get copied from lists and rarely pruned.

Our takeDon't copy a newspaper's robots.txt. The Times, CNN and the BBC are negotiating with AI companies over content they sell, so blocking AI search is a bargaining chip for them. If you sell software or services, the same file would hide you from ChatGPT, Claude and Perplexity in exchange for nothing. Medium's pattern (block training agents by name, leave search crawlers on the default rules) is the better starting point for almost everyone else.

Should you block GPTBot?

Block GPTBot if your content is what you sell; allow it if your content exists to sell something else.

Block GPTBot (and ClaudeBot, CCBot, Google-Extended) when at least one of these is true:

  1. Your pages are the product: paywalled journalism, paid research, course material, proprietary datasets.
  2. You license content to AI companies or plan to, and an open door would weaken that negotiation.
  3. A legal or contractual obligation (client agreements, rights you don't own) bars machine reuse of the content.

Keep GPTBot allowed when all of these are true:

  1. You sell software, a service or a physical product, and your public pages are marketing, docs or help content. Most B2B SaaS companies already sit here: only 4 of the 95 SaaS sites in our study block GPTBot.
  2. You want ChatGPT to describe you correctly when it answers without searching. In Slate's 1,000-question test, 372 of ChatGPT's answers carried no link at all23, so the model's training data was the only thing those answers had.
  3. Nothing on the allowed paths is confidential (if it is, robots.txt is the wrong tool; put it behind a login).

Mixed cases get path rules. A SaaS company with a paid research library can allow GPTBot on / and disallow /reports/, the way the BBC carves out /storyworks in the other direction:

User-agent: GPTBot
Disallow: /reports/
Allow: /

Whatever you decide about GPTBot, don't touch OAI-SearchBot unless you actually want out of ChatGPT search. Blocking it removes you from ChatGPT search answers and has no effect on training, which GPTBot governs1. If you're a B2B software company unsure whether ChatGPT already names you for your category, run your buyer prompts through the free Arobis AI visibility checker before and after any robots.txt change.

Does blocking Google-Extended remove you from AI Overviews?

No. Google's crawler documentation states that "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search"4. AI Overviews and AI Mode are Search features, and Google's AI features guide names Googlebot's robots.txt rules as the control8.

Google-Extended manages whether content may train future Gemini models and be used for grounding (supplying content to the model at prompt time) in the Gemini app and Grounding with Google Search on Vertex AI4. So blocking it is a clean training opt-out with one real cost: Gemini app citations. CNN, the Times and the BBC all block it; the Guardian and Medium don't name it. Across the 261 sites in our study, 53 (20.3%) block Google-Extended, the second least-blocked of the eight training agents we tested.

To limit what appears in AI Overviews specifically, Google points to nosnippet, data-nosnippet, max-snippet and noindex8. These also trim your regular search snippets, and a page has to be snippet-eligible to be an AI supporting link at all8.

Which robots.txt rules should I use?

Most sites that want AI visibility should allow search crawlers and user fetchers and make the training call separately. These three recipes cover the common positions. Swap in your sitemap URL and test the result with our robots.txt AI checker.

Recipe 1: allow AI search, block training

# Training crawlers and training-control tokens: blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
User-agent: MistralAI-Training
Disallow: /

# AI search crawlers and user fetchers: allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: DuckAssistBot
User-agent: MistralAI-Index
User-agent: MistralAI-User
User-agent: Amzn-SearchBot
Allow: /

# Everyone else, including Googlebot, bingbot and Applebot
User-agent: *
Allow: /

Sitemap: https://example.com/sitemap.xml

Blocking Google-Extended also removes you from grounding in the Gemini app, so drop that line if Gemini citations matter. Applebot and Amazonbot may feed training too, per Apple and Amazon, but blocking them also hits Apple search features and Amazon products, so they stay allowed here59. Stacking several User-agent lines over one rule is valid under RFC 930912.

Recipe 2: allow everything

User-agent: *
Allow: /

Sitemap: https://example.com/sitemap.xml

This is the same as having no rules, training crawlers included. It's the right file for most B2B software sites. Check that your CDN's bot-management settings aren't challenging AI crawlers at the network level, because robots.txt can't override a firewall.

Recipe 3: block all AI crawlers

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: meta-externalfetcher
User-agent: CCBot
User-agent: Amazonbot
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: DuckAssistBot
User-agent: MistralAI-User
User-agent: MistralAI-Index
User-agent: MistralAI-Training
Disallow: /

User-agent: *
Allow: /

This is roughly what CNN does. You'll disappear from ChatGPT search, Claude search, Perplexity and DuckDuckGo's AI answers. Googlebot and bingbot stay allowed, so AI Overviews, AI Mode and Copilot can still cite you; blocking those would take you out of Google and Bing entirely. User fetchers that ignore robots.txt need a firewall or CDN rule.

Where robots.txt stops working

Robots.txt is voluntary. Operators that honor it say so in their docs, but anyone can put "GPTBot" in a user agent string, which is why the major operators publish IP ranges (next section).

Changes take time to land. OpenAI says about 24 hours for ChatGPT search1, Amazon says about 24 hours, and DuckDuckGo says up to 72910. Each subdomain needs its own file, as Anthropic and Amazon both point out, so docs. and help. hosts are easy to forget29. Blocks don't reach back in time either: Anthropic describes a ClaudeBot block as a signal about future materials2.

Compliance is disputed for at least one operator. In August 2025 Cloudflare reported that Perplexity used undeclared crawlers to get around no-crawl rules; Perplexity rejected the findings13. And a disallowed URL can still be listed: OpenAI's publisher FAQ says it may show the link and title of a blocked page found through a third-party search provider, and recommends noindex if you want it gone14. If you'd also like to hand AI tools a curated map of your site, see our explainer on llms.txt.

How do I verify AI crawlers in my server logs?

Find the hits by user agent, then prove they're real by IP. Start by counting requests per bot and IP in your access log (this assumes the common combined log format, where the IP is the first field):

awk 'match($0, /GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|bingbot|Googlebot/) {
  print $1, substr($0, RSTART, RLENGTH) }' access.log | sort | uniq -c | sort -rn | head -30

For Google, Bing, Apple and Common Crawl, use the two-step DNS check. Reverse-resolve the IP, confirm the hostname is on the operator's domain, then forward-resolve that hostname and confirm it returns the same IP. We ran it on two IPs from the published ranges:

$ host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
$ host 157.55.39.1
1.39.55.157.in-addr.arpa domain name pointer msnbot-157-55-39-1.search.msn.com.

OpenAI, Anthropic and Perplexity don't document reverse DNS, so check the IP against their JSON lists instead. A few lines of Python do it:

import ipaddress, json, urllib.request
req = urllib.request.Request("https://openai.com/searchbot.json", headers={"User-Agent": "Mozilla/5.0"})
ranges = json.load(urllib.request.urlopen(req))["prefixes"]
nets = [ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix")) for p in ranges]
ip = ipaddress.ip_address("203.0.113.7")   # the IP from your log
print(any(ip in n for n in nets))
OperatorPublished IP listReverse DNS should end in
OpenAIgptbot.json, searchbot.json, chatgpt-user.json1Not documented; use the IP lists
Anthropicclaude.com/crawling/bots.json2Not documented; use the IP list
Perplexityperplexitybot.json, perplexity-user.json3Not documented; use the IP lists
Googlecommon-crawlers.json and the other files listed on Verifying Googlebot24googlebot.com, google.com or googleusercontent.com
Microsoftbingbot.json, plus the Verify Bingbot toolsearch.msn.com25
Appleapplebot.json5applebot.apple.com
Common Crawlccbot.json, linked from its CCBot page11crawl.commoncrawl.org (IPv4 only)

Then read the status codes. A verified OAI-SearchBot or PerplexityBot IP getting 403s or challenge pages means a firewall rule is overriding your robots.txt; Perplexity's docs tell Cloudflare users to create an Allow rule for its IP ranges3. Verified hits from ChatGPT-User or Claude-User are worth watching for another reason: each one is a real person asking an assistant about that URL. Pair the log work with prompt testing, since a crawler visit shows access and only a citation shows use; our guide to measuring AI visibility covers that half.

Which user agents did we leave out?

We include a token only when its operator documents it. Three widely copied tokens miss that bar. anthropic-ai isn't in Anthropic's current crawler documentation, which lists ClaudeBot, Claude-User and Claude-SearchBot2, yet five of the eight major sites above still block it. We couldn't find first-party documentation from ByteDance for Bytespider or from Cohere for cohere-ai. Blocking them does no harm, but we can't tell you what they do in their operators' own words.

Frequently asked questions

What is the difference between GPTBot and OAI-SearchBot?

GPTBot crawls content that may be used to train OpenAI's foundation models. OAI-SearchBot crawls content to surface websites in ChatGPT's search features. OpenAI says the settings are independent, so you can block GPTBot to opt out of training and still allow OAI-SearchBot, which keeps your pages eligible to appear and be cited in ChatGPT search answers.

Does blocking Google-Extended remove my site from AI Overviews?

No. Google says Google-Extended does not affect inclusion in Google Search and is not a ranking signal. AI Overviews and AI Mode are part of Search and are crawled by Googlebot. Google-Extended controls whether your content trains Gemini models and grounds answers in the Gemini app and Vertex AI, so blocking it costs Gemini app visibility only.

Do AI user agents like ChatGPT-User follow robots.txt?

Not reliably. OpenAI says robots.txt rules may not apply to ChatGPT-User because a user started the visit, Perplexity says Perplexity-User generally ignores robots.txt, and Meta says meta-externalfetcher may bypass it. Anthropic says its bots honor robots.txt. If you need a hard block on user-triggered fetchers, use a CDN or firewall rule matched to the operator's published IP ranges.

Should I block GPTBot?

Block it if your content is the product you sell or license, as with paywalled news, paid research or course material. Keep it allowed if you sell software or services and want ChatGPT to describe you accurately when it answers without searching, which happened on 372 of 1,000 answers in one 2026 test. Either way, keep OAI-SearchBot allowed if you want ChatGPT search citations.

How do I know a request claiming to be GPTBot is real?

Match the request IP against the operator's published list: openai.com/gptbot.json for GPTBot, with separate files for OAI-SearchBot and ChatGPT-User. Google, Bing, Apple and Common Crawl also support reverse DNS: the IP should resolve to their crawler domain, and that hostname should resolve back to the same IP. A user agent string alone proves nothing, since anyone can send it.

Which bot does Microsoft Copilot use?

Copilot sends generated search queries to Bing, so bingbot is the crawler that matters. Blocking bingbot removes your pages from Bing and from Copilot citations. To limit how Copilot uses a page while staying in Bing results, Bing supports the NOCACHE and NOARCHIVE robots meta values. Real bingbot traffic reverse-resolves to a search.msn.com hostname.

How many websites block AI crawlers?

About a third of large sites block at least one. In AI Ranked's study of 261 major websites, 35% disallow at least one AI training crawler and 21% disallow at least one AI search crawler. News publishers account for most of it, with 88% blocking a training bot, while 94 of 95 B2B SaaS sites leave all three AI search crawlers open.

Sources

  1. OpenAI, Overview of OpenAI crawlers
  2. Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?
  3. Perplexity, Perplexity crawlers
  4. Google Search Central, Google's common crawlers (Google-Extended, Google-CloudVertexBot)
  5. Apple Support, About Applebot
  6. Microsoft Learn, Data, privacy, and security for web search in Microsoft Copilot
  7. Bing Webmaster Blog, New options for webmasters to control usage of their content in Bing Chat (September 2023)
  8. Google Search Central, AI features and your website
  9. Amazon Developer, Amazonbot and Amazon crawlers
  10. DuckDuckGo Help, DuckAssistBot
  11. Common Crawl, CCBot
  12. IETF, RFC 9309: Robots Exclusion Protocol (September 2022)
  13. Cloudflare, Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives (4 August 2025)
  14. OpenAI Help Center, Publishers and developers FAQ
  15. Meta for Developers, Meta web crawlers
  16. Mistral AI, Robots and user agents
  17. Bing Webmaster Help, Which crawlers does Bing use?
  18. Brave Search Help, Brave Search crawler
  19. Axios, NYT reaches AI licensing deal with Amazon (30 May 2025)
  20. OpenAI, OpenAI and Guardian Media Group launch content partnership (14 February 2025)
  21. Search Engine Land, Microsoft confirms Reddit blocked Bing search (July 2024)
  22. Google, Google expands partnership with Reddit (22 February 2024)
  23. Slate, How to get cited by ChatGPT, Perplexity, Claude and Gemini (1,000-question study, 27 August 2026)
  24. Google Search Central, Verifying Googlebot and other Google crawlers
  25. Bing Webmaster Blog, How to verify that Bingbot is Bingbot (31 August 2012)
  26. Profound, AI Platform Citation Patterns (data August 2024 to June 2025)
Cite this page AI Ranked Editorial. "AI crawlers list: every user agent for robots.txt." AI Ranked, October 11, 2026. https://airanked.ai/guides/ai-crawlers-robots-txt