- 170 of 261 sites (65%) block none of the 14 AI user agents we tested. Only 12 (4.6%) block all 14.
- Half of GPTBot blockers keep ChatGPT search open: 31 of the 59 sites that disallow GPTBot still allow OAI-SearchBot.
- News is the outlier. 42 of 48 news sites (88%) block a training bot and 33 (69%) block a search bot, against 11 of 95 (12%) and 1 of 95 for B2B SaaS.
- Every one of the 56 sites that blocks an AI search bot also blocks at least one training bot. Nobody blocks search while allowing training.
- 81 of 241 sites (34%) serve a valid llms.txt. B2B SaaS leads at 64%; news sits at 8%.
How many top websites block AI crawlers?
About one in three blocks something, and about one in five blocks a bot that decides whether the site can be cited in AI answers. We fetched robots.txt from 285 domains and could read 261. Of those, 91 (35%) disallow at least one of the eight training agents we tested, and 56 (21%) disallow at least one of the three AI search agents: OAI-SearchBot, Claude-SearchBot and PerplexityBot.
Blocking tends to be all or nothing. 26 sites block all eight training agents, while 170 block none. The middle is thin: just 12 sites block exactly one training agent, and for eight of those the one bot is Bytespider or CCBot.
Googlebot is our control, and it behaves like one. Exactly 1 of 261 robots.txt files we received blocks Googlebot at the homepage, and that file (Reddit's) blocks everyone. More on that below.
Which AI crawlers get blocked most often?
Training crawlers top the list, led by Common Crawl's CCBot and ByteDance's Bytespider at 75 sites each. The three AI search crawlers sit near the bottom, and OpenAI's OAI-SearchBot is the least blocked AI agent of all.
/. Bars drawn to scale on a 0 to 30% axis. Googlebot is a control. Data collected October 2026 by AI Ranked; download the CSV.The table splits each count by how the block happens. "Named group" means the file has a User-agent line for that bot. "Fallback" means the bot isn't named, so under RFC 9309 it obeys the * group, and that group says Disallow: /8.
| User agent | Operator | Type | Blocked at / | By a named group | By the * fallback |
|---|---|---|---|---|---|
Bytespider | ByteDance | Training | 75 (28.7%) | 62 | 13 |
CCBot | Common Crawl | Training | 75 (28.7%) | 63 | 12 |
ClaudeBot | Anthropic | Training | 67 (25.7%) | 59 | 8 |
GPTBot | OpenAI | Training | 59 (22.6%) | 53 | 6 |
Applebot-Extended | Apple | Training | 58 (22.2%) | 48 | 10 |
meta-externalagent | Meta | Training | 58 (22.2%) | 47 | 11 |
Google-Extended | Training | 53 (20.3%) | 46 | 7 | |
Amazonbot | Amazon | Training | 52 (19.9%) | 42 | 10 |
PerplexityBot | Perplexity | AI search | 50 (19.2%) | 42 | 8 |
ChatGPT-User | OpenAI | User fetch | 43 (16.5%) | 35 | 8 |
Claude-User | Anthropic | User fetch | 43 (16.5%) | 32 | 11 |
Perplexity-User | Perplexity | User fetch | 40 (15.3%) | 28 | 12 |
Claude-SearchBot | Anthropic | AI search | 38 (14.6%) | 27 | 11 |
OAI-SearchBot | OpenAI | AI search | 30 (11.5%) | 22 | 8 |
Googlebot | Control | 1 (0.4%) | 0 | 1 |
ClaudeBot (67) is blocked more often than GPTBot (59), even though both are training crawlers and both operators document a separate search bot. The fallback column is large for the newer agents: 11 of the 38 Claude-SearchBot blocks and 12 of the 40 Perplexity-User blocks come from a site that never named the bot at all. Those sites wrote rules for the older tokens and a catch-all deny for everything else.
Do sites block training bots but allow AI search bots?
Yes, and that's the most useful pattern in the data. OpenAI, Anthropic and Perplexity each run separate agents for training, search indexing and user-requested fetches, and they document them separately123. A site can refuse training and still be eligible for citations.
- 59 sites block GPTBot. 31 of them (53%) still allow OAI-SearchBot.
- 67 sites block ClaudeBot. 31 of them (46%) still allow Claude-SearchBot.
- 91 sites block at least one training agent. 35 of them (38%) allow all three AI search agents.
- 56 sites block at least one AI search agent, and all 56 also block a training agent. We found zero sites that allow training and block search.
- 21 sites (8%) block all three AI search agents.
The vendors are not treated equally. 13 sites block PerplexityBot but allow GPTBot, and another 13 block Claude-SearchBot but allow OAI-SearchBot. Only 5 sites do the reverse (block OAI-SearchBot, allow Claude-SearchBot): Gartner, CNBC, bbc.com, bbc.co.uk and the Sydney Morning Herald. Across the three search agents, PerplexityBot is blocked on 50 sites, Claude-SearchBot on 38 and OAI-SearchBot on 30.
Which kinds of websites block AI crawlers?
News and social platforms block the most; developer, government and B2B SaaS sites block almost nothing. The gap is large enough that a single blended number hides it, so here it is per category.
| Category | Sites read | Block 1+ training bot | Block 1+ AI search bot | Block GPTBot | Block OAI-SearchBot | Block ClaudeBot |
|---|---|---|---|---|---|---|
| News & media | 48 of 53 | 42 (88%) | 33 (69%) | 27 (56%) | 15 (31%) | 37 (77%) |
| Social & community | 17 of 17 | 12 (71%) | 9 (53%) | 9 (53%) | 8 (47%) | 10 (59%) |
| Video & streaming | 9 of 10 | 4 (44%) | 2 (22%) | 1 (11%) | 1 (11%) | 1 (11%) |
| Reference & education | 20 of 25 | 8 (40%) | 4 (20%) | 6 (30%) | 1 (5%) | 6 (30%) |
| Travel & local | 8 of 8 | 3 (38%) | 1 (12%) | 3 (38%) | 1 (12%) | 3 (38%) |
| Search & portals | 11 of 11 | 4 (36%) | 3 (27%) | 4 (36%) | 2 (18%) | 4 (36%) |
| E-commerce & marketplaces | 20 of 23 | 5 (25%) | 3 (15%) | 5 (25%) | 1 (5%) | 4 (20%) |
| B2B SaaS | 95 of 102 | 11 (12%) | 1 (1%) | 4 (4%) | 1 (1%) | 2 (2%) |
| Government & public sector | 13 of 14 | 1 (8%) | 0 (0%) | 0 (0%) | 0 (0%) | 0 (0%) |
| Tech & developer | 20 of 22 | 1 (5%) | 0 (0%) | 0 (0%) | 0 (0%) | 0 (0%) |
| All sites | 261 of 285 | 91 (35%) | 56 (21%) | 59 (23%) | 30 (11%) | 67 (26%) |
Percentages use the sites we could read in each category as the base. Small categories (8 to 11 sites) swing a lot on one site, so read them as direction, not precision.
Nearly all the blocking happens in news. Of the 48 news sites we read, only six block none of the 14 AI agents: Fox News, Time, The Independent, Sky, Hindustan Times and the South China Morning Post. Nine block training while leaving all three search agents open, including CBS News, ABC News, Business Insider, Axios, the Los Angeles Times and TechCrunch.
B2B SaaS is the opposite. 84 of the 95 SaaS sites we read block none of the AI agents. The 11 that block something mostly target scrapers: Bytespider alone at Amplitude, Postman, G2 and Capterra, CCBot at Calendly, Amazonbot at Notion. Canva and Figma block several training agents. Gartner, which we grouped with B2B software, is the only one in the group that blocks an AI search agent (OAI-SearchBot).
Among tech and developer sites, the only block is GitHub's on Bytespider, and among government sites it's the World Health Organization's on CCBot. Wikipedia, Microsoft, Mozilla and every US and UK government domain we read allow all 14.
How many top websites have an llms.txt file?
34%: 81 of the 241 domains that gave us a definite answer. llms.txt is a proposed standard for a Markdown file at the site root that points language models to a site's key content10. We counted it only if the response was HTTP 200, the content type was text/plain or text/markdown, and the file opened with a Markdown H1 (# followed by a space). That rules out sites that redirect /llms.txt to an HTML page.
| Category | Sites with a definite answer | Valid llms.txt | Adoption |
|---|---|---|---|
| B2B SaaS | 92 | 59 | 64% |
| Tech & developer | 18 | 9 | 50% |
| Travel & local | 7 | 2 | 29% |
| E-commerce & marketplaces | 18 | 5 | 28% |
| Video & streaming | 8 | 1 | 12% |
| Reference & education | 19 | 2 | 11% |
| News & media | 40 | 3 | 8% |
| Search & portals | 10 | 0 | 0% |
| Social & community | 16 | 0 | 0% |
| Government & public sector | 13 | 0 | 0% |
| All sites | 241 | 81 | 34% |
"Definite answer" means the server returned 200, 404 or 410. We left out 44 domains that answered with 403, 429, another error or a timeout, because a refusal tells you nothing about whether the file exists. OpenAI's and Perplexity's own domains returned 403 to our fetcher, so they are not in the base.
Adoption tracks the people who sell to developers and marketers. 59 of 92 SaaS sites have one, including Stripe, HubSpot, Atlassian, Vercel and Shopify. Six more SaaS sites (Salesforce, Notion, Twilio, Greenhouse, BILL and PagerDuty) serve a plain text file at /llms.txt that doesn't start with an H1, so our rule didn't count them. Among news sites, only Fox News, The Times of India and Hindustan Times pass.
Ten sites publish an llms.txt and also block at least one training crawler, including GitHub, Coursera, ZoomInfo and Loom. Those aren't contradictory. llms.txt tells a model where the useful pages are; robots.txt decides which bots may fetch them.
Want one for your own site? Our llms.txt guide covers what the file does and doesn't do, and the llms.txt generator builds a valid file from your page list.
What do the actual robots.txt rules look like?
Here's how ten well-known sites express their policy. Every rule below is quoted from the file our fetcher received.
- The New York Times gives each AI agent its own group ending in
Disallow: /, includingUser-agent: OAI-SearchBot,Claude-SearchBotandPerplexityBot. Googlebot shares the*group, which only blocks paths like/ads/. Result: all 14 AI agents blocked, Google allowed. - The Guardian lists about 30 agents in a single group that ends in
Disallow: /: ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Applebot-Extended, meta-externalagent, Amazonbot and others. GPTBot, OAI-SearchBot and ChatGPT-User are not in that group, so they fall back to the permissive*rules and can crawl. - BBC (bbc.com) blocks GPTBot and OAI-SearchBot with
Disallow: /plusAllow: /storyworks, and blocks ClaudeBot outright. It never names Claude-SearchBot or Claude-User, so both are allowed through the*group. - Bloomberg labels its groups with comments,
# Trainingabove GPTBot and# Search indexingabove OAI-SearchBot. The search group readsDisallow: /followed byAllow: /$and a few section allows. The homepage is open to ChatGPT search; articles are not. - Facebook and Instagram end with
User-agent: */Disallow: /. Both name GPTBot, ClaudeBot and PerplexityBot with path-level rules, so those can reach the homepage. OAI-SearchBot, Claude-SearchBot and even Meta's ownmeta-externalagentaren't named, so the catch-all blocks them. - Netflix runs default-deny (
User-agent: */Disallow: /), then opens a group for GPTBot, ChatGPT-User, OAI-SearchBot, Google-Extended, ClaudeBot, Claude-User, Claude-SearchBot and PerplexityBot withAllow: /. Agents it didn't list, such as Perplexity-User and Applebot-Extended, stay blocked. - LinkedIn blocks GPTBot, ClaudeBot and PerplexityBot with
Disallow: /, but gives OAI-SearchBot its own group of path-level disallows such as/public-profile/. ChatGPT search can reach the homepage; OpenAI's trainer cannot. - Reddit served our fetcher a two-line policy:
User-agent: */Disallow: /. Taken literally that blocks Googlebot too, and it's the only Googlebot block in the dataset. Reddit pages still rank in Google, which tells you the file a generic fetcher sees isn't always the whole access story. - Canva groups GPTBot, ClaudeBot, CCBot and Applebot-Extended under
Disallow: /withAllow: /create/, while OAI-SearchBot, Claude-SearchBot and PerplexityBot stay allowed. A clean training-only block. - GitHub names GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot together with
Crawl-delay: 1and a list of allowed sections. It slows them down instead of shutting them out; Bytespider is the only agent it blocks at the homepage.
The fallback trap shows up again and again. 15 of the 261 sites run a default-deny * group, and on those sites any AI agent the file forgets to name is blocked. Facebook names three AI crawlers and blocks the search agents it left out. That might be intentional. It's also exactly what happens when a file gets a rule for GPTBot and then nobody touches it again.
What should you do with your own robots.txt?
Start by checking what your file does today: paste it into our robots.txt AI crawler checker, which uses the same parser as this study. Then apply these rules.
- If you sell software or services and want AI engines to recommend you, allow all three search agents and the user fetchers. 94 of the 95 SaaS sites we read leave OAI-SearchBot, Claude-SearchBot and PerplexityBot open. Blocking them puts you behind every competitor that doesn't.
- Block training agents only if your content is the product. That's the pattern among publishers whose pages are what they sell. If your pages exist to win customers, a training block protects nothing you'd charge for.
- If you block a vendor, name all of its agents explicitly, in both directions. BBC blocks ClaudeBot but, by leaving Claude-SearchBot unnamed, allows it. If that's what you want, write
User-agent: Claude-SearchBot/Allow: /so the next person editing the file can see it. - If your
*group saysDisallow: /, list every AI agent you want to allow. Under default-deny, every new agent is blocked until you name it. Re-check the list whenever OpenAI, Anthropic or Perplexity publish a new token. - Test a deep URL, not just the homepage.
Allow: /$opens the homepage and nothing else. When we re-ran every check against a generic article path, OAI-SearchBot blocks rose from 30 to 34 sites. - Don't rely on robots.txt for user fetchers. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores robots.txt, because a person asked for the page13. If you must stop those, do it at the CDN or firewall.
For the full list of AI user agents and what each one does, see our AI crawler user agent reference. Access is only step one. To see whether ChatGPT, Perplexity and Gemini actually mention your brand once they can reach you, run your buyer prompts through Arobis AI's free AI visibility checker, or compare the free and paid options in our roundup of AI visibility checkers.
How we collected and parsed the data
Sample. We built the list ourselves rather than downloading a ranking: 285 domains in 10 categories. 130 are large global sites across search and portals (11), social and community (17), video and streaming (10), e-commerce and marketplaces (23), travel and local (8), reference and education (25), tech and developer (22) and government and public sector (14). 102 are well-known B2B SaaS companies (the group also includes G2, Capterra and Gartner). 53 are news and media publishers from the US, UK, Europe, Asia and Australia. Each domain appears once. The full list with categories is in the dataset.
Collection. Data collected October 2026. A Python script requested https://<domain>/robots.txt and https://<domain>/llms.txt with the user agent CitedResearch/1.0 (robots.txt study) (the publication's working name before AI Ranked), a 10-second timeout, redirects followed and at most 8 requests in flight. Failed requests were retried once, including on the www. host.
What counted as read. A 200 response with a text body was parsed. A 404 or 410 means the site has no robots.txt, which RFC 9309 treats as permission to crawl8; two sites (disneyplus.com and telegram.org) fell here. 24 domains were excluded from every percentage: 17 refused our fetcher with an HTTP error (403, 405, 418 or 429), 3 returned an HTML challenge or redirect page instead of a text file, 3 timed out and 1 refused the connection. That leaves n = 261.
Parser. We ported the parser from our robots.txt checker to Python and evaluated each agent against the path /. Consecutive User-agent lines form one group, groups naming the same agent are merged, an agent with no named group falls back to *, the longest matching rule wins, Allow wins a tie, and * and $ work as wildcards. Agent names are matched case-insensitively. An agent counts as blocked if / is disallowed.
Agents. Training: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot, Bytespider, Amazonbot. AI search: OAI-SearchBot, Claude-SearchBot, PerplexityBot. User fetch: ChatGPT-User, Claude-User, Perplexity-User. Control: Googlebot. The groupings follow each operator's documentation123456711. ByteDance doesn't publicly document Bytespider's purpose, so we grouped it with training crawlers based on common practice.
Limitations
- robots.txt is a request, not an enforcement mechanism. A site that allows a bot here may still block it at the firewall or CDN. Cloudflare, for example, began blocking AI crawlers by default for new customers on 1 July 20259, and that kind of block never shows up in robots.txt.
- Some sites serve different robots.txt files to different requesters. We saw one file per site, the one served to a self-identified research fetcher. We can't see what Reddit serves to Googlebot, for example.
- We tested the homepage path only. Sites that block sections (or allow only
/$) are under- or over-counted for deep pages; the article-path rerun raised each agent's count by 1 to 4 sites. - The 24 excluded domains skew toward sites with aggressive bot protection, so the true blocking rate among very large sites is probably somewhat higher than what we measured.
- This is a curated list of well-known sites, not a random sample of the web. Our numbers describe big brands, not the average website. For wider context, see our AI search statistics.
Frequently asked questions
What percentage of top websites block GPTBot?
In our October 2026 sample, 59 of 261 major websites (22.6%) disallow GPTBot at the homepage path in robots.txt. News publishers drive most of it: 27 of 48 news sites block GPTBot, compared with 4 of 95 B2B SaaS sites and none of the 20 tech and developer sites we could read.
Which AI crawler is blocked most often?
CCBot (Common Crawl) and Bytespider (ByteDance) tie at 75 of 261 sites each, or 28.7%. ClaudeBot follows at 67 sites (25.7%), then GPTBot at 59 (22.6%). The least blocked AI agent is OAI-SearchBot, the crawler behind ChatGPT search, at 30 sites (11.5%).
Do sites that block GPTBot also block ChatGPT search?
About half do not. Of the 59 sites that block GPTBot, 31 still allow OAI-SearchBot, so they opt out of OpenAI training while staying eligible for ChatGPT search citations. Anthropic sees the same split: 31 of the 67 sites blocking ClaudeBot still allow Claude-SearchBot.
Which industries block AI crawlers the most?
News and media, by a wide margin. 42 of the 48 news sites we could read (88%) block at least one AI training crawler and 33 (69%) block at least one AI search crawler. Social platforms come next at 71% for training bots. B2B SaaS sits near the bottom: 11 of 95 block a training bot and only 1 blocks a search bot.
How many big websites have an llms.txt file?
81 of the 241 sites that gave a clear answer (34%) serve a valid llms.txt, meaning HTTP 200, a text/plain or text/markdown content type and a file that opens with a Markdown H1. B2B SaaS leads at 59 of 92 (64%). Only 3 of 40 news sites have one, and no government, search or social site in our sample does.
Can I download the data behind this study?
Yes. The full dataset is a CSV with one row per domain: category, whether the robots.txt fetch worked, the allowed or blocked status for all 15 user agents at the homepage path, and llms.txt status. It covers all 285 domains we tried, including the 24 we could not read, so you can rerun any number in this report.
Sources
- OpenAI, Overview of OpenAI crawlers
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity, Perplexity crawlers
- Google Search Central, Google's common crawlers (Google-Extended)
- Apple Support, About Applebot
- Meta for Developers, Meta web crawlers
- Common Crawl, CCBot
- IETF, RFC 9309: Robots Exclusion Protocol (September 2022)
- Cloudflare press release, Cloudflare just changed how AI crawlers scrape the internet-at-large (1 July 2025)
- llmstxt.org, The /llms.txt file proposal
- Amazon Developer, Amazonbot and Amazon crawlers
Dataset: ai-crawler-access-data.csv (285 rows, one per domain, including the 24 we could not read).
AI Ranked Editorial. "Which AI crawlers do top websites block? A study of 261 sites." AI Ranked, October 11, 2026. https://airanked.ai/research/ai-crawler-access-report