AI Ranked
Free tool

Free robots.txt checker for AI crawlers

Paste your robots.txt and see, bot by bot, whether ChatGPT, Claude, Perplexity, Gemini and 17 other AI user agents can reach your pages.

Browsers block one site from reading another site's files (CORS), so this page cannot fetch your robots.txt for you. Open it, copy everything, paste it above. It always lives at the root: https://yourdomain.com/robots.txt

Result for /

User agentPurposeStatus

"Rule" shows the line that decided it. "No group" means the file never names that bot and has no * group, so everything is allowed.

Recommended config

Policy

Search crawlers decide whether you can be cited in ChatGPT search, Claude, Perplexity and similar answers. Training crawlers only affect what future models learn. Blocking training does not remove you from AI search results.

Before adding this, delete any existing groups for these user agents. Groups that name the same bot get merged, so an old Disallow: / would still apply. Rules from your * group are copied into the allowed groups, because a bot that has its own group ignores *. Googlebot, Bingbot and Applebot are left out on purpose: they are general search crawlers that should keep following your normal rules.

Add to robots.txt


  

How the checker decides allowed or blocked

It applies the Robots Exclusion Protocol as written in RFC 9309, the same rules Google documents for its crawlers. Consecutive User-agent lines form one group. A bot obeys only the group that names it and falls back to * only when nothing names it. Groups naming the same bot are merged.

Inside a group the longest matching path wins, and on a tie Allow beats Disallow. * matches any run of characters and $ anchors the end, so Disallow: /*.pdf$ blocks /guide.pdf but not /guide.pdf?v=2. The Status column shows the exact line that decided each bot, so you can find it in your file.

Worked example: The Guardian lets ChatGPT in and keeps Claude out

The Guardian's live robots.txt is a good lesson in how much a file can say without saying it. The top carries a comment that LLM and AI uses of its content aren't permitted. Then comes one long group of named bots, ending in a single rule:

User-Agent: PerplexityBot
User-Agent: yacy
User-agent: anthropic-ai
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
...
User-agent: meta-externalagent
User-agent: Amazonbot
User-agent: DuckAssistBot
Disallow: /

Paste the full file above and test an article path. The checker reports 10 of 21 AI user agents allowed. PerplexityBot, all three Anthropic agents, DuckAssistBot and Amazon's bots are blocked by that group. GPTBot, OAI-SearchBot and ChatGPT-User aren't named anywhere, so they fall through to the * group, which only blocks paths like /search and /discussion/. Google-Extended isn't named either.

That's a business decision written as a robots.txt. Guardian Media Group signed a content partnership with OpenAI in February 2025, with Guardian reporting shown in ChatGPT with attribution. Two things to take from it. The comment at the top has no effect on any crawler, since only User-agent and Allow/Disallow lines count. And silence is permission. A bot you don't name gets your * rules.

Compare The New York Times, whose file gives GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Perplexity-User and Google-Extended each their own Disallow: /. That blocks the search agents too, which means opting out of being a cited source in ChatGPT search and Perplexity. For a publisher negotiating licences that can be the point. For a SaaS company it would be self-harm.

Which blocked rows should you fix first?

Fix a blocked AI search agent or a blocked Googlebot or Bingbot first; a blocked training crawler can wait, because it may be a deliberate choice. The table mixes three kinds of agent, a split documented by OpenAI, Anthropic and Perplexity.

If this is blockedWhat you loseDo this
An AI search agent (OAI-SearchBot, Claude-SearchBot, PerplexityBot) on pages you want foundEligibility to be cited in that engine's answersFix today, unless you blocked it on purpose for licensing reasons
Googlebot or BingbotGoogle Search, AI Overviews and AI Mode, or Bing and CopilotTreat it as an outage
A user fetch agent (ChatGPT-User, Claude-User)Pages a person pastes or asks about may not loadAllow unless the path is private; several of these ignore robots.txt anyway
A training crawler (GPTBot, ClaudeBot, CCBot) or token (Google-Extended)Influence on what future models know without searchingA policy choice, see below

Don't test only /. Put your pricing page, a top blog post and a docs page in the path box. Old rules such as Disallow: /blog/*? or a blanket block on /docs/ tend to hit exactly the pages engines would cite.

Should you block GPTBot? A decision rule

Block training crawlers if your content is the product: you sell subscriptions or licences, or a model that memorized your pages would replace a visit. Allow them if you're a brand that wants models to describe your category with you in it, which covers most B2B software companies. Either way, keep the search agents open.

That's what most large sites that block training already do. In our robots.txt study of 261 top sites, 31 of the 59 that disallow GPTBot still allow OAI-SearchBot, and every site that blocked an AI search bot also blocked a training bot. B2B SaaS sites barely block at all: 11 of 95 block any training bot.

Our takeFor a SaaS company, blocking GPTBot and ClaudeBot buys almost nothing. Your pricing and feature pages aren't a licensable asset, and models that never saw them will describe you from third-party reviews and old forum threads instead. Block training only when you can name what you're protecting.

Four mistakes that block AI search by accident

A leftover staging rule. User-agent: * plus Disallow: / shipped to production blocks every bot you haven't named, including all the AI search agents. The checker shows this as a wall of red rows reading "Disallow: / (* group)".

Giving a bot its own group and losing your other rules. Add User-agent: GPTBot with Allow: / and GPTBot stops reading your * group, so it can now crawl /admin/ and /cart. The recommended config above copies your * rules into the named group for that reason.

Expecting Google-Extended to keep you out of AI Overviews. It doesn't. Google says the token only covers Gemini training and grounding and has no effect on Search. AI Overviews and AI Mode are part of Search, and Google's AI features documentation points to nosnippet, data-nosnippet, max-snippet or noindex to limit what appears there.

A clean file and a blocking firewall. robots.txt is a request, and your CDN may refuse bots before they ever read it. Cloudflare began blocking known AI crawlers by default for new domains on July 1, 2025. If the checker says allowed but you never see the bot, search your access logs for the user agent (for example grep -c "OAI-SearchBot" access.log) and check your bot-management settings.

Once access is right, the next questions are whether a page is worth quoting, which the citability grader scores, and whether engines name you at all. The free Arobis AI visibility checker tests the second one against real buyer prompts, and our comparison of 10 free AI visibility checkers covers the alternatives. The full user-agent list, with IP verification, is in our AI crawler and robots.txt guide, and the OAI-SearchBot glossary entry has the short version.

Frequently asked questions

Why can't this tool fetch my robots.txt automatically?

Browsers enforce CORS, a security rule that stops a page on one domain reading files on another unless that site allows it. Most sites don't, so a browser-only tool can't load your file. Open yourdomain.com/robots.txt in a new tab, copy the whole file and paste it into the box. Nothing you paste leaves your browser.

Does blocking GPTBot remove my site from ChatGPT?

No. OpenAI documents GPTBot as the crawler that collects content which may train its models. ChatGPT search uses a different agent, OAI-SearchBot, and pages a user asks about are fetched by ChatGPT-User. To stay citable in ChatGPT search while opting out of training, disallow GPTBot and leave OAI-SearchBot allowed.

Do AI user agents always obey robots.txt?

The indexing crawlers are documented as respecting it: GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot among them. User-initiated fetchers are a different story. OpenAI says robots.txt may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores it, and Meta says meta-externalfetcher may bypass it, because a person asked for the page.

If I name a bot in its own group, does it still follow my * rules?

No. Under RFC 9309 a crawler obeys only the group that names it and ignores the * group entirely. If your * group blocks /admin/ and you add a group allowing GPTBot, GPTBot can now crawl /admin/. Copy any rules you still need into the named group, which is what the recommended config on this page does.

How long do robots.txt changes take to apply?

Usually about a day. RFC 9309 lets crawlers cache robots.txt for up to 24 hours, OpenAI says OAI-SearchBot changes take about 24 hours to show in search, and Meta gives the same figure. DuckDuckGo is slower: it says DuckAssistBot stops about 72 hours after an opt-out. Pages already indexed may stay cited until the engine re-crawls.