Robots.txt for AI crawlers

Published

AI assistants now crawl the web with bots of their own, and robots.txt is where you decide what each one may do: cite your pages, train on them, both, or neither. It is one small text file at the root of your site. There is no single "AI bot" to wave in or turn away: each company runs several crawlers split by job, so a rule that keeps you out of one AI product can quietly keep you out of another. And a bot you allow may still never reach you, because your CDN or firewall can block it before robots.txt is ever read.

What robots.txt is

robots.txt is a plain-text file at the root of a site (https://example.com/robots.txt) that tells crawlers which URLs they may fetch. It is a request, not a lock: a well-behaved crawler checks it and follows the rules it finds there.

The file is built from a handful of lines. User-agent names the crawler a group of rules applies to. Disallow lists paths that crawler may not fetch. Allow carves an exception out of a disallowed path. Sitemap points to your sitemap file. You will also see Crawl-delay in the wild; it is non-standard, and it gets its own section later.

Crawling is allowed by default. Google's spec puts it plainly: "By default, there are no restrictions for crawling" (Google robots.txt spec). You only write Disallow to take something away. Allow exists solely to punch a hole in a path you already blocked, so writing Allow: / for a bot you never disallowed does nothing at all.

The file also has important limits. Compliance is voluntary: as Google's introduction notes, "it's up to the crawler to obey them," and the same page warns that robots.txt "is not a mechanism for keeping a web page out of Google" (Google robots.txt intro). Crawling and indexing are different things: "A page that's disallowed in robots.txt can still be indexed if linked to from other sites." Whether a reachable page then actually gets indexed is a separate subject, which the companion guide to indexability covers in full. And because the file is public, it is the opposite of a privacy tool; it announces the paths you would rather not have visited.

Finally, the file only works if it loads and parses. Google's crawlers "treat all 4xx errors, except 429, as if a valid robots.txt file didn't exist," which means a broken file reads as "no rules, crawl everything." Size matters too: "Google enforces a robots.txt file size limit of 500 kibibytes (KiB). Content which is after the maximum file size is ignored" (per the spec). Keep it small, syntactically clean, and returning a normal response.

Why robots.txt matters for AI now

AI answer engines crawl the web the way search engines always have, and robots.txt is where you set the rules they follow. Block the wrong bot and you quietly disappear from AI answers; allow the wrong one and your content feeds training runs you meant to opt out of.

One mental model unlocks everything that follows: each AI company runs several bots, split by job. At OpenAI and Anthropic the work splits three ways: one bot gathers training data, one builds the search index used for citations, and one fetches a page when a user asks for it. Perplexity documents the last two and no training crawler at all. Blocking one does not touch the others. OpenAI states this outright: "Each setting is independent of the others" (OpenAI bot docs), and Anthropic runs three separately named crawlers on exactly this split.

Even if AI referral traffic is modest today, blocking a search crawler can remove your pages from that product's answers entirely, so the decisions you make now are worth making deliberately.

The AI crawlers, by platform

Here are the AI crawlers that matter most, grouped by the platform you actually think about, with what each one does and what blocking it costs you.

OpenAI (ChatGPT)

OpenAI documents four bots as of July 2026 (OpenAI's crawler list). Three of them matter for organic visibility:

  • GPTBot is the training crawler, "used to crawl content that may be used in training our generative AI foundation models." Disallowing it tells OpenAI that content crawled from your site should not be used in training. It does not affect ChatGPT search.
  • OAI-SearchBot is the search indexer. "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links." Blocking this one is what removes you from ChatGPT's cited answers.
  • ChatGPT-User fetches a page when a person asks ChatGPT about it. "Because these actions are initiated by a user, robots.txt rules may not apply."

The fourth, OAI-AdsBot, is an advertising bot, not an organic one: it "is used to validate the safety of web pages submitted as ads on ChatGPT," it "only visits pages submitted as ads," and the data it collects "is not used to train generative AI foundation models." OpenAI documents robots.txt controls for two bots specifically: it "uses OAI-SearchBot and GPTBot robots.txt tags to enable webmasters to manage how their sites and content work with AI." It does not say the same about OAI-AdsBot.

Anthropic (Claude)

Anthropic runs three independent crawlers, and this is the detail most guides get wrong. ClaudeBot collects content for model training. Claude-SearchBot indexes for Claude's search features; per Anthropic, disabling it "prevents our system from indexing your content for search optimization." Claude-User fetches pages when a Claude user asks about them (Anthropic's crawler docs).

Because the three are independent, a robots.txt group that disallows ClaudeBot alone leaves Claude-SearchBot and Claude-User crawling exactly as before. Each bot needs its own directive.

Google (Gemini)

Google-Extended is not a crawler at all. It is a robots.txt token with no bot of its own: Google's regular crawlers fetch your pages either way, and the token governs two things. Google documents it as controlling "whether content Google crawls from their sites may be used for training future generations of Gemini models that power Gemini Apps and Vertex AI API for Gemini and for grounding (providing content from the Google Search index to the model at prompt time to improve factuality and relevancy) in Gemini Apps and Grounding with Google Search on Vertex AI" (Google's crawler list). That second job matters: disallowing the token does not only opt you out of training, it also switches off Gemini-Apps grounding on your content. Google adds that it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."

The corollary trips people constantly: blocking Google-Extended does not remove you from AI Overviews or AI Mode. Those features draw on the regular Search index, which Googlebot builds. The only robots.txt lever that removes you from them is the one that removes you from Google Search itself.

Perplexity

Perplexity documents two bots (Perplexity's bot docs). PerplexityBot is the search indexer, "designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models." Perplexity-User handles on-demand fetches, and Perplexity is unusually direct about it: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules."

One dispute is worth knowing. In August 2025 Cloudflare reported that Perplexity was using undeclared, stealth crawlers to evade sites' no-crawl directives, and delisted its bots from Cloudflare's verified list (Cloudflare's report). Perplexity called the report a "publicity stunt" and said Cloudflare had conflated user-driven requests with automated crawling (TechCrunch's coverage). The disagreement turns on the very line this article draws: whether a fetch made on a user's behalf counts as crawling that robots.txt should govern.

Other AI crawlers

A few more names you will meet in server logs. CCBot belongs to Common Crawl (Common Crawl's CCBot page), whose public dataset is reused widely, including by AI companies for training. Applebot-Extended is Apple's robots.txt name for opting out of training. meta-externalagent crawls for Meta, and Bytespider for ByteDance. Rounding out the list: Amazonbot, cohere-ai, xAI's Grok bots, YouBot (You.com) and DuckAssistBot (DuckDuckGo).

Quick reference

BotOperatorJobHonors robots.txt?
GPTBotOpenAITrainingYes
OAI-SearchBotOpenAISearchYes
ChatGPT-UserOpenAIUser fetchMay not apply
OAI-AdsBotOpenAIAd landing-page checksNot documented
ClaudeBotAnthropicTrainingYes
Claude-SearchBotAnthropicSearchYes
Claude-UserAnthropicUser fetchYes
Google-ExtendedGoogleTraining and Gemini-grounding controln/a (a setting, not a crawler)
PerplexityBotPerplexitySearchYes
Perplexity-UserPerplexityUser fetchDisputed
CCBotCommon CrawlTraining dataYes

Allow, block, or something in between

The decision hinges on training versus visibility. Blocking training bots is usually low-cost: each vendor treats the rule as a signal that future crawls should not feed its models. It does not erase content already collected, and one token, Google-Extended, carries a visible trade-off in Gemini grounding. Blocking search and user bots is different; that is what removes you from AI answers, which is rarely what a site chasing visibility wants.

Three starting points, using Disallow only.

1. Allow everything. A minimal allow-all file can just declare your sitemap:

User-agent: *
Disallow:

Sitemap: https://example.com/sitemap.xml

2. Allow the search crawlers, limit training collection. Disallow the training bots and leave the search and user bots unlisted (unlisted means allowed):

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: CCBot
Disallow: /

Disallowing Google-Extended also turns off grounding in Gemini Apps and Vertex AI, with the trade-off set out in the Google section above. Drop that line if you want Gemini grounding.

3. Block the dedicated AI bots, keep classic search. Add the search and user bots to the list above; Googlebot and Bingbot have no groups here, so ordinary search is untouched:

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: Claude-User
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: Perplexity-User
Disallow: /

User-agent: CCBot
Disallow: /

This does not remove you from AI features built on the ordinary search indexes. AI Overviews and AI Mode run on the index Googlebot builds, and Apple's Siri and Search answers run on the index Applebot builds. No robots.txt rule removes you from those while keeping full search access.

Add any of the other crawlers from the previous section (meta-externalagent, Bytespider, Amazonbot and the rest) on the same pattern if you want them out too. One caveat: the training-versus-search line holds only as long as each vendor keeps those bots separate, and vendors change their bot lineups without much notice. Treat these templates as a starting point, then verify against the vendors' current documentation.

Common robots.txt mistakes

The section above was about what to choose. This one is about what not to break while writing it down.

The User-agent: * trap. A broad Disallow: / in the default group blocks every crawler that has no group of its own, which is most bots on this page. Matching happens per URL: the rule with the most specific path wins, and when rules conflict, the least restrictive one does.

Naming a bot removes it from your * rules. Google is explicit: "User agent specific groups and global groups (*) are not combined" (Google's grouping rules). The moment you add a GPTBot group, GPTBot stops reading your * group entirely, and you must repeat in its group anything you still want to apply to it.

Location and case. The file must sit at the root of the exact host, protocol and port being crawled; a robots.txt on www does not cover a bare domain or a subdomain. Path values are case-sensitive: Disallow: /Private/ does not block /private/.

Getting a dozen bot names, groups and paths exactly right by hand is tedious, and a single typo fails silently. Two subtler problems remain, though, and neither one is a typo: a directive that means different things to different engines, and a file that says one thing while your server does another.

Crawl-delay: the throttle that ages you out

Crawl-delay does not block a bot; it asks it to wait a number of seconds between requests. Set it high and you throttle your own re-crawl: in principle, the copy of your site inside an AI index refreshes more slowly, so recent changes take longer to show up in answers. It is a slow fade rather than a block, which makes it easy to miss.

It is also not a standard. The directive appears nowhere in RFC 9309, the Robots Exclusion Protocol specification, and engines disagree about it. Google ignores it entirely: "other fields such as crawl-delay aren't supported" (Google's field list), so it has no effect on Googlebot or any Google AI feature. Anthropic honors it: "we support the non-standard Crawl-delay extension to robots.txt" (Anthropic's support page). The same line in the same file means nothing to one engine and a real throttle to another. If you need it for server capacity, keep the value small and treat it as a hint that some bots will ignore.

"Allowed" isn't "reachable": when the edge blocks the bot

Your robots.txt can say "come in" while your CDN or firewall answers the same bot with a challenge page or a block. The file is a declaration; the edge is enforcement, and nothing keeps the two in agreement.

This disagreement is not hypothetical. On July 1, 2025, Cloudflare announced it was "changing the default to block AI crawlers unless they pay creators for their content" (Cloudflare's announcement). A managed firewall rule can turn away GPTBot or ClaudeBot while your robots.txt still reads "allowed," and nobody on the content team knows it is happening.

One limit is worth stating plainly: from the outside, you cannot prove whether a specific AI crawler can reach someone else's page. Real crawlers are verified by source IP address and reverse DNS hostname, not by the user-agent string (Google's verification docs), because anyone can fake a user-agent. Common Crawl warns about crawlers "falsely identifying themselves as CCBot" (CCBot documentation), and the Perplexity dispute above involved exactly this kind of rotating identity. So a test that sends a bot's user-agent from an ordinary server only tells you how your edge treated that one request. It is a signal worth investigating, never proof.

For certainty about your own site, use tools that fetch as the real bot or watch the real traffic: Google Search Console's URL Inspection, Bing Webmaster Tools, Cloudflare's AI Crawl Control, and your own server logs.

And one thing not to do: do not try to block AI bots by IP address instead of robots.txt. Anthropic warns that IP blocking "may not work correctly or persistently guarantee an opt-out, as doing so impedes our ability to read your robots.txt file" (Anthropic's opt-out guidance).

FAQ

Does blocking GPTBot hurt my Google rankings?

No. Google Search crawls with Googlebot, a different bot from a different company. A robots.txt group for GPTBot applies to GPTBot and nothing else.

What's the difference between GPTBot and ChatGPT-User?

GPTBot gathers content for model training on OpenAI's schedule. ChatGPT-User fetches one page, on demand, because a person asked ChatGPT about it. If it is citation in ChatGPT search you care about, the relevant bot is a third one, OAI-SearchBot.

Can I allow AI search but block AI training?

Yes. Disallow the training bots and leave the search and user bots unlisted; that is exactly what the second template above does.

Should I block Google-Extended?

Only if you accept both consequences. Google documents the token as covering training of future Gemini models and grounding in Gemini Apps, so blocking it also stops your content being supplied to Gemini at prompt time. It has no effect on your Google Search rankings, and it does not remove you from AI Overviews or AI Mode, since those are built from Google's regular index.

Do I need a separate rule for each of a vendor's bots?

Yes. OpenAI, Anthropic and Perplexity each run several independently named bots, and a rule for one does not apply to the others. To opt a whole company in or out, list each of its bots.

Does robots.txt actually stop AI bots?

It stops the crawlers that comply, and OpenAI, Anthropic, Google and Common Crawl all document that their automatic crawlers obey it; the user-triggered fetchers such as ChatGPT-User and Perplexity-User may not, as the table above shows. For anything that will not comply, enforcement lives at your CDN or firewall, with the "allowed isn't reachable" trade-off that comes with it.

Once your rules are written, confirm them against reality: check that every AI bot you meant to allow or block is actually listed the way you intended, and that the file itself loads, parses and stays within the limits that make it count.

See if AI can read, trust, and cite your site

Add to Chrome

Free · No signup · Every issue links back to a guide like this one