AI content controls: training opt-outs vs AI answers

Published

There is no switch that walls AI systems out of your content. What a site owner actually has is a set of declarations: signals placed in robots.txt, meta tags, HTTP headers, and well-known files that state how you want your pages used. Those declarations carry very different weight, and two questions decide which one you need. First, what are you trying to stop: AI systems training on your pages, or AI products quoting them in answers? Second, does the lever you are about to set actually do anything?

One boundary before anything else. Stopping a crawler from reaching your site at all is an access question, handled in robots.txt, and it is its own topic. This guide covers the other half: controlling how a page is used once a crawler can already fetch it. The honest headline is that only Google and Microsoft document a lever that changes what their AI products do with a page they can already fetch, and both levers cost you something.

What you can and cannot control

Every usage control serves one of three goals: telling AI systems not to train on your pages, telling them not to quote you in AI answers, or making sure your own signals agree with each other. Deciding which one you want is the whole of it, because the way you opt out of training is not the way you opt out of answers. Sorting a lever into the right goal matters, because a training opt-out does nothing about AI answers, and the reverse.

You also need to judge how much weight each declaration carries. From strongest to weakest: documented first-party answer controls (Google's nosnippet, Bing's noarchive); vendor training tokens in robots.txt (GPTBot, Google-Extended, ClaudeBot); advisory first-party syntax (Cloudflare's Content Signals); a legal rather than technical reservation (TDMRep); and de-facto, nonstandard tags (noai, noimageai). None of this is technical enforcement. As the IPTC puts it, "robots.txt is only a recommendation to site crawlers, and it does not guarantee that it will be followed by AI providers in any jurisdiction" (IPTC opt-out guidelines). A real standard is being drafted at the IETF's AI Preferences Working Group, but today's controls are a patchwork. The rest of this guide walks them strongest first.

Stay indexed, stay out of AI answers

Google and Microsoft are the only two companies that document a post-fetch control, one you set on a page you still let them crawl, that they say keeps your text out of an AI answer. OpenAI, Anthropic and Perplexity document no equivalent; their documented levers are crawler-level. Both piggyback on snippet directives that predate generative AI.

On Google, the lever is nosnippet. Google's documentation states: "This applies to all forms of search results (at Google: web search, Google Images, Discover, AI Overviews, AI Mode) and will also prevent the content from being used as a direct input for AI Overviews and AI Mode." (Google robots meta tag doc) The softer max-snippet limits how much text Google may show or use rather than suppressing it completely: "This applies to all forms of search results (such as Google web search, Google Images, Discover, Assistant, AI Overviews, AI Mode) and will also limit how much of the content may be used as a direct input for AI Overviews and AI Mode." If you only want to shield part of a page, the data-nosnippet attribute scopes the same exclusion to a single HTML element instead of the whole document. The page-level version is one line:

<meta name="robots" content="nosnippet">

On Bing, the page-level equivalents are noarchive and nocache. Microsoft's announcement is explicit: "Content tagged NOARCHIVE will not be included in Bing Chat answers, not be linked to in the answers. Going forward, for content in our Bing Index that is labeled NOARCHIVE, we will not use the content for training Microsoft's generative AI foundation models." (Bing content controls) The gentler option keeps you visible in a reduced form: "Content with the NOCACHE tag may be included in Bing Chat answers. We will only display URL/Snippet/Title in the answer." Precedence is also documented: "If content has both NOCACHE and NOARCHIVE tags, we will treat it as NOCACHE." Since October 2025 Bing also supports data-nosnippet for excluding a single section: "Any content marked with data-nosnippet is still indexed normally, but it will be excluded from snippets and AI summaries." (Bing data-nosnippet announcement)

Now the cost, stated plainly. These directives govern preview text in general, not AI features specifically. Setting nosnippet removes your page from AI Overviews and AI Mode, and it also deletes the ordinary description under your blue link in regular search results. There is no AI-only opt-out: the price of leaving the AI answer box is a bare, snippetless listing everywhere else. That trade is the whole decision, and it is worth making deliberately.

Tell AI not to train on you

The tag most blog posts reach for is the de-facto one: <meta name="robots" content="noai, noimageai">. Its provenance is documented by the IPTC: "The additional "noai" and "noimageai" directives, defined by DeviantArt in 2022 and used by other sites including fab.com". That is the full extent of its standing. It is nonstandard, and the major AI companies do not document honoring the noai meta tag. Lists that show OpenAI, Google, and Anthropic all "honoring" it point to no vendor documentation, because none of those companies mentions the tag in its crawler docs. Publish it if you want your intent on the record, but treat it as a statement, not a control.

The training opt-out the vendors themselves document lives in robots.txt, as a per-vendor user-agent token. OpenAI's is GPTBot: "Disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models." (OpenAI bots doc) Google's training and grounding token is Google-Extended, and Anthropic's training crawler is ClaudeBot: "When a site restricts ClaudeBot access, it signals that the site's future materials should be excluded from our AI model training datasets." (Anthropic's crawler docs) Each is disallowed the same way as GPTBot. On the Microsoft side, the noarchive tag quoted above does double duty: it removes you from Copilot answers and from Microsoft's generative AI training in one directive. The mechanics of writing those robots.txt rules are an access topic of their own. Perplexity is the notable absence from this list: it is a retrieval-only answer engine that, by its own documentation, does not build foundation models, so there is no Perplexity training opt-out to set.

One myth needs correcting here, because it inverts the two goals. Blocking Google-Extended does not remove you from AI Overviews or AI Mode. Those features draw on the regular Search index, built by Googlebot, and Google's own documentation frames the token narrowly: "To limit AI training and grounding in some of Google's other systems, read more about Google-Extended." (Google AI features doc) If your goal is to leave AI Overviews, the lever is nosnippet in the previous section, not a training token.

Cloudflare Content Signals

Content Signals is Cloudflare's attempt to make robots.txt say something about use, not just access. It adds a Content-Signal: line carrying up to three yes/no preferences: search, where Cloudflare specifies that "Search does not include providing AI-generated search summaries."; ai-input, covering content fed into AI answers; and ai-train, covering "training or fine-tuning AI models." (Cloudflare Content Signals). A typical line looks like this:

Content-Signal: search=yes, ai-train=no

Cloudflare's own framing is a request to crawlers, not a gate: the signals are advisory, and no engine is technically bound by them. What makes this lever worth knowing about is the default. Cloudflare ships the line as part of that feature: if you turned on Cloudflare's managed robots.txt, it generates search=yes, ai-train=no for you, with ai-input left unset. Cloudflare states how widely that feature is enabled: "Cloudflare customers have already turned on our managed robots.txt feature for over 3.8 million domains". The wording is Cloudflare's, and Cloudflare treats enabling the feature as your choice to disallow AI training, so read your own robots.txt and confirm the inherited preference is still what you mean. Note also what the default does not say: search=yes excludes AI-generated summaries by definition, so it is not a grant of AI-answer use.

Reserve your rights with TDMRep

The TDM Reservation Protocol (TDMRep) is a different kind of lever: its force is legal, not technical. It lets a site declare a text and data mining reservation in one of three carriers: an HTTP tdm-reservation header, a JSON file (the spec states: "This specification defines a JSON file named tdmrep.json, which MUST be hosted in the /.well-known repository of a Web server."), or an HTML meta tag. The value 1 means rights are reserved; 0 means they are not.

tdm-reservation: 1

Its standing needs stating precisely. TDMRep is a W3C Community Group Final Report, and the document says of itself: "It is not a W3C Standard nor is it on the W3C Standards Track." (W3C TDMRep report). Its purpose is to give machine-readable form to a reservation under Article 4 of the EU's Copyright in the Digital Single Market Directive (2019/790), which allows text-and-data-mining rights to be reserved when that is done "in an appropriate manner, such as machine-readable means" (Directive 2019/790). The Directive names no specific protocol, and TDMRep is one proposal for satisfying it. No crawler is technically required to read tdmrep.json, and adoption is unmeasured. What the file provides is a dated, machine-readable record that the reservation was declared. That is a description of what the Directive and the spec say, not legal advice.

Make your declarations agree

The levers above live in five different places: robots.txt lines, meta tags, HTTP headers, a well-known file, and for some sites an llms.txt. Because they are set at different times by different people, a site can easily end up contradicting itself. The classic case is publishing an llms.txt that invites AI systems in while robots.txt blocks the very crawler that would read the invitation. That one is self-defeating by construction: a crawler you blocked cannot fetch the welcome you published for it. The quieter case is a Content Signals line declaring ai-input=no while the marketing pages actively court AI visibility. Those two point opposite ways about one use: whether your content may be fed into AI answers.

The problem with a contradiction is not that engines resolve it in some exotic way. It is simpler: you have told the web two different things, so at least one of them is not what you meant, and any engine that reads your signals may act on either. The fix is an audit, not a trick: list every AI-related declaration your site currently publishes, across all five carriers, and check that they describe one coherent intent.

One rendering caveat belongs in that audit. A meta tag injected only by client-side JavaScript may never reach the crawlers it is meant to instruct, since many fetch raw HTML and do not render scripts. Usage controls belong in the served HTML or in HTTP headers, where every fetcher sees them.

Frequently asked questions

Can I actually stop AI from training on my content?

You can declare it: robots.txt vendor tokens like GPTBot and ClaudeBot, Bing's noarchive, or the de-facto noai tag. The major labs document these tokens as instructions about future training material, not as technical enforcement, and none of them describes a robots.txt directive that reaches content already collected.

Does nosnippet keep me in Google Search?

You stay indexed and ranked, but you lose your snippet. Because nosnippet and max-snippet govern preview text generally, opting out of AI Overviews this way also removes the ordinary description under your result.

Does blocking Google-Extended remove me from AI Overviews?

No. AI Overviews and AI Mode draw on the regular Search index through Googlebot. Google-Extended governs training and grounding in some of Google's other systems, not those answer surfaces. If leaving AI Overviews is the goal, the documented lever is nosnippet.

Does the noai meta tag actually work?

It is a nonstandard directive, defined by DeviantArt in 2022, and the major AI companies do not document honoring the meta tag. Their documented opt-outs run through robots.txt user-agent tokens instead. Treat noai as a public statement of intent, not a guarantee of anything.

Is a Content Signal in my robots.txt something I chose?

Possibly not in those words. Cloudflare's managed robots.txt generates search=yes, ai-train=no for the domains that turned the feature on, and Cloudflare treats enabling it as your choice to disallow AI training. Read your robots.txt and check that the inherited preferences match what you actually want.

See if AI can read, trust, and cite your site

Add to Chrome

Free · No signup · Every issue links back to a guide like this one