Indexability: crawling vs indexing
A crawler reaching your page is only half the job. Whether the page then gets indexed is a separate decision. For Google, indexing is the gate: a page must be in the Search index to appear in AI Overviews or AI Mode. Other answer engines reach content through their own indexes, through search partners, or by fetching a page live when a user asks, so their eligibility rules have to be checked one at a time. A page can be perfectly reachable, serve without a single error, and still never make it into an index.
Crawling vs indexing
Crawling and indexing are two separate stages, and search engines treat them that way. Crawling is a bot fetching your page: it requests the URL, receives the response, and reads what came back. Access at that stage is governed mainly by robots.txt, which the companion article on AI crawler access covers in full. It is not absolute: both OpenAI and Perplexity document user-triggered fetchers that may bypass robots.txt because a person asked for the page. Indexing happens afterward: the engine decides whether to keep the page and make it eligible to appear in results. A successful fetch commits the engine to nothing. Google's documentation puts it plainly: "For Google Search, an HTTP 2xx (success) status code doesn't guarantee indexing" (Google on HTTP status codes). Reachable does not mean indexed.
This distinction is what makes indexability an AI visibility question, not just a search one. Google is explicit that its AI answers sit on top of the ordinary index: "our generative AI features on Google Search are rooted in our core Search ranking and quality systems. These features rely on AI techniques to highlight content from our Search index," and to appear in them "a page must be indexed and eligible to be shown in Google Search with a snippet" (Google's AI features guide). The same guide adds a caution worth keeping in mind for everything below: "Indexing and serving aren't guaranteed." So if a page is not in the index, it is not in AI Overviews or AI Mode either. There is no separate door.
When a reachable page fails to index, this article covers two kinds of cause: an explicit instruction that you, or your CMS, put there on purpose, and technical conditions that disqualify the page silently. Canonical selection, duplication and Google's own quality judgement can also keep a crawlable page out of the index.
The noindex directive: telling engines to stay out on purpose
For Google Search, noindex is the deliberate way to keep a page out of results, and it is decisive. In Google's words: "When Googlebot crawls that page and extracts the tag or header, Google will drop that page entirely from Google Search results, regardless of whether other sites link to it" (Google's noindex guide). Links pointing at the page do not save it. That is the point: it exists for pages you never want competing in results.
It comes in two forms with the same effect. In an HTML page, put the robots meta tag in the <head>:
<meta name="robots" content="noindex"> Or return it as the X-Robots-Tag HTTP response header, sent by the server alongside the page:
X-Robots-Tag: noindex Google treats the two as equivalent: "Any rule that can be used in a robots meta tag can also be specified as an X-Robots-Tag" (Google's robots meta tag docs). The header earns its keep on non-HTML files: a PDF or an image has no <head> to hold a meta tag, and per the noindex guide it "can be used for non-HTML resources, such as PDFs, video files, and image files."
One confusion is so common it needs settling here: noindex is not nofollow. noindex keeps the page itself out of results. nofollow tells the crawler "Do not follow the links on this page ..." (same robots meta tag docs). They are independent rules: a page can be indexed while its links are nofollowed, or dropped from results while its links are still followed. If your goal is keeping a page out of results, nofollow does nothing for you.
Typical pages worth a deliberate noindex: staging environments, thank-you and internal-search-results pages, and thin pages you do not want in results at all. Duplicates are a different case: consolidate those with a canonical, because noindex removes a page rather than choosing between versions of it.
Do AI crawlers actually respect noindex?
For Google's AI features, yes, by inheritance: noindex drops a page from the Search index, and AI Overviews and AI Mode draw only on that index. There is also a narrower, related lever: the nosnippet rule, which per the robots meta tag documentation "will also prevent the content from being used as a direct input for AI Overviews and AI Mode." That keeps an indexed page's content out of Google's AI answers while leaving the page itself in Search, at the cost of its ordinary snippet too.
For the non-Google engines the picture is uneven, and each has to be read separately (checked July 2026; these are living pages). OpenAI documents noindex for one narrow purpose: if it learns of a disallowed page from a third-party search provider it may still surface the link and title, and it tells publishers "If you do not want this to happen, use the noindex meta tag", adding that "in order for our crawler to read a meta tag, it must be allowed to crawl the relevant page(s)" (OpenAI's publisher FAQ). Anthropic goes further, listing the tag among its removal levers: "The noindex robots meta tag is a rule that tells our partners not to index your content so that they don't send it to us in response to your web search query", and covered content "will not appear in Claude outputs that use web search" (Anthropic's content removal guide). Note that this works through the search partners Claude queries, not through ClaudeBot. Perplexity is the exception: its bot guide names only robots.txt, saying "Webmasters can use the following robots.txt tags to manage how their sites and content interact with Perplexity" (Perplexity's bot docs), and says nothing about noindex. That silence is not evidence it ignores the tag, but robots.txt is the only lever Perplexity actually documents.
The confusion also runs the other way, and this one is a verified myth: putting noindex rules inside robots.txt. "Specifying the noindex rule in the robots.txt file is not supported by Google." Google itself honored it unofficially for years, then announced the end of that support in July 2019, in a post titled "A note on unsupported rules in robots.txt". Worse, blocking a page in robots.txt is actively self-defeating if your goal was deindexing. The noindex guide warns: "For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file ... the crawler will never see the noindex rule, and the page can still appear in search results ..." A blocked page cannot deliver its own removal instruction. So never combine the two: to deindex a page, let it stay crawlable and serve noindex.
When a page deindexes itself by accident
Not every indexability problem is deliberate. The page carries no noindex, nobody decided to keep it out, and yet it never appears. These are the mechanisms that disqualify a page silently. Together, these failures can drop a page from the index, make it look empty, or send Google to a different URL than the one you meant. And because an AI engine has to fetch and read a page before it can cite it, the same failure can starve it at the source: an OAI-SearchBot, Claude-SearchBot or PerplexityBot that requests a 404, a soft-404 shell, or the wrong content type receives nothing usable to index or quote. The mechanics below are Google's documented behavior; the AI consequence follows from the same fetch, not from a separate AI rule.
The wrong HTTP status
A normal page you want indexed should answer with 200. Google's status code documentation is just as clear about errors as it is about success: "Google doesn't index URLs that return a 4xx status code, and URLs that are already indexed and return a 4xx status code are removed from the index." Server errors are more forgiving but only for a while: with 5xx responses, "already indexed URLs are preserved in the index, but eventually dropped." The trap is that a page can render perfectly in your browser and still serve the wrong status to a crawler: an error handler that fires only for certain user agents, a misconfigured rewrite, an overloaded origin that returns 500 to bots at crawl time. What you see on screen is not evidence of what the crawler received.
Soft 404s
The inverse failure is a page that looks like an error but claims success. "A soft 404 error is when a URL that returns a page telling the user that the page does not exist and also a 200 (success) status code" (Google on soft 404s). The consequence is blunt: "Such pages are excluded from Search." This is especially common in JavaScript applications, where the server happily returns 200 for any path and the client-side code then renders an empty state or a "not found" message. The user sees an error page; the crawler sees a successful response wrapped around error content, concludes the URL is worthless, and drops it. Real missing pages should return a real 404.
Meta refresh redirects
A <meta http-equiv="refresh"> tag that forwards visitors to another URL is a weak substitute for a real redirect. Google does try to interpret it: per its redirect documentation, "Google Search interprets instant meta refresh redirects as permanent redirects" and delayed ones as temporary, but "a server side redirect has the highest chance of being interpreted correctly by Google" (Google's redirect docs). A meta refresh only works if the crawler fetches the page, parses the HTML, and draws the intended conclusion; a server-side 301 states the intent in the response itself, before any parsing. If a page has permanently moved, redirect it with a 301 at the server and let the meta refresh trick stay in the past.
The wrong content type
Google decides how to handle a file by what the server says it is. "The file type is determined by the Content-Type HTTP header returned when Google crawls the file, though in some cases Google may use the file extension or re-parse the file ... if the Content-Type header is missing or incorrect" (Google's file type docs). The practical consequence follows from that mechanism: if an HTML page is served with a wrong header such as Content-Type: text/plain, the crawler has been told it is looking at plain text, not a web page, and the page may not be parsed and indexed as one. The fallback re-parsing exists, but it is a recovery path, not something to rely on. Serving text/html for HTML is a one-line server fix that removes the ambiguity entirely.
How to check a page is actually indexable
Start with what every crawler, AI crawlers included, actually receives, because these are facts about your server's response, independent of any one search engine. Two things cover most of this article. Read the raw response first: look in the page source for a robots meta tag carrying noindex, then read the response headers: the real status code (is it actually 200?), any X-Robots-Tag, and the Content-Type. Your browser's network panel shows them, and so do these commands. Use a normal GET rather than curl -I, because -I sends a HEAD request and a server can answer it differently from the GET a crawler makes:
curl -sD - -o /dev/null https://example.com/page
curl -sD - -o /dev/null -A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" https://example.com/page The second command repeats the check as a named crawler, because a server can vary its response by user agent. If JavaScript injects the robots meta tag at runtime, check the rendered DOM separately and treat it as a weaker signal: Vercel's study of AI crawler traffic found that "none of the major AI crawlers currently render JavaScript" (Vercel's AI crawler study), so a tag that exists only after rendering is one most of them never see.
If those response-level checks are clean, the page is indexable as far as your server is concerned. For Google specifically, you can go further: URL Inspection in Search Console fetches the page as Google and reports whether it is indexed and, if not, why. Bing Webmaster Tools offers its own URL Inspection that does the same for Bing's index. Each is one tool for one engine, not the whole answer. A site: search is only a rough hint, not proof, in either direction. None of the vendor crawler documents cited above describes a comparable URL-level inspection tool, which is why the response-level checks come first.
FAQ
Does noindex keep my page out of ChatGPT and other AI answers?
It depends on the engine. OpenAI tells publishers to use noindex to stop a disallowed page surfacing as a link and title, and Anthropic lists the tag as a way to keep content out of Claude answers that use web search, though it works through Claude's search partners rather than through its own crawler. Perplexity documents only robots.txt. So noindex is a documented lever for two of the three, not a universal one (checked July 2026).
Does noindex remove me from Google's AI Overviews?
Yes, indirectly. AI Overviews and AI Mode run on the Search index, and noindex drops the page from that index, so there is nothing left for them to surface.
Should I use robots.txt or noindex to keep a page out of search?
They do different jobs: robots.txt stops crawling, noindex stops indexing. To deindex a page, use noindex and leave the page crawlable so the crawler can actually read the rule.
What's the difference between noindex and nofollow?
noindex keeps the page itself out of search results; nofollow tells crawlers not to follow the links on it. They are independent, and neither implies the other.
Why does Search Console say "crawled, currently not indexed"?
Google fetched the page but chose not to index it. Reachability is not the problem, so look for an overlooked noindex, thin or soft-404 content, or a status or redirect issue.
See if AI can read, trust, and cite your site
Add to ChromeFree · No signup · Every issue links back to a guide like this one