Crawl budget and delivery speed

Published

Before ChatGPT, Perplexity, Claude, or Google's AI Overviews can use anything you publish, a crawler has to fetch the page. Those crawlers now arrive in volume, and they read whatever HTML your server hands them. They also waste more of what they fetch than Googlebot does: Vercel and MERJ measured roughly a third of ChatGPT's and Claude's fetches landing on 404 pages, against 8.22% for Googlebot. How much of your site they carry away comes down to two things you control: how fast and cheap each page is to fetch, and whether the visit gets spent on pages that matter or burned on duplicates and dead ends. Getting both right decides whether your pages get fetched at all, and how fresh the fetched copies stay.

Most AI crawlers read the raw HTML your server returns and never run your JavaScript; what that means for how you build pages is covered in our guide to how AI crawlers render pages. This article picks up at the moment a crawler connects: what limits how much it fetches, and how to keep it from wasting the visit.

What crawl budget is

Google opens its crawl budget guide with a caveat, and it belongs at the top here too: "If your site doesn't have a large number of pages that change rapidly, or if your pages seem to be crawled the same day that they are published, you don't need to read this guide" (Google's crawl budget guide). The guidance is aimed at large sites of a million or more unique pages, and at medium sites of ten thousand pages or more whose content changes daily; Google adds that "These are not exact thresholds." A third signal comes from Search Console: when large numbers of your URLs sit in "Discovered, currently not indexed", crawling is often the bottleneck. If none of that describes your site, the practices below are still good hygiene, but they are not an emergency. A fifty-page site that gets picked up the day it publishes has no crawl budget problem to solve.

The model itself is compact: "A site's crawl budget is determined by two main elements: crawl capacity limit and crawl demand." Capacity is how much crawling your server can absorb without degrading. Demand is how much of your site a crawler considers worth fetching, which in Google's description reflects its perceived inventory of your URLs, how popular they are, and how stale its last copy has become. Both halves have to be there, as Google's summary makes explicit: "Google defines a site's crawl budget as the set of URLs that Google can and wants to crawl. Even if the crawl capacity limit isn't reached, if crawl demand is low, Google will crawl your site less."

That document describes Google's own crawling. OpenAI, Anthropic, and Perplexity publish no equivalent model for how GPTBot, ClaudeBot, or PerplexityBot pace themselves. But the mechanics on your side of the connection do not care whose bot is asking: a fast, healthy server is cheaper to crawl for every visitor, a bloated page costs every visitor the same extra bytes, and a duplicate URL wastes any crawler's time.

Capacity: make your pages cheap to fetch

Capacity follows server health, and Google describes the feedback loop plainly: "If the site responds consistently and its response times (including latency and Time-to-First Byte) remain stable or improve, the limit goes up, meaning more connections can be used to crawl. If the site slows down (latency increases or response times become longer), or responds with server errors (5xx HTTP status codes) or rate-limiting signals (such as HTTP 429), the limit goes down and Google crawls less" (Google Search Central). Notice what is missing: a number. Google publishes only the direction of the effect: faster, healthier responses earn more crawling, and slower or error-prone responses earn less. No vendor publishes the thresholds behind it. The precise figures that circulate in SEO writing, a timeout of so many seconds or a throttle at so many milliseconds, are invented. The practical instruction in the same guide is a single line: "Improve loading speed: Optimize your server response times and resources to make pages load faster." Elsewhere Google puts the payoff just as simply: "If Google can load and render your pages faster, we might be able to read more content from your site."

For a crawler that never renders, the page is the HTML document, byte for byte. That makes size and compression direct levers. HTML text compresses well; gzip, Brotli, or zstd typically make a document three to five times smaller than its uncompressed size, so serving uncompressed HTML spends the crawl on byte transfer instead of pages. Markup bloat has a second cost beyond transfer time: when most of the document is wrapper code, the content a crawler came for sits diluted inside it. There is a hard cap on the fetch, too. Googlebot "crawls the first 2MB of a supported file type" and drops the rest, a limit "applied on the uncompressed data" (Google's Googlebot documentation); other crawlers set their own. A lean document is fetched faster, parsed cheaper, and read further.

The third capacity lever is the re-crawl. Crawlers do not fetch a page once; they come back to check whether it changed, and you can make those return visits nearly free: "Use HTTP caching: Support 304 (Not Modified) HTTP status codes. If a page hasn't changed since Google last crawled it, returning a 304 code tells Google to reuse the cached version, saving your server bandwidth and resources" (Google's caching guidance). The mechanism is a conditional request. Your server sends an ETag or Last-Modified header with the page, the crawler presents it on the next visit, and if nothing changed the whole exchange is a status line instead of a download:

GET /pricing HTTP/1.1
If-None-Match: "v42"

HTTP/1.1 304 Not Modified

None of this is a Core Web Vitals exercise. LCP, CLS, and INP measure what a user experiences in a browser, so they matter for exactly that reason. They also matter for rendering engines like Google that load your page the way a browser does. A related nuance: render-blocking scripts and stylesheets delay a rendering crawler, because it has to fetch them before it can finish the page, but a crawler that never renders is not held up by them. Vercel and MERJ found the non-rendering bots still request JavaScript files (11.50% of ChatGPT's requests, 23.84% of Claude's); they simply never execute them, so there is no render pass to block. What does reach the non-rendering crawlers is delivery itself: time to first byte, response size, compression, and cache headers. No vendor documents Core Web Vitals as an AI citation signal, and a page can pair a fast raw response with a poor Vitals score or the reverse. Fix delivery for the crawlers; fix Vitals for your users.

Demand: don't make crawlers waste the visit

The second half of the budget is what a crawler chooses to spend it on, and Google is unusually direct about where your leverage sits: "Perceived inventory: Without guidance from you, Google tries to crawl all or most of the URLs that it knows about on your site. If many of these URLs are duplicates, or you don't want them crawled for some other reason ... this wastes a lot of Google crawling time on your site. This is the factor that you can positively control the most" (Google's crawl budget doc).

The waste has familiar shapes. Faceted navigation and URL parameters mint near-duplicate pages by the thousand: the same product list re-sorted, re-filtered, and re-paginated under different query strings. Calendars, session IDs, and open-ended filter combinations go further and create infinite URL spaces a crawler can wander for as long as you let it. Consolidation fixes the duplicates: choose a canonical URL for each piece of content and point the variants at it. For sections you never want crawled, robots.txt is the right tool, because it stops the request before it is spent:

User-agent: *
Disallow: /search
Disallow: /*?sort=

Cleanup covers the rest. Removed pages should return 404 or 410 rather than a soft 404, a "not found" page served with a 200 status, which invites crawlers to keep revisiting nothing. Long redirect chains spend a request per hop before any content arrives, so flatten them. And a sitemap kept current, with an accurate <lastmod> on each URL, points crawlers at the fresh content instead of leaving them to rediscover it by brute force.

One common move backfires: marking unwanted pages noindex to save crawl. "Don't use noindex, as Google will still request, but then drop the page when it sees a noindex meta tag or header in the HTTP response, wasting crawling time." The tag controls indexing, not crawling. The page is fetched, then discarded, and the budget is spent either way. If the goal is to keep scheduled crawling out of a section entirely, block it in robots.txt. That governs conforming crawlers, not user-triggered fetchers: Perplexity documents that Perplexity-User generally ignores robots.txt.

AI crawlers, shared capacity, and freshness

Everything above would be worth doing for Google alone. What has changed is who else is in the queue. A December 2024 analysis by Vercel and MERJ measured the major AI crawlers arriving at serious scale. It found that they fetch the raw HTML and do not render JavaScript. The crawlers it covered were GPTBot and OAI-SearchBot from OpenAI, ClaudeBot from Anthropic, and PerplexityBot from Perplexity. Anything absent from the HTML your server returns is not recovered by a rendering pass. Slow, bloated, or uncompressed delivery hits them with nothing to soften it.

Crowding is part of the picture too. Google notes that "the crawl capacity limit is shared across all crawlers. This means that high demand from one crawler can reduce the capacity available for others" (Google's documentation). The same squeeze plays out at the level of your server, whoever the bots answer to: a server running near its limits while absorbing heavy AI bot traffic has less headroom left for the crawls that keep your important pages fresh, and if it starts answering slowly or throwing 5xx errors, Google documents that its own crawling drops off. What OpenAI, Anthropic, and Perplexity do in that situation is not documented. The capacity work above is also how you stay comfortable under the extra load. Your CDN is where that traffic becomes visible: Cloudflare's AI Crawl Control, for one, reports requests broken down by individual crawler, status code and path (Cloudflare's AI traffic docs).

Freshness is the quieter half of the payoff. Some sites respond to bot volume by setting a high Crawl-delay in robots.txt. Support is uneven and worth checking before you rely on it: Anthropic documents that it respects the non-standard Crawl-delay extension, while Google states its crawlers do not process the field at all, and OpenAI and Perplexity document no handling of it. Where a crawler does honor it, spacing out its requests also stretches how long a correction, a price change, or a rewritten page takes to be re-fetched into the index that answer draws from. If the goal is lighter load rather than slower updates, cheap re-crawls get you there without the staleness: a correct 304 for unchanged pages and honest <lastmod> dates let crawlers stay current at a fraction of the cost.

What none of this buys is a citation. No vendor and no study documents a speed or crawl-budget lift in how often AI engines cite a page. The stakes are earlier and plainer: whether your pages are in the index an answer engine consults at all, and how stale the copy is when it gets used. Do this work for coverage, freshness, server cost, and the humans on your site.

FAQ

Does my small site need to worry about crawl budget?

Usually not. Google says that if your pages are crawled the same day they are published, you do not need to worry about it. Crawl budget matters mainly for large or rapidly updated sites, roughly 10,000 pages and up, or for sites with many URLs stuck in "Discovered, currently not indexed".

Does page speed help me get cited by AI?

Not as a citation lever. Speed decides whether crawlers can fetch your pages cheaply and how many they take: Google documents that a consistently healthy server earns more crawling and a slow or error-prone one earns less, and no AI vendor publishes an equivalent rule. The payoff is coverage and freshness, not a documented citation boost.

How do AI crawlers change crawl budget?

Most AI crawlers arrive in volume and read your raw HTML without rendering it, so slow or bloated delivery hurts them directly. Gemini, which runs on Googlebot's infrastructure, AppleBot, and Bing are the measured exceptions to that no-rendering pattern. Google says its crawl-capacity limit is shared across Google's own crawlers. Separately, all bot traffic competes for the same server resources, so heavy AI crawling on a slow server leaves less headroom for everything else.

What wastes crawl budget the most?

Duplicate and parameter URLs, which Google calls the factor you can control the most. Soft 404s, long redirect chains, and infinite URL spaces add to the waste, as does putting noindex on pages you should have blocked in robots.txt, since Google still crawls them to see the tag. Consolidate or block them.

Do speed scores matter to AI crawlers?

There is no documented evidence that they are. Core Web Vitals matter for users and for rendering engines like Google. For the non-rendering AI crawlers the levers are delivery speed, byte size, compression, and cache headers, not an LCP or CLS score.

See if AI can read, trust, and cite your site

Add to Chrome

Free · No signup · Every issue links back to a guide like this one