XML sitemaps for AI crawlers
An XML sitemap is the file where you hand crawlers a list of the pages you want them to find. It is one of the most misunderstood tools in technical SEO: half of what people put in one is ignored outright, and it does not do what many site owners believe it does. It will not get pages indexed and it will not move rankings. Whether it helps in AI answers depends on the engine: Microsoft documents sitemaps as a signal for Bing Copilot, while OpenAI, Anthropic, and Perplexity do not document whether their crawlers read one at all.
What an XML sitemap is
An XML sitemap is a machine-readable file that lists the URLs on your site you want search engines to know about. Google's definition is plain: "A sitemap is a file where you provide information about the pages, videos, and other files on your site ... Search engines like Google read this file to crawl your site more efficiently" (Google's sitemaps guide). A minimal, valid sitemap is short. This is Google's own example, and it is a complete file:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/foo.html</loc>
<lastmod>2022-06-04</lastmod>
</url>
</urlset> One point of naming confusion is worth clearing immediately: this is not the same thing as an HTML sitemap or a visual sitemap, which are human-facing pages or planning diagrams. The XML file exists for crawlers, and only for crawlers.
Be clear about what it buys you. A sitemap is a discovery aid, nothing more. It tells crawlers which URLs exist so they do not have to find every page by following links. It does not rank anything, and it does not force anything into the index. Google says so directly: "A sitemap helps search engines discover URLs on your site, but it doesn't guarantee that all the items in your sitemap will be crawled and indexed."
Which raises the honest question most guides skip: do you even need one? Not always. The same guide notes that "If your site's pages are properly linked, Google can usually discover most of your site," and that you can likely skip a sitemap when "Your site is 'small'. By small, we mean about 500 pages or fewer ...". A sitemap earns its keep on large sites, new sites with few inbound links, sites with pages that internal navigation does not reach, and media-heavy sites. On a small, well-linked site, it is optional. It is usually near-free when your CMS generates it, though Google notes an XML sitemap "can be cumbersome to work with" and "complex to maintain the mapping on larger sites, or sites where the URLs change often." It is a convenience for crawlers, not a requirement.
What actually belongs in it
A sitemap entry has one required tag and three optional ones. Only one of the optional tags still matters.
<loc> is the URL, and the discipline around it is what matters most. List only URLs you actually want indexed: the canonical version of each page, returning 200, not blocked by robots.txt, not carrying a noindex, not a redirect, not an error page. The sitemap is you telling crawlers "these are my pages." If that list includes URLs your other signals say to skip, you are sending mixed messages, and every junk URL you list is crawl attention you asked for and wasted, attention that could have gone to pages that matter. This is the same crawl-versus-index discipline that governs indexability generally: the sitemap should agree with everything else your site says.
<lastmod> is the one optional tag Google actually uses, and only if it tells the truth. Per Google's build-sitemap doc, "Google uses the <lastmod> value if it's consistently and verifiably (for example by comparing to the last modification of the page) accurate." The same doc defines what counts: "The <lastmod> value should reflect the date and time of the last significant update to the page ... an update to the copyright date is not." Changes to the main content, structured data, or links are significant; bumping a footer year is not. The incentive here is real but conditional: an accurate <lastmod> is a genuine freshness signal that helps crawlers prioritize recrawling, while a <lastmod> that updates on every deploy regardless of content teaches Google to ignore yours entirely. Gaming this tag simply does not work.
Then there are the two dead tags. Google's position is a single sentence: "Google ignores <priority> and <changefreq> values." The protocol itself never promised much more: the sitemaps.org spec says priority "is not likely to influence the position of your URLs in a search engine's result pages," and calls <changefreq> a hint rather than a command. Sitemap generators commonly emit both tags anyway. Don't bother tuning them. They add bytes, they invite busywork, and Google states outright that it ignores them. For a plain URL sitemap, <loc> is required and an honest <lastmod> is the only optional core tag Google says it uses. Image, video, and news sitemap extensions are separate and separately documented; add them only if your site contains the corresponding content.
Does a sitemap help AI find and cite your pages?
Indirectly for Google's AI features, directly according to Microsoft, and not in any documented way for the rest. That distinction matters, so take the engines one at a time.
Google's AI: yes, through the index. AI Overviews and AI Mode do not run on a separate crawl of the web. Google states that "our generative AI features on Google Search are rooted in our core Search ranking and quality systems. These features rely on AI techniques to highlight content from our Search index," and that a page "must be indexed and eligible to be shown in Google Search with a snippet" (Google's AI features guide). Note what that guide does not say: it never mentions sitemaps. The connection is a chain you can trace, not a sentence Google wrote. A sitemap feeds Googlebot's discovery, discovery feeds the Search index, and the Search index is what Google's AI features draw from. Every link in that chain is documented; the sitemap simply sits at the start of it.
Non-Google AI crawlers: not a documented lever. The companies behind ChatGPT, Claude, and Perplexity each publish documentation on how their crawlers work and how site owners can control them. Read those pages and the same pattern emerges for OpenAI's GPTBot and OAI-SearchBot, Anthropic's ClaudeBot and Claude-SearchBot, and PerplexityBot alike: it is robots.txt and published IP ranges, every time. OpenAI writes that "OpenAI uses OAI-SearchBot and GPTBot robots.txt tags to enable webmasters to manage how their sites and content work with AI" (OpenAI's bot docs). Anthropic writes that "Anthropic's Bots respect "do not crawl" signals by honoring industry standard directives in robots.txt" (Anthropic's crawler docs). Perplexity writes that "we recommend allowing PerplexityBot in your site's robots.txt file and permitting requests from our published IP ranges ..." (Perplexity's bot docs).
None of those pages mentions reading an XML sitemap. That is a silence, not a denial: it would be wrong to claim these crawlers read your sitemap, and equally wrong to claim they ignore it. What the silence does tell you is where not to place your bet. For these engines, the documented lever is being crawlable in robots.txt; being reachable through ordinary links is the likely discovery path, though none of these vendors documents how its crawler finds URLs. A sitemap is not among the levers they document. If your goal is visibility in ChatGPT, Claude, or Perplexity, the work is access and internal linking, not sitemap tuning.
Microsoft: the one vendor that says it outright. Bing published a post in July 2025 titled "Keeping Content Discoverable with Sitemaps in AI Powered Search", which states that "sitemaps remain a foundational signal for ensuring comprehensive URL coverage across your site" and recommends pairing them with IndexNow for faster notification of changes (Bing Webmaster Blog). Since Copilot is grounded in the Bing index, this is the clearest first-party statement tying sitemaps to an AI-powered search surface.
So the honest summary: an XML sitemap is a classic search-engine discovery aid whose AI payoff runs through the search indexes behind Google's AI features and Bing's Copilot, and which OpenAI, Anthropic and Perplexity simply do not discuss. It is worth having for exactly that reason. It is not a switch that turns on AI visibility.
Making sure crawlers can find your sitemap
A sitemap nobody can find does nothing. Google documents two ways to make yours known, and they are not equivalent.
The first is a Sitemap: line in your robots.txt file. Google's instruction: "Insert the following line anywhere in your robots.txt file, specifying the path to your sitemap."
Sitemap: https://example.com/my_sitemap.xml This is the passive, universal method. Anything that reads your robots.txt sees the reference, which makes it the one discovery hook a non-Google crawler could plausibly encounter, since robots.txt is the file the AI crawlers above do document reading. It costs one line and is visible to every crawler that reads your robots.txt.
The second is submitting the sitemap in Google Search Console, through the Sitemaps report or the Search Console API. This one is Google-specific, but it is the only method that talks back: Search Console reports back on the sitemap's status, whether it could be fetched and whether it parsed. That feedback loop is the practical reason to submit even if your robots.txt already carries the reference. (Google once also offered a ping endpoint you could hit to announce sitemap updates; that method is gone from Google's documentation, so don't build anything on it.)
Keep expectations calibrated even after submitting. Per Google's submission docs, "submitting a sitemap is merely a hint: it doesn't guarantee that Google will download the sitemap or use the sitemap for crawling URLs on the site." Do both: the robots.txt line for universal discoverability, the Search Console submission for feedback.
Keeping it valid and trustworthy
A sitemap fails in two ways: as XML, or as a promise. Both are avoidable.
As XML first. A malformed file can be rejected wholesale, and every URL in it loses the discovery benefit at once. The requirements are the basics of well-formed XML: a proper <urlset> root with the http://www.sitemaps.org/schemas/sitemap/0.9 namespace, correctly nested tags, and URLs with special characters properly encoded. If your sitemap is generated by a framework or plugin, this is usually handled; if it is built by custom code, validate the output rather than assuming.
Size has hard documented ceilings. Per Google's build-sitemap doc: "All formats limit a single sitemap to 50MB (uncompressed) or 50,000 URLs ... You can optionally create a sitemap index file ...", which lets you split a large site across multiple sitemap files and submit one index that references them all. The sitemaps.org protocol sets the same bounds, down to the byte: no more than 50,000 URLs and no larger than 50MB, which it specifies as 52,428,800 bytes. Cross either line and the file is out of spec.
The harder failure is the promise. Every URL in your sitemap is a claim that the page is worth a crawler's time. A sitemap padded with noindexed pages, non-canonical variants, redirects, and 404s makes that claim false at scale, and a fabricated <lastmod> does the same for freshness. Keep the list to canonical URLs you actually want indexed and the file keeps doing what it is for. A small sitemap that is entirely accurate beats a large one that is partly fiction.
Keep it valid, keep it current, list only what you genuinely want indexed, and it quietly does its one job: making sure that when a crawler wants to find your pages, nothing is left to luck.
FAQ
Do I need an XML sitemap?
Not always. Google itself says a site of about 500 pages or fewer with proper internal linking can usually be discovered without one; the file pays off most on large, new, or poorly linked sites.
Does a sitemap help my pages rank or get indexed?
No. It helps crawlers discover URLs, which is a precondition for indexing, but Google is explicit that inclusion in a sitemap guarantees neither crawling nor indexing, and it plays no role in ranking.
Do AI engines like ChatGPT use my sitemap?
Not that they document. OpenAI, Anthropic, and Perplexity document robots.txt controls and IP verification rather than how their crawlers find URLs, and none mentions sitemaps. Microsoft does: Bing recommends sitemaps for discovery in AI powered search, and Copilot draws on the Bing index. Google's AI features draw on the Search index, which a sitemap can help feed.
Should I set priority and changefreq?
No. Google ignores both tags outright, and the protocol only ever offered <priority> as a possible tiebreak between URLs on your own site, never as a ranking lever. An accurate <lastmod> is the one optional tag worth maintaining.
Where do I put my sitemap so it's found?
Reference it with a Sitemap: line in your robots.txt, which any crawler can see, and submit it in Google Search Console, which reports back on whether it was fetched and parsed.
See if AI can read, trust, and cite your site
Add to ChromeFree · No signup · Every issue links back to a guide like this one