Canonical URLs and duplicate content

Published

Most pages end up living at more than one URL without anyone deciding they should. A canonical tag is how you tell search engines which of those addresses is the real one. For Google it is a hint rather than an instruction, and outside Google none of the AI crawler docs mention the tag at all.

What a canonical URL is

A canonical URL is the one you nominate as the master copy of a page that exists at several addresses. The duplicates accumulate on their own: the same content is reachable over http and https, at www and non-www, at a URL with tracking parameters bolted on, as a print version, or as the same product filed under two category paths. To a person these are all one page. To a crawler they are several distinct pages, and the ranking signals they earn get split across them.

In Google's own definition, "a canonical URL is the URL of a page that Google chose as the most representative from a set of duplicate pages" (Google's canonicalization overview). You declare your preference with a rel="canonical" link in the page's <head>:

<link rel="canonical" href="https://example.com/product">

The job is consolidation. Instead of five variants each collecting a share of the links and signals, the duplicates all point at one URL and the signals gather there. A self-referencing canonical, a page pointing at its own URL, is normal and fine; it simply states that this address is the master.

Do AI crawlers respect canonical tags?

The honest answer splits in two. Google's AI features respect your canonical by inheritance. The engines behind ChatGPT, Claude and Perplexity do not document the tag at all.

Google's AI: yes, through the index. Canonicalization is a Google Search mechanism: Search picks one canonical from each set of duplicates and indexes that one. Google's AI features, such as AI Overviews and AI Mode, then draw on that same index. Google states that "our generative AI features on Google Search are rooted in our core Search ranking and quality systems," and that a page "must be indexed and eligible to be shown in Google Search with a snippet" to qualify (Google's AI features guide). So the URL Google selects as canonical is the version eligible to be surfaced or cited in Google's AI answers. Get the canonical wrong and Google may surface the wrong duplicate, or may not surface the content at all.

Non-Google AI crawlers: not a documented lever. OpenAI, Anthropic, and Perplexity all publish crawler documentation for bots including GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, and PerplexityBot, and it covers robots.txt user agents and published IP ranges, with no mention of rel=canonical. Perplexity's page is representative of the genre: "Webmasters can use the following robots.txt tags to manage how their sites and content interact with Perplexity" (Perplexity's crawler docs). OpenAI's crawler docs and Anthropic's crawler docs follow the same pattern: user agent strings, robots.txt directives, IP ranges, nothing about canonical handling. That was still true when all three pages were checked in July 2026. None of those vendor pages says its crawlers use rel=canonical to consolidate duplicate URLs. That does not mean they ignore it. It means the behavior is undocumented, so the tag is not a reliable way to steer which of your URLs they use.

The lever those engines do document is robots.txt, which governs their automated crawlers: GPTBot and OAI-SearchBot at OpenAI, ClaudeBot and Claude-SearchBot at Anthropic, PerplexityBot at Perplexity. It does not cover everything they fetch: OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them, because both are triggered by a person asking a question. Beyond that, the linking on your own site shapes which URL they find first: not a documented rule, just how discovery works. Practically, if two near-identical URLs are both crawlable, an engine like Perplexity or ChatGPT may fetch either one. The signals you can actually control are your own: keeping links, sitemap, and redirects pointed at one URL is a defensive habit, not a published rule about which duplicate gets fetched.

The clean way to hold all this: canonical is a search-indexing tool. It reaches Google's AI through the index and does its real work there.

How Google picks the canonical

Even inside Google, your canonical is a preference, not an instruction. Google is explicit about this: "You can indicate your preference to Google using these techniques, but Google may choose a different page as canonical than you do, for various reasons. That is, indicating a canonical preference is a hint, not a rule" (Google's canonicalization docs).

Google weighs several signals, and they are not equal. In its guidance on specifying canonicals, a redirect is "A strong signal" that the target should become canonical, a rel="canonical" annotation is likewise "A strong signal," and sitemap inclusion is only "A weak signal" (Google on specifying canonicals). None of them are mandatory: per the same page, "if you don't specify a canonical URL, Google will identify which version of the URL is objectively the best version to show to users in Search."

To see which URL Google actually chose for a page, use URL Inspection in Search Console. It reports which URL Google selected as canonical, and that is not always the one you declared. Search Console is also where a message that scares a lot of people appears: "Alternate page with proper canonical tag." This is not an error. Google describes AMP, mobile, and desktop alternates this way: "This page correctly points to the canonical page, which is indexed, so there is nothing you need to do" (Page indexing report). The duplicate is doing its job, and the canonical it points at is indexed.

The duplicate content penalty myth

The most durable myth in this corner of SEO is that duplicate content earns a penalty. It does not, and Google has said so for years. A Google post that has been live since 2008 says, "There's no such thing as a 'duplicate content penalty,'" and it adds that duplicate content is "not grounds for action on that site unless it appears that the intent of the duplicate content is to be deceptive and manipulate search engine results" (Google's 2008 blog post). Current guidance says the same thing: "Some duplicate content on a site is normal and it's not a violation of Google's spam policies" (Google's canonicalization overview).

Don't swing to the other extreme either. The same 2008 post notes that duplicates "can potentially affect your site's performance." Signals split across the versions, the crawler spends effort on near-identical URLs, and Google may index a version you did not want people to land on. That is a reason to canonicalize. It is not a punishment to fear.

Canonical vs noindex vs redirect

Three tools get confused with each other because they all deal with "extra" URLs, and they do very different things.

A canonical keeps both URLs live and names one as the master, so signals consolidate there. It is the right tool for genuine duplicates you want to keep reachable: the printable version, the parameterized URL, the second category path.

noindex removes a page from search entirely, and it is the wrong tool for choosing between duplicates. Google again: "We don't recommend using noindex to prevent selection of a canonical page within a single site, because it will completely block the page from Search. rel="canonical" link annotations are the preferred solution" (Google's consolidation guidance). Whether a page can be crawled and whether it can be indexed are separate questions with their own machinery; the point here is that noindex does not pick a winner among duplicates, it deletes a contestant.

A 301 redirect removes the old URL from circulation and sends everyone, users and crawlers alike, to the new one. Reach for it when the duplicate should not exist as its own page, for example after a URL change. A redirect is likewise a strong canonicalization signal in Google's list, but it is not interchangeable with rel="canonical": a redirect collapses two URLs into one, a canonical keeps both alive and names a master.

The takeaway travels well beyond Google. Pick one URL per page, point every duplicate at it with the canonical, redirect the URLs that should not exist, and keep your links and sitemap agreeing with that choice. Google treats the hint as a strong signal, and the engines that document nothing about canonicals (OpenAI, Anthropic, Perplexity) will at least keep meeting the same consistent answer everywhere they look.

FAQ

Do AI engines like ChatGPT use my canonical tag?

Not that they document. The crawler docs from OpenAI, Anthropic, and Perplexity cover robots.txt user agents and IP ranges, not canonical handling; Google's AI is the exception because it runs on a Search index that already honors your canonical.

What is a canonical tag?

A rel="canonical" link in a page's <head> naming one URL as the master among a set of duplicates, so engines consolidate their signals on that one URL.

Does duplicate content hurt my SEO?

There is no penalty; Google consolidates duplicates and indexes one version. Duplicates can still waste crawl effort and split signals across URLs, which is the practical reason to canonicalize.

What does "Alternate page with proper canonical tag" mean in Search Console?

It is a normal state, not an error: the page points to a canonical that Google has indexed, so there is nothing to fix.

Does a canonical guarantee that URL gets indexed?

No. The tag is a hint, Google can select a different canonical than the one you declared, and indexing itself is never guaranteed.

See if AI can read, trust, and cite your site

Add to Chrome

Free · No signup · Every issue links back to a guide like this one