HTTPS for AI crawlers

Published

HTTPS gets discussed as a ranking signal, which misses what it actually does for your visibility in AI answers. Every AI crawler is an HTTP client: a page it cannot fetch because the certificate is broken never reaches the model by that route, no matter how good the content is. And a valid certificate on its own is not the finish line.

What HTTPS actually is

HTTPS is ordinary HTTP carried over an encrypted TLS connection. Everything that travels between the client and your server is encrypted in transit, so it cannot be read or tampered with along the way. The padlock in the address bar signals it, and its absence now works against you: major browsers label plain HTTP pages "Not Secure".

Getting HTTPS is a solved problem: a TLS certificate (free options exist) plus a server configured to serve the site over https://. The interesting part is not obtaining the certificate. It is what a crawler does when the certificate, or the page behind it, is broken.

Do AI crawlers need HTTPS?

None of the major AI crawlers' docs list HTTPS as a requirement. But every AI crawler has to fetch your page before its content can be read, indexed, or cited, and a broken certificate can make that fetch fail. That distinction, requirement versus reachability, is the whole answer.

Take the crawlers by name: GPTBot and OAI-SearchBot from OpenAI, ClaudeBot and Claude-SearchBot from Anthropic, and PerplexityBot from Perplexity. Their published documentation covers robots.txt rules and IP ranges and says nothing about HTTPS or TLS (OpenAI's crawler docs are typical). So there is no documented lever here: serving HTTPS earns you nothing these vendors have promised. What you can do is lose everything with a broken certificate. Each of these crawlers is an HTTP client, and to fetch an https:// page it first has to complete a TLS handshake. An expired certificate, a self-signed one, or one issued for a different hostname makes that handshake fail for any client that validates the certificate, which standard TLS clients do by default, and a page a crawler cannot fetch is a page it cannot read, index, or cite. No vendor documents that failure mode because none needs to: it is how the transport works.

Google documents its own eligibility conditions for AI answers such as AI Overviews and AI Mode. Google states that its "generative AI features on Google Search are rooted in our core Search ranking and quality systems", that "a page must be indexed and eligible to be shown in Google Search with a snippet", and that you should "ensure your content is crawlable" (Google's AI features guide). That guide never names HTTPS. The concrete lever sits one layer down, in how Google chooses which version of a page to index. Google "prefers HTTPS pages over equivalent HTTP pages as canonical, except when there are issues or conflicting signals such as the following:", and one of those issues is that "The HTTPS page has an invalid SSL certificate" (Google's canonicalization guidance). The same document is blunt about the failure case: "Avoid bad TLS/SSL certificates and HTTPS-to-HTTP redirects because they cause Google to prefer HTTP very strongly. Implementing HSTS cannot override this strong preference." A broken certificate does not merely fail to help. It flips Google to your HTTP version, and that version is what then competes, or fails to compete, for indexing and for eligibility in AI answers.

And the ranking effect? Real, but small enough that Google itself, announcing it in 2014, called HTTPS "only a very lightweight signal", one "affecting fewer than 1% of global queries, and carrying less weight than other signals such as high-quality content" (Google's 2014 announcement). It is a tiebreaker, not a strategy. If ranking were the only reason to care about HTTPS, it would barely be worth the effort.

Mixed content: the HTTPS mistake that breaks your page

A valid certificate is not the end of the job. Mixed content is what happens when a page served over HTTPS pulls in some of its resources over plain HTTP: MDN defines it as "securely loaded web pages that use resources to be fetched via HTTP or another insecure protocol" (MDN on mixed content).

The consequence appears in the client. Per the same MDN reference, browsers "mitigate the risks of mixed content by auto-upgrading image, video, and audio mixed content requests from HTTP to HTTPS, and block insecure requests for all other resource types". Blocked means blocked: an HTTP <script>, stylesheet, or <iframe> on an HTTPS page never loads, so the script does not run, the styles do not apply, and the page renders broken or partial. That is what a human visitor sees. Google renders with an evergreen version of Chromium, so a crawler on a browser engine would in principle hit the same blocked requests, though Google does not document its renderer's mixed-content behavior. That browser-side failure does not affect crawlers that only download raw HTML. The crawlers behind ChatGPT, Claude and Perplexity were measured, in Vercel and MERJ's December 2024 study, not to execute page resources (Vercel's crawler study), so GPTBot, ClaudeBot and PerplexityBot receive the raw HTML untouched. So mixed content costs you with human visitors and with rendering crawlers; a crawler that never executes the page never notices. A green padlock can sit above a half-broken page.

Google also counts it against you at the canonicalization step. Its canonicalization guidance flags an HTTPS page that "contains insecure dependencies (other than images)". Mixed content therefore costs you twice: the page renders broken, and Google leans back toward the HTTP version.

Prevention is unglamorous: serve every resource over HTTPS. For assets on your own domain, Google's 2014 tip still holds: "Use relative URLs for resources that reside on the same secure domain". For everything hosted elsewhere, write explicit https:// URLs, and skip protocol-relative URLs that start with //; that pattern is dated advice. In practice, fixing mixed content is usually a find-and-replace on http:// references in templates and content, plus a check of third-party embeds.

Delivery mistakes that make crawlers miss your content

The failures that keep content away from Google and AI crawlers alike are delivery failures: a certificate that does not validate, redirects that point the wrong way, pages that render broken. Moving from HTTP to HTTPS cleanly means closing every one of those gaps, so that every crawler lands on the same working HTTPS URL.

  1. Serve a valid certificate that matches the host. A certificate error stops the fetch for any client that validates the certificate. The requirement is explicit: "The certificate must match your complete site URL, or be a wildcard certificate that can be used for multiple subdomains on a domain" (Google's certificate guidance). A common break is serving one host's certificate on another host, for example the bare domain's certificate on a subdomain it does not cover.
  2. 301-redirect every HTTP URL to its HTTPS equivalent, one to one, with no chains. Google notes that "a server side redirect has the highest chance of being interpreted correctly by Google", and a 301 is a "strong signal" that the move is permanent where a 302 is a "weak" one (Google's redirects documentation). Never redirect the other way: an HTTPS-to-HTTP redirect is one of the cases that triggers the strong HTTP preference quoted earlier.
    # nginx: send every HTTP request to its exact HTTPS equivalent
    server {
        listen 80;
        server_name example.com;
        return 301 https://example.com$request_uri;
    }
    
  3. Update canonical tags and internal links to the https:// versions. Your own signals should agree with each other, so that you are not relying on the redirect to correct every link and canonical you publish.
  4. Keep the HTTPS site crawlable and indexable. A classic way migrations fail: the new HTTPS site goes live behind a robots.txt rule that blocks it, or with a leftover noindex tag. Google's guidance is direct: "Don't block your HTTPS site from crawling using robots.txt", and "Allow indexing of your pages by search engines where possible. Avoid the noindex robots meta tag." How robots.txt directives and indexing controls work is its own topic; for the migration, the rule is simply that neither may fence off the new site.
  5. Fix mixed content as part of the move, not as a follow-up ticket, and verify the HTTPS property in Search Console so you can watch the new URLs get picked up.

Done this way, there is one clean HTTPS site, and it is the same site for every client: Googlebot, GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, and PerplexityBot alike.

FAQ

Do AI crawlers require HTTPS?

OpenAI, Anthropic and Perplexity document no HTTPS requirement for their crawlers, but OAI-SearchBot, Claude-SearchBot, and PerplexityBot all have to complete a successful fetch before your content can appear in an answer. A broken certificate fails that fetch, so the content never reaches the model.

Does a bad SSL certificate hurt my SEO or AI visibility?

Yes, concretely. For any crawler a certificate error can block the fetch outright, and for Google an invalid certificate or an HTTPS-to-HTTP redirect makes it prefer your HTTP version very strongly.

What is mixed content?

An HTTPS page that loads some of its resources, such as scripts or stylesheets, over plain HTTP. Browsers block those insecure requests, which can break the page, and Google holds insecure dependencies against your HTTPS page when choosing a canonical.

Is HTTPS a ranking signal?

Yes, but a very lightweight one: when Google announced it in 2014 it said the signal affected fewer than 1% of global queries and carried less weight than content quality. Treat it as a tiebreaker between otherwise equal pages, not a growth lever.

How do I move from HTTP to HTTPS cleanly?

Install a certificate that matches your hostname, 301-redirect every HTTP URL to its exact HTTPS twin, update canonical tags and internal links, fix mixed content, and leave the HTTPS site open to crawling and indexing. Then confirm the HTTPS version in Search Console.

See if AI can read, trust, and cite your site

Add to Chrome

Free · No signup · Every issue links back to a guide like this one