HTML for AI crawlers: encoding, doctype and structure
Before an AI engine can quote a sentence from your page, the document carrying that sentence has to survive up to three steps. First it is decoded: the bytes arriving over the network are turned into text. Then it is parsed: the text is turned into a structure a machine can walk. For the consumers that go that far, it is then rendered: the parsed structure is laid out visually. Seven document-level conditions can decide whether those steps succeed: the character encoding declaration, freedom from mojibake, the doctype, the head and body order, the declared page language, the base URL, and the viewport tag.
This is the document shell, not the content. Whether a bot runs your JavaScript at all, how you write meta tags and headings, structured data, and semantic HTML are all separate topics. The promise here is narrow and mechanical: get these seven right and you remove the common document-shell obstacles to decoding, parsing, and rendering; get them wrong and even perfect content can arrive garbled, misread, or truncated.
Three things read your document, and they fail differently
Three layers of processing stand between your server and an AI answer, and the things that read your document stop at different layers. Decoding turns bytes into text, and every crawler has to do it before anything else. Parsing turns text into structure. Rendering lays the page out visually, and almost nothing in the AI pipeline does it.
Most of the crawlers behind AI answers read raw HTML exactly as served. OAI-SearchBot (which feeds ChatGPT search), PerplexityBot and CCBot do not run a browser over your page. In December 2024, Vercel and MERJ measured live AI crawler traffic and reported that none of the major AI crawlers they observed render JavaScript (Vercel crawler study). Anthropic's Claude-SearchBot was not in that test, and Anthropic documents no rendering behavior for it; what it does document, "The web fetch tool currently does not support websites dynamically rendered with JavaScript" (Anthropic web fetch docs), covers a developer tool in its API, not a crawler. A strict, non-browser parser is also less forgiving of broken structure than a browser, which repairs almost anything.
The crawler that matters most here does render: Googlebot Smartphone, which fetches the mobile page, lays it out, and feeds the index behind Google's AI Overviews and AI Mode.
That gives the seven fundamentals a map, built from the standards rather than from vendor statements. Charset and mojibake act at decoding, and reach every consumer. Head and body order, the base URL, and the declared page language act during parsing, which is where the raw-HTML bots live. Viewport and the layout half of the doctype act during rendering, which among the consumers here means Googlebot and the other browser-based crawlers such as AppleBot.
Character encoding: declare UTF-8
The character encoding declaration tells a parser how to turn your bytes into characters. The preferred form is a <meta charset="utf-8"> tag in the head; when the server also sends the declaration in an HTTP header (Content-Type: text/html; charset=utf-8), the header takes precedence over the tag.
The choice of encoding is not open. The WHATWG Encoding Standard is blunt: "Authors must use the UTF-8 encoding and must use its (ASCII case-insensitive) "utf-8" label to identify it" (WHATWG Encoding Standard).
Placement matters because of how parsers find the declaration: a parser prescans roughly the first 1024 bytes of the document looking for it. Put <meta charset="utf-8"> first inside <head>, before the title or anything else that could push it past that window. A minimal correct opening looks like this:
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>Page title</title>
</head> When no encoding is declared anywhere, every consumer has to guess, and guesses differ: some default to UTF-8, others to the legacy windows-1252. Plain ASCII survives either guess, so an all-English page can look fine everywhere. A page with accented letters, Cyrillic or CJK text, emoji, or typographic quotes can decode differently for different crawlers. This is the decode layer, so it hits every crawler, rendering or not, before parsing even begins.
Mojibake: garbled text
Mojibake is what a wrong decode looks like on the page: multi-byte UTF-8 sequences read as single-byte characters. An accented e becomes é, a right single quote becomes ’, and bytes that cannot be decoded at all become the replacement character � (U+FFFD).
The obvious cause is a missing or wrong charset declaration. The subtle case is worse: a page can declare utf-8 correctly and still ship mojibake, because the text was double-encoded upstream, for example by a CMS re-encoding content that was already UTF-8. The declaration looks right; the bytes themselves are already wrong.
For AI consumption this is a plain readability failure, not a ranking penalty. Garbled text cannot be tokenized sensibly, quoted, or cited, and a raw-HTML crawler reads exactly the broken bytes you served. There is no rendering step to paper over them.
The fix lives at the source: repair the encoding in the database or publishing pipeline and serve consistent UTF-8 end to end. When you verify, inspect the raw served bytes rather than the page in a browser, because the browser view can mask the damage.
The doctype and quirks mode
<!doctype html> is the entire modern doctype. MDN is explicit about what it does: "The only purpose of <!doctype html> is to activate no-quirks mode." Without it, a layout engine falls back to legacy behavior kept for old pages: "There are now three modes used by the layout engines in web browsers: quirks mode, limited-quirks mode, and no-quirks mode" (MDN quirks mode guide).
Note the word layout. Quirks mode is a rendering behavior: it changes how CSS sizes and positions boxes. For a crawler that renders and lays pages out, such as Googlebot, a missing or legacy doctype can drop the page into quirks mode and change the layout it indexes. For a raw-HTML crawler that only extracts text, the doctype barely changes what it reads. A missing doctype does not make ChatGPT unable to read your page.
Ship it anyway: it is only fifteen characters, protects the rendering path, and costs nothing. Put it before everything except an optional byte order mark or comment, and use the short form; the long legacy doctypes from the HTML4 era buy you nothing.
Head and body order
The document shell is <html> containing <head> and then <body>. What counts as broken is narrower than most checklists suggest: HTML5 permits omitting the <html>, <head>, and <body> open tags entirely, and a parser will synthesize the missing elements. A page without an explicit <body> tag is not, by itself, faulty.
The real fault is broken order, or content in the wrong container: a <head> that opens after <body> has started, or body content sitting where head metadata belongs. Browsers practice aggressive error recovery. They hoist stray content into <body>, merge a late <head>, and present a tidy document tree. Strict, non-browser parsers used in parts of the crawling and training stack are less forgiving. On a badly ordered document they can miss metadata that appears late, for example a <meta name="robots"> that turns up after <body> has opened, or stop processing early. This is a mechanism and a risk, not a documented policy of any named vendor; no engine publishes its parser's recovery rules.
The practical consequence: raw-HTML crawlers read your served bytes, but your browser shows you the repaired result, so the rendered view hides exactly this class of problem. Check the raw served HTML, not the DOM inspector.
Declare the page language
This is the one condition of the seven that does not decide whether your document can be read. It decides what a machine does with the text once it has it. The lang attribute on the root element, <html lang="en">, declares the primary language of the document: one value per page, such as en, uk, or de, with a region subtag when it matters (en-GB) (MDN lang attribute docs). Screen readers use it to load the correct pronunciation rules (W3C WCAG guidance), browsers use it for hyphenation and typography, and a parser gets the language stated instead of having to guess it from the text.
A wrong value is worse than a missing one. Declare lang="en" over Ukrainian text and every consumer that trusts the attribute is actively misled; the value must match the visible content. It is not a search or AI-citation lever: "Google uses the visible content of your page to determine its language. We don't use any code-level language information such as lang attributes, or the URL" (Google multilingual site docs). And it is not hreflang, which points to alternative language versions of a page rather than declaring the language of this one.
The base URL
Most pages never need a <base> element, which is part of why it does outsized damage when it is wrong. MDN: "The <base> HTML element specifies the base URL to use for all relative URLs in a document. There can be only one <base> element in a document." And: "If multiple <base> elements are used, only the first href and first target are obeyed, all others are ignored" (MDN base element docs). It must sit in <head> before any element that uses a URL, and data: and javascript: values are not allowed in its href.
One wrong <base> silently rewrites every relative URL in the document at once: navigation links, a relative canonical, image sources. The classic bug is a base that carries a path with no trailing slash, because the last segment is then replaced rather than kept:
<base href="https://example.com/docs">
<a href="page">A link</a>
<!-- resolves to https://example.com/page, not /docs/page -->
<base href="https://example.com/docs/">
<a href="page">A link</a>
<!-- resolves to https://example.com/docs/page --> A raw-HTML crawler resolves your relative links and canonical against <base> exactly as written. If the value is wrong, the crawler follows broken URLs and can attribute the page to a canonical you never intended. The rule is simple: most sites should leave <base> out, and any site that uses it must get the origin and trailing slash exactly right. One boundary worth knowing: Open Graph tags ignore <base> and need absolute URLs regardless.
Viewport and mobile-first indexing
<meta name="viewport" content="width=device-width, initial-scale=1"> tells a layout engine to size the page to the device instead of pretending to be a desktop window. It is the standard mark of a mobile-friendly page.
It belongs squarely to the render layer. In the December 2024 Vercel and MERJ measurements, OAI-SearchBot and PerplexityBot did not render, so a viewport tag changes nothing they see. No vendor publishes this as a standing guarantee. The consumer it clearly matters for is Google: "Google uses the mobile version of a site's content, crawled with the smartphone agent, for indexing and ranking. This is called mobile-first indexing" (Google mobile-first docs). That index is the one AI Overviews and AI Mode draw from.
Keep the payoff proportionate. Google states that its AI features set no extra bar: "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary" (Google AI features docs). The viewport tag is not an AI trick. It is one line of ordinary mobile hygiene that keeps the mobile-first indexing path healthy, and that path happens to feed Google's AI answers.
FAQ
Do AI crawlers care about my doctype?
Mostly one of them does. Googlebot renders and lays out the page, so a missing doctype can drop it into quirks mode and change the layout Google indexes. Raw-HTML crawlers such as PerplexityBot barely notice. Ship <!doctype html> anyway: it is free, and it is the correct default.
Do I need a charset declaration if my page is already UTF-8?
Yes. The declaration is what stops consumers falling back to locale-dependent guesses. Without one, crawlers guess, and some guess the legacy windows-1252 rather than UTF-8, which garbles any non-ASCII character. Declare <meta charset="utf-8"> as the first element in <head>, or send the equivalent HTTP header.
Will a missing <html> or <body> tag break AI crawling?
Not by itself. HTML5 permits omitting the <html>, <head>, and <body> open tags, and parsers synthesize them. The real risk for strict raw-HTML parsers is broken order: head metadata that appears after body content has started can be missed. Check the served source, not the browser's repaired view.
Does the viewport tag matter for ChatGPT?
No. In the December 2024 Vercel and MERJ measurements, ChatGPT's search crawler, OAI-SearchBot, did not render the page, so a viewport tag changes nothing it sees. Viewport matters for Googlebot Smartphone and Google's mobile-first index, which is what feeds AI Overviews and AI Mode.
Is <base href> worth adding?
Usually not. Most sites resolve relative URLs correctly without it. If you do use it, a single mistake, typically a missing trailing slash, silently rewrites every relative URL in the document, including a relative canonical. Get the origin and path exactly right, or leave the element out.
See if AI can read, trust, and cite your site
Add to ChromeFree · No signup · Every issue links back to a guide like this one