Semantic HTML: landmarks and accessible names
Two kinds of machine now read your page. Answer engines such as ChatGPT, Perplexity, Bing Copilot and Google extract your content and cite it in their answers. AI agents go a step further: they operate the page, clicking buttons and filling forms on a user's behalf. Where either one reads your HTML or the accessibility tree built from it, what it gets depends on the structure of your markup, not on what the page looks like. This guide covers the two structural surfaces that carry that weight: the landmark elements that mark out your main content, and the accessible names that tell a machine what each control does. Title and meta tags, structured data and heading structure matter too, but they are separate topics with their own guides.
Two ways AI reads your structure
Semantic HTML serves two different AI audiences, and they use it differently. Answer engines extract; agents operate. Keeping the two apart makes it clear what each element of your markup is actually doing for you.
Extraction: answer engines separating content from boilerplate
Raw HTML is mostly not content. A typical page wraps a few hundred words of substance in navigation menus, ads, cookie banners, related-article widgets and a footer. Before an answer engine can chunk a page and cite it, it has to separate the main content from that boilerplate. Main-content extraction routinely keys on landmarks: open-source extractors such as trafilatura select <main> and <article> directly and drop <header> and navigation regions. <main> says "this is the content," and document-level <header>, <footer> and <nav> say "this is not." A page built from nested <div>s gives no structural signal at all, and leaves the extractor to infer where the article starts from visual heuristics. This is a mechanism, not a measured ranking lift: content extraction is a long-standing, documented step in how machines process pages, but no answer-engine vendor documents whether landmarks influence its own extraction, and no study has isolated landmarks as a citation factor.
Operation: agents reading the accessibility tree
The second audience does not just read your page. It clicks it. One route from your markup to an operating agent is the accessibility tree, the semantic model the browser builds from your HTML. In the words of the MDN glossary: "Browsers then create an accessibility tree based on the DOM tree, which is used by platform-specific Accessibility APIs to provide a representation that can be understood by assistive technologies, such as screen readers." Every node in that tree carries a role (button, link, textbox), a name (what it is called) and a state (checked, expanded, disabled). Screen readers have navigated by this model for decades. AI agents driving a browser can work three ways: from a screenshot of the rendered page, from the raw HTML, or from that accessibility tree.
Which method a given agent uses is contested
Different agents use different inputs, and their public documentation varies in detail. OpenAI's publisher FAQ describes its ChatGPT Atlas browser as using ARIA tags, the same labels and roles that support screen readers, to interpret page structure and interactive elements. Yet OpenAI's Operator, its computer-using agent, works primarily from screenshots. Microsoft's Playwright MCP is reported to work purely from the accessibility tree, with no vision at all. Anthropic's Claude computer use, by its own documentation, takes the opposite approach: screenshots plus pixel coordinates. Coverage of Google's Gemini-based agent describes a hybrid, with no clean first-party statement. So the input varies by product, and the through-line is that native semantic HTML directly serves two of those three inputs: it gives clean markup to the HTML readers and builds a correct tree for the tree readers. Screenshot-based agents work from the rendered interface instead, and swapping a <div> for a <main> changes nothing they can see, so they are served by a clear visual layout rather than by your element names. Sound semantic markup plus a clear visual layout is the robust bet across all three approaches. The ceiling on the evidence is worth stating plainly: there is one clean vendor statement (OpenAI, about Atlas).
Landmarks: mark your main content
The landmark set is small. <main> wraps the dominant content of the page, the thing the URL exists for. Document-level <header> and <footer> mark the site chrome. <nav> marks navigation blocks. <article> and <section> divide the content itself into self-contained and thematic pieces. Of these, the <main> element is the headline signal, because it draws the one boundary an extractor cares most about: content on the inside, boilerplate on the outside.
Most pages still do not draw it. The WebAIM Million survey of the top million home pages found in 2026: "A <main> element or main landmark was present on 46.1% of home pages, up from 42.6% in 2025." Adoption is rising, but more than half of the most popular pages on the web still give machines no explicit marker for where the main content is.
Note the phrase "or main landmark" in that finding. A <div role="main"> is a valid fallback that lands in the accessibility tree exactly where <main> does. If your site already uses a working role="main", there is no need to rip it out to chase the element; the signal is the same. Reach for the native element when you write fresh markup, because it carries the role with zero extra attributes.
One caveat applies to everything in this section: landmarks only help crawlers that see them. If your <main> and <nav> only exist after JavaScript builds the page in the browser, a crawler that does not run JavaScript never sees them. Put the structure in the server-rendered HTML, where every reader gets it.
A minimal skeleton covers most pages:
<body>
<header>
Site logo and name
<nav>Primary navigation links</nav>
</header>
<main>
<article>
The content this page exists for
</article>
</main>
<footer>Legal, contact, secondary links</footer>
</body>
Give every control a name
In the accessibility tree, every control has a name field, and that field can be empty. A <button> containing only an icon, with no text and no label, is a nameless node: an agent can see that a button exists but has no way to identify what it does. The requirement has been written down for years as WCAG success criterion 4.1.2, Name, Role, Value: every user interface component needs a programmatic name and role that software can read.
The web fails this at scale. Among the most common failures WebAIM detected across the million home pages:
- "Empty links" on 46.3% of home pages: a link with no text content and no label.
- "Empty buttons" on 30.6% of home pages: a button an agent can press but cannot identify.
- Missing form input labels on 51% of home pages: fields with nothing to say what goes in them.
The fixes are small and mechanical:
- Give every button and link real text where possible. For icon-only controls, add an
aria-labelthat says what the control does ("Search", "Close dialog", "Open menu"). - Give every form input a
<label>element tied to it. Placeholder text is not a label; it vanishes as soon as the user types.
And do not hide what can be focused. W3C's Using ARIA states it as the Fourth Rule of ARIA: "Do not use role="presentation" or aria-hidden="true" on a focusable element." A tabbable element inside an aria-hidden region is a trap: invisible to any agent reading the tree, but still reachable by a human user's keyboard, so the two audiences now experience different pages.
ARIA names are often applied by JavaScript after load, so a crawler that does not render sees the unlabeled version of your page. The names that end up in the rendered tree are what an operating agent gets, but only an agent that runs your scripts gets them.
Native HTML first, ARIA for the gaps
Everything above might read as an argument for sprinkling ARIA attributes everywhere. It is the opposite. The First Rule of ARIA is to avoid it when HTML already does the job: "If you can use a native HTML element or attribute with the semantics and behavior you require already built in, instead of re-purposing an element and adding an ARIA role, state or property to make it accessible, then do so." A <button> is a button in the accessibility tree with no extra work: role, name from its text, keyboard behavior, focusability, all built in. A <div> with a click handler is a nameless, roleless node until you hand-wire every one of those things back with ARIA, and each attribute is a chance to get it wrong.
WebAIM's data shows how often it goes wrong in practice: "Home pages with ARIA present had significantly more errors (59.1 on average) than pages without ARIA (42 on average)." That correlation needs its own caveat, and WebAIM supplies it: "This does not necessarily mean that ARIA introduced these errors (these pages were also more complex)." ARIA is not poison; complex pages use more of it and have more of everything, errors included. The honest reading is narrower: reach for ARIA only where native HTML cannot express the pattern, because an empty or incorrect ARIA attribute fills the accessibility tree with confident, wrong information, and a machine acting on wrong information is worse off than one facing an honest gap.
The same survey gives the wider trend, and it is the real motivator here. The web is getting less accessible, not more: WebAIM found "95.9% of home pages had detected WCAG 2 failures," up from 94.8% the year before, a rise it attributes partly to "automated or AI-assisted coding practices ('vibe coding')." Most of those detected failures are visual, but the two that bear directly on machine reading, empty links and empty buttons, are among the six most common. Against that backdrop, landmarks and accessible names are small, markup-level fixes with an unusual property: the same change serves the human visitor using a screen reader and the machine visitor reading the tree, at the same time, for the cost of an element name.
FAQ
Do AI crawlers actually use semantic HTML?
In two ways. Main-content extractors key on landmarks like <main> to separate your content from nav and footer boilerplate before a page is cited, though no answer-engine vendor documents its own extraction, and AI agents read the accessibility tree your markup builds to know what they can click. There is no measured citation lift, so do it for reliable parsing, not for a promised number.
What is the accessibility tree?
A semantic model the browser builds from your DOM, listing each element's role, name and state. It has powered screen readers for years, and some AI agents read it too, though others work from screenshots instead.
Should I add ARIA everywhere?
No. Use native HTML first and add ARIA only where HTML cannot express the pattern. In WebAIM's 2026 data, pages with more ARIA had more errors on average, and an empty or wrong ARIA attribute misleads a machine worse than an honest gap.
Is a <div role="main"> as good as <main>?
For the accessibility tree, yes: the role gives the same signal, so a working role="main" is fine and not worth ripping out. Where you are writing fresh markup, the native <main> element is simpler and harder to get wrong.
Do these tags help if my content loads with JavaScript?
They can still help agents that run your scripts, but many AI crawlers do not run JavaScript, so landmarks and labels injected client-side may never reach them; put the structure in the server-rendered response.
See if AI can read, trust, and cite your site
Add to ChromeFree · No signup · Every issue links back to a guide like this one