llms.txt, Markdown and the agent web

Published

A small family of files now promises to make your site easier for AI to read, sitting at your site root beside robots.txt and sitemap.xml: an llms.txt index, a Markdown copy of each page, a manifest describing the actions an agent can take on your site. Guides call this the agent web, and most of them tell you to ship all of it. The question that decides whether it is worth your time is simpler: does any AI actually read these files?

This page is about the files you publish for AI. Blocking a crawler's access with robots.txt, and controlling how AI may use content it already has, are separate topics. The answer you can act on now: today the AI answer engines mostly do not consume these files, and the one file with a real, present-day consumer serves AI coding tools, not the search engines that cite you in answers.

Who actually reads these files

The content files in this family are aimed at one of two very different readers, and almost every misunderstanding about the agent web comes from mixing them up. Agent manifests are a third case and are covered further down: they address software that acts on your site rather than reads it.

The first audience is the answer engines: ChatGPT, Perplexity, Gemini and Copilot, the systems that quote and cite web pages inside generated answers. None of them documents consuming llms.txt, llms-full.txt or an agent manifest when deciding what to cite. The second audience is coding agents: Cursor, GitHub Copilot, or Claude running inside an editor, tools that work against a site's documentation while writing code. These do fetch Markdown and docs indexes, because plain text costs far fewer tokens than rendered HTML.

One detail makes the split easy to remember. OpenAI, Anthropic and Perplexity all publish an llms.txt on their own developer documentation sites, yet none of them says its crawlers consume yours when building search answers. They publish these files for their own documentation; none of them documents using yours to select citations.

The practical takeaway for everything below is that these files are cheap, low-risk machine-readability hygiene, not a visibility lever, and no citation number exists to promise you. The sections below separate the most-hyped files from the one format with a present-day consumer and the action standards still being proposed.

llms.txt, and what the data says

llms.txt is a Markdown file at your site's root that lists your most important pages for language models. Its author, Jeremy Howard of Answer.AI, frames it on llmstxt.org as "A proposal to standardise on using an /llms.txt file to provide information to help LLMs use a website at inference time." It is a proposal, not a standard, and it is deliberately light. The format requires almost nothing: "An H1 with the name of the project or site. This is the only required section". Everything else, the summary blockquote and the link sections, is optional. A minimal file looks like this:

# Example Company

> Example Company builds scheduling software for small clinics.

## Documentation

- [Getting started](https://example.com/docs/start.md): Installation and first setup
- [API reference](https://example.com/docs/api.md): Endpoints, parameters and errors

The proposal is easy to implement, which is why adoption ran well ahead of evidence. When Ahrefs looked at actual behavior in its llms.txt study, it "studied 137,000 domains" and "found that 28% publish an llms.txt file". Then it read the server logs: "Of the ~38,000 domains in our study with a valid llms.txt, 97% received zero requests for it in May 2026. No bots, no humans, nothing." Other measurements point the same way. SE Ranking, in a 2026 study across roughly 300,000 domains, found no correlation between publishing the file and AI visibility. Trakkr, also in 2026, across 37,894 domains, found "zero measurable advantage". Otterly's 2026 log analysis put llms.txt at 0.1% of AI-bot visits. None of this is mysterious: a file that is almost never fetched has almost no way to reach an answer, which is exactly what the citation studies measure.

Google's position matches the logs. Its AI optimization guide says: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." And on publishing them anyway: "Doing so will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them." Asked about the file's future, Google's John Mueller was just as direct, as reported by Search Engine Journal: "I don't think anyone knows, it's purely speculative for now (the file has existed for years, yet none of the AI systems use it, what does it mean?)."

There is one place where Google appears to point the other way. Chrome's Lighthouse tool added an "Agentic browsing" category with an audit of llms.txt structure. Its documentation says: "The llms.txt file is an emerging convention used to provide a machine-readable summary of a website's content, specifically designed for LLMs and AI agents." and "Without this file, agents may spend more time crawling the site to understand its high-level structure and primary content." That is not a contradiction of the mythbusting above; it is a different scope. Google Search is talking about its ranking and citation pipeline, which ignores the file. Lighthouse is a developer tool scoring how ready a page is for browsing agents, and it labels the whole category clearly: "The Agentic Browsing category and WebMCP support are experimental and based on proposed standards." An experimental agent-readiness audit is not a search endorsement.

llms-full.txt

llms-full.txt is usually described as the companion file: instead of an index of links, a single file bundling all of your page content. What matters most is where it comes from. The llmstxt.org proposal does not define it; the file appears nowhere in the spec. It is a convention popularized by the documentation platform Mintlify, whose llms.txt docs describe it plainly: "The llms-full.txt file combines your entire documentation site into a single file as context for AI tools." The same page calls llms.txt "an industry standard", which overstates it; it is a proposal. Any guide telling you llms-full.txt is part of the spec is repeating a myth.

The file is also unmeasured, and it inherits the smaller file's problem. If crawlers barely fetch a short index at your root, there is no reason to expect them to fetch a much larger bundle of everything you have written. Publish it only if a specific tool you use documents that it wants the file.

Markdown versions of your pages

This is the one lane with a real, present-day consumer. The llmstxt.org proposal also covers per-page Markdown: "We furthermore propose that pages on websites that have information that might be useful for LLMs to read provide a clean markdown version of those pages at the same URL as the original page, but with .md appended." In practice there are three ways to offer it: a parallel .md URL, an advertisement in the page head via <link rel="alternate" type="text/markdown">, or content negotiation, where the server returns Markdown to any client that sends Accept: text/markdown.

The case for it is token economics. Markdown strips navigation, ads and wrapper markup, leaving the text a model actually needs. Cloudflare measured the difference on one of its own pages in Markdown for Agents: "This blog post you're reading takes 16,180 tokens in HTML and 3,150 tokens when converted to markdown. That's a 80% reduction in token usage." That is a single page, not a study, but the mechanism is plain: less markup means fewer tokens, and fewer tokens make a page cheaper to ingest.

The part most guides skip is who asks for it. AI coding agents such as Cursor and OpenCode request Markdown when they work against a site's documentation. The search and citation crawlers do not. Checkly's February 2026 tests found only Claude Code, Cursor and OpenCode asking for text/markdown, and the acceptmarkdown.com tracker still had ChatGPT's browsing fetcher taking raw HTML in April 2026. Serving Markdown is a genuine efficiency for developer tools reading your docs, not a citation lever for the answer engines.

Agent manifests: MCP, WebMCP, agents.json

Agent manifests are a different kind of file altogether. llms.txt and Markdown describe content; a manifest describes actions, the things an agent could do on your site: call an endpoint, add an item to a cart, fill and submit a form. Nothing in this category is a shipped standard. Discovery of MCP servers through a well-known URL lives in unmerged proposals (SEPs), not in the protocol itself. agents.json is an early v0.1.0 format from a single startup. And OpenAI's ai-plugin.json manifest, the closest thing the category had to a deployed convention, was retired when ChatGPT plugins were sunset.

The direction of travel is WebMCP, a proposed browser API, currently experimental in Chrome, that lets a page expose its tools and forms directly to an in-browser agent. Google's John Mueller, in the same exchange reported by Search Engine Journal, said "I like the WebMCP approach", and added that the most basic agentic optimization is simply not blocking agents from your site. This corner of the agent web has the clearest goal, an agent that can act rather than just read, but it is still a proposal, not something to rebuild your site around today.

If you publish an llms.txt anyway

None of the evidence above makes llms.txt harmful. The file is cheap to create and close to risk-free, though Ahrefs notes a curated index also makes your content easier for competitors to scrape. It has a forward-looking case too: if agent traffic grows, a curated index is exactly the kind of file an agent would want. There is also an immediate internal benefit even if nobody reads it: writing a good llms.txt forces you to decide which pages matter most and to state what your site is for in plain language.

A good file, by the proposal's own format, has four properties:

  • A summary that reads like the opening of a Wikipedia entry, factual and specific, not a marketing hero line.
  • A meaningful one-line description after each link, saying what the page actually answers.
  • A curated table of contents of your key pages, not a dump of your whole sitemap.
  • An ## Optional section for secondary links, which the proposal marks as the material a model can skip when its context budget is tight.

Keep it Markdown, keep it at your root, and keep it current when your key pages change. Treat the file as a curation exercise and cheap optionality, not as a route to citations.

FAQ

Does publishing an llms.txt get me cited by ChatGPT or Google?

There is no evidence that it does. Multiple large studies found no citation benefit, Google says its Search ignores the file, and John Mueller calls it purely speculative.

Do AI crawlers even fetch my llms.txt?

Mostly not. In Ahrefs' study of 137,000 domains, 97% of the published llms.txt files received zero requests in a month, from bots or from anyone else, and Otterly's logs put the file at 0.1% of AI-bot visits.

Is llms-full.txt part of the standard?

No. The llms.txt proposal does not define it. It is a convention from the docs platform Mintlify, which bundles an entire documentation site into one file for AI tools. Treat any guide that calls llms-full.txt "part of the spec" with caution.

Should I serve a Markdown version of my pages?

It is worth it mainly for AI coding tools such as Cursor, which request Markdown to save tokens when reading documentation. ChatGPT's browsing fetcher was still taking plain HTML rather than Markdown in April 2026, so treat it as developer-agent efficiency rather than a play for citations.

What is WebMCP, and should I care yet?

WebMCP is a proposed browser API that lets a page expose its actions, such as tools and forms, to in-browser AI agents. It is experimental in Chrome, and Google's John Mueller says he prefers it to llms.txt, but it is still a proposal.

See if AI can read, trust, and cite your site

Add to Chrome

Free · No signup · Every issue links back to a guide like this one