Structured data for AI

Published

HTML tells a machine how your page is arranged, not what the words on it mean. The schema.org documentation illustrates the gap with a single word: "'Avatar' could refer to the hugely successful 3D movie, or it could refer to a type of profile picture" (schema.org getting started). A human resolves that ambiguity instantly from context. A crawler parsing millions of pages may infer the right sense from the surrounding prose, or may not. Structured data is how you tell it, and it is one of the few purely technical changes that state your meaning outright rather than leaving a machine to infer it.

What structured data is

Structured data is a standardized format for labeling what the content of a page means. Google's introduction to the topic states the purpose in one sentence: "You can help us by providing explicit clues about the meaning of a page to Google by including structured data on the page" (Google structured data guide). The clues are typed properties: this string is a product name, this number is its price, this person is the author of this article. Instead of hoping a crawler infers the right meaning from surrounding prose, you state the meaning outright.

Three names circulate for this practice, and they fit together simply. "Structured data" is the format itself. Schema.org is the vocabulary nearly everyone writes it in; the project describes itself as providing "a collection of shared vocabularies webmasters can use to mark up their pages in ways that can be understood by the major search engines: Google, Microsoft, Yandex and Yahoo!" And "schema markup", or just "schema", is the everyday name for marking pages up with that vocabulary. In practice all three terms point at the same work.

The vocabulary can be written in three formats: JSON-LD, microdata, and RDFa. Microdata and RDFa weave attributes into your existing HTML tags, while JSON-LD sits in its own block, separate from the visible markup. All three are valid, but the default choice is settled: "In general, Google recommends using JSON-LD for structured data if your site's setup allows it, as it's the easiest solution for website owners to implement and maintain at scale". A minimal JSON-LD block for an article looks like this:

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "How to Brew Better Coffee at Home",
  "author": {
    "@type": "Person",
    "name": "Maria Chen"
  },
  "datePublished": "2026-05-12"
}
</script>

The @context line declares which vocabulary the block uses, and @type declares what kind of thing it describes. Everything else is properties of that thing, stated as plain key-value pairs a machine can read with no language understanding at all.

How AI engines use structured data

The value of those explicit labels appears when a machine consumes the page. Google is direct about what it does with them: "Google uses structured data that it finds on the web to understand the content of the page, as well as to gather information about the web and the world in general, such as information about the people, books, or companies that are included in the markup." The second half of that sentence matters more than it first appears. Your markup describes more than one page; it feeds the entity knowledge that connects your organization, your authors, and your products across the web.

Search Engine Land's analysis of schema in AI search describes the same mechanism from the consuming system's side: markup gives a machine "a set of explicit entity, brand, product, price, author, and topic fields that a system can map to, rather than inferring everything from unstructured prose" (Search Engine Land). Language models are good at inferring, but inference from prose can misattribute an author, confuse two similarly named products, or read a crossed-out price as the current one. An explicit field removes the guess.

What each engine says for itself differs, and the differences are worth knowing. Google documents its use plainly, as quoted above, and is equally plain about the limits: structured data is not a ranking factor. What it still buys in Google Search is rich results: enhanced listings that can show ratings, prices, and dates. Google also publishes case studies reporting the click-through gains those features produce. In one of them, Rotten Tomatoes reported a 25% higher click-through rate on pages with markup. That figure is a rich-result gain in classic Search, not an AI citation number, but it is a documented first-party payoff from the company that also runs AI Overviews, AI Mode, and Gemini.

Microsoft is the AI vendor most clearly on record telling publishers to use it. Writing about how content gets included in AI search answers, a Bing principal product manager states that "Schema is a type of code that helps search engines and AI systems understand your content," describes it as "turning plain text into structured data that machines can interpret with confidence," and puts it in the core checklist: "Structure your content: Use schema, clear headings, and modular layouts" (Bing AI search guidance). That is guidance from inside an AI answer pipeline, from the search engine behind Microsoft Copilot. The same post stays honest about the ceiling, noting "there's no secret sauce that guarantees selection in AI answers".

For the pure-LLM engines the record is thinner, and it is worth saying plainly. OpenAI does not document that ChatGPT reads the schema.org markup on your pages. Anthropic does not document that Claude does. Perplexity does not document it either. Plenty of third-party guides assert confidently that all the AI engines use schema; none of those claims traces back to vendor documentation. Marking up your pages may well help these systems the same way it helps any machine reader, but for ChatGPT, Claude, and Perplexity that is an inference from mechanism, not a documented behavior.

One recurring confusion deserves its own paragraph. OpenAI and Anthropic both offer an API feature called "Structured Outputs", and it is regularly pointed to as proof that these companies rely on schema markup. It is unrelated. In OpenAI's words, "Structured Outputs is a feature that ensures the model will always generate responses that adhere to your supplied JSON Schema, so you don't need to worry about the model omitting a required key, or hallucinating an invalid enum value" (OpenAI structured outputs). That feature constrains the JSON the model sends back to a developer. It says nothing about reading schema.org markup on your web pages.

Does structured data get you cited?

Google says explicitly that structured data is not required for its generative AI features: "Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add. However, it's a good idea to continue using it as part of your overall SEO strategy, as it helps with being eligible for rich results on Google Search" (Google AI optimization guide). Its AI features documentation is broader still: "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary" (Google AI features docs). No other vendor covered above documents page-level schema.org as a requirement either.

The measured evidence for a citation lift is where skepticism belongs, because the claims in circulation run well ahead of the data. The same Search Engine Land analysis sums up the state of the research: "To date, there are no peer-reviewed studies on schema's impact on AI search visibility, or controlled experiments on LLM citation behavior and schema markup." It also cites counter-evidence: "A December 2024 study from Search/Atlas found no correlation between schema markup coverage and citation rates." The confident figures you will meet elsewhere, claims that some large share of AI-cited pages carry markup or that adding it lifts citations by a fixed percentage, are third-party correlations drawn from pages that differ in dozens of other ways. None of them is a controlled result.

That leaves the payoff you can actually count on: machine clarity and real Search features. Google's guidance for AI features still lists "Making sure your structured data matches the visible text on the page" among the fundamentals worth doing. Structured data earns its place the way clean HTML does. Google and Bing document reading it, and any other system that parses schema.org gets an unambiguous version of what you meant instead of a guess. That is the case for doing it, whether or not a citation percentage ever materializes.

Getting the markup right

Structured data only works when it parses. JSON-LD is code, and crawlers treat it like code: a block with a syntax error, a stray trailing comma, or a truncated brace fails to parse and is dropped in its entirety, silently. The page renders normally, so nothing looks wrong, but the entity you described never reaches the engine. A missing top-level @context is the subtler version of the same failure: property names no longer expand to schema.org terms, so headline and author carry no meaning even though the JSON is valid. Nested items inherit that context, so they do not each need their own. @type is a different matter: it is optional in JSON-LD, and a node without it still parses, it is simply untyped, which usually means no engine feature can act on it. Markup that does not parse is not partially credited, it is simply gone. Markup that parses but is incomplete is a milder case: it is still read, but it will not qualify for the Search features that depend on the missing properties.

The second failure mode is markup that parses but does not match the page. Google's guideline is blunt: "don't add structured data about information that is not visible to the user, even if the information is accurate." Markup is a machine-readable restatement of what the page already shows, not a side channel for claims the visitor never sees. Ratings the page does not display and prices that differ from the visible ones make your markup untrustworthy exactly where trust is the point.

A third failure mode is prioritizing breadth over accuracy. Google again: "it is more important to supply fewer but complete and accurate recommended properties rather than trying to provide every possible recommended property with less complete, badly-formed, or inaccurate data." A small block that states five facts correctly beats a sprawling one that guesses at twenty.

Before shipping, validate. Google's Rich Results Test reports whether your markup parses and qualifies for Search features, and the Schema.org validator checks it against the vocabulary itself. Between them the two tools catch syntax and vocabulary errors and tell you whether Google considers the item eligible for a feature. Neither can tell you whether the markup matches what the page actually shows, so check that yourself.

What to mark up

Mark up the real entities the page is about. A company site should describe its Organization, with its name, logo, and contact points. An article should carry Article markup with its headline, author, and dates. A product page should state the Product, its price, and its availability. The test is whether a machine, reading only your markup, could recite the basic facts of the page without guessing at any of them.

Entity connections extend the same idea beyond your site. Google notes that it "can make general use of the sameAs property and other schema.org structured data. Some of these elements may be used to enable future Search features, if they are deemed useful." The sameAs property ties your Organization or Person to its profiles elsewhere on the web, which should in principle help a machine confirm that the entity on your page is the same one it already knows about. That disambiguation benefit is an industry inference from how knowledge graphs reconcile entities, not a use Google has confirmed.

One thing not to chase: FAQ and HowTo rich results no longer appear in Google Search, so marking up those blocks will not buy the result features they once did. The markup remains valid, and question-and-answer content is still worth structuring well for its own sake, but a rich-result payoff is no longer among the reasons.

Frequently asked questions

Does structured data help you get cited by AI?

It is not a documented or measured citation lever, and no controlled study shows a lift. What is real: Google and Bing use it to understand your content, and it powers rich results in Search. Add valid structured data for that clarity, not for a promised citation number.

Which format should I use: JSON-LD, microdata, or RDFa?

JSON-LD. It keeps your markup in one block instead of scattered through your HTML. Microdata and RDFa remain valid if implemented correctly.

Is "structured data" the same as "schema"?

Effectively yes. Structured data is the practice and the format; schema.org is the shared vocabulary almost everyone uses; "schema markup" is the everyday name for marking your pages up with it.

Do ChatGPT, Claude, and Perplexity read my structured data?

None of the three documents doing so, so treat any benefit there as unproven. Google and Bing are the engines on record using structured data.

Is structured data required to appear in AI Overviews?

No. Google says there is no special schema.org markup you need to add and no additional requirements for AI Overviews or AI Mode. It is still worth doing, done correctly, for machine understanding and rich results.

See if AI can read, trust, and cite your site

Add to Chrome

Free · No signup · Every issue links back to a guide like this one