Lists, tables and other formats

Published

Structured formats spell out the relationships on a page: a list declares that its items are parallel, and a table declares which value belongs to which label. Explicit labels remove the guesswork about which value belongs to which thing, though whether a given engine's citation choice turns on that is not something any vendor documents. The measured support is narrower than the advice around it. Use lists and tables because they make relationships explicit and read better, not as a documented citation lever.

The numbers behind that caution are worth having. GEO-SFE, a 2026 preprint, reported a 17.3% aggregate citation-rate gain across six engines from a package of structural changes rather than from lists or tables on their own, and separately reported a 43% advantage for structured formats over prose on extraction accuracy, which measures how cleanly text comes out rather than whether it gets cited. A separate 2026 controlled study by Vishwakarma et al. found that formatting-only edits had little impact on which of two competing sources got cited first.

Format is a tool for the shape of your data, not a decoration to pile on. The study that measured this most directly puts both halves in one sentence: structured formats extracted better than equivalent prose, and past roughly a third of the page being structured elements, reading suffered. So there is an optimal amount of structure, not a maximum, and it depends on the markup being real.

Why structure extracts better than prose

An AI engine does not read a page the way a person does. It works at the level of sections and chunks. Google says as much about its own systems in its ranking systems guide: "Passage ranking is an AI system we use to identify individual sections or 'passages' of a web page to better understand how relevant a page is to a search." GEO-SFE groups the engines it tested by architecture: Google SGE and Bing Chat retrieve in batches, Perplexity searches over multiple rounds, and ChatGPT and Claude retrieve in real time. The vendors publish no common rule for how a list or a table is extracted.

The mechanism is straightforward, though no engine documents it: a list item encodes membership and a table header associates a label with a cell, so the relationship is already in the markup rather than something a reader or a model has to reconstruct from a sentence.

The effect has been measured. The GEO-SFE study tested structural optimization on its own, independent of what the words say: "Evaluation across six generative engines demonstrates consistent 17.3% citation improvements, with subjective assessments revealing 18.5% average enhancement in perceptual quality." Format diversity, the mix of structural elements on a page, was among the features the study weighted most heavily for the ChatGPT and Claude class of engines. The same paper puts a number on the format itself and on its limit in one sentence: "Structured formats (lists, tables) demonstrate 43% higher extraction accuracy than equivalent prose in our experiments, but excessive structure (F_d > 0.35) disrupts reading flow and reduces human comprehension."

The result does not show that more structure is better, and the study says where the turn is: past roughly a third of the page being structured elements, the same paper reports that reading flow and human comprehension suffer.

A September 2025 benchmark by Improving Agents, 1,000 questions over synthetic records, found a markdown table (51.9%) barely beat plain prose (49.6%), with markdown key-value pairs best at 60.7% and CSV worse than prose at 44.3%. The real win from lists and tables is extraction and labeling reliability, not a comprehension miracle. The widely repeated "tables are cited 2.5 times as often" figure comes from an agency blog post rather than a controlled study, and no published measurement supports it.

This article covers formats only; writing citable passages, structuring headings, building FAQ sections and sourcing facts are separate subjects in this series.

Lists

Use a real ordered list for a sequence, anything where the order carries meaning: steps in a setup, stages in a process, a ranking. Use a real unordered list for a set of parallel options or points, where order is arbitrary.

The common failure mode is a list that is not really a list: bullet glyphs typed into a paragraph, or sentences separated by line breaks with no <li> markup behind them. Visually it reads as a list, but a real <li> encodes list membership in the markup and a typed bullet glyph encodes nothing, so in the HTML there is nothing marking the items as items. The fix is simple: if the content is a list, mark it up as one, with real <ul> or <ol> elements. That is a sound reason to use real list elements, though no published measurement isolates a citation effect from it.

Three habits make a list useful to readers and machines:

  • One idea per item. An item that bundles two points is a chunk with two subjects, and neither quotes cleanly.
  • Keep items parallel in structure. If the first item is a noun phrase, make them all noun phrases; parallel items read as a coherent set.
  • Do not force prose into list form. A flowing argument chopped into bullet fragments is not a list, it is a paragraph with its connective tissue removed.

Tables

A data table is the natural format for a multi-attribute comparison, because headers keep every value tied to a labeled row and column. An engine reading it can cite the number and state what the number means. That is why a three-way comparison belongs in a table rather than three paragraphs: in prose, the engine has to reassemble the grid; in a table, the grid is the markup.

The value only materializes if the table is machine-readable, and the common failure is a table with no header cells. <th> cells state which label belongs to which axis; without them, that association is absent from the HTML and has to be inferred from position. A minimal extractable table looks like this:

<table>
  <thead>
    <tr>
      <th>Plan</th>
      <th>Monthly cost (USD)</th>
      <th>Storage (GB)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Basic</td>
      <td>10</td>
      <td>100</td>
    </tr>
  </tbody>
</table>

A few construction rules follow from the same logic. Put units in the header labels: "Monthly cost (USD)" beats "Price", because the unit travels with the value when a cell is quoted alone. Keep cell contents plain text; a checkmark image conveys nothing to a model reading the HTML. Avoid merged cells, which break the row and column labeling. Put the primary entity in the leftmost column, where readers and machines both look for it. And make sure the table exists in the page's delivered HTML rather than being rendered late by JavaScript.

Never save a table as an image. A screenshot keeps the look and discards the structure. Some systems are multimodal and may read the pixels, but the row and column relationships are no longer in the HTML, so publish the table as a table.

Definition lists and blockquotes

Two further formats are semantically explicit but low-stakes: worth using where the content genuinely calls for them.

A definition list is the most explicit way to mark up a glossary. Each term sits in a <dt> element and its definition in the <dd> that follows, so the term and its definition are paired in the markup rather than left to be worked out from a sentence:

<dl>
  <dt>Crawler</dt>
  <dd>A program that fetches pages so a search or AI system can index them.</dd>
</dl>

That precision is worth having on a page that actually is a glossary or a reference of terms. It is not worth contorting ordinary prose into a definition list, and the absence of a <dl> is never a defect.

A <blockquote> marks quoted material as exactly that: someone else's words, presented as attributable evidence. Use it for a named source's words rather than folding them into plain prose, because the boundary between your claims and your source's words is then explicit in the markup. A page that quotes nobody loses nothing.

Comparisons on product and commercial pages

On a page whose job is to help someone choose, an explicit comparison earns its keep: a comparison table, a clear "X vs Y" treatment, or a pros-and-cons section. In a controlled two-document test, Vishwakarma et al. found that a page containing comparisons was cited ahead of an otherwise identical page without them, in at least four of the six models tested. The study varied whether comparison content was present, not how it was formatted.

The scope is narrower than most advice admits. Comparison content belongs on commercial and product pages, where a choice is actually being made. An explainer that never compares anything is not deficient for it. Where you do compare, prefer the table form, so the contrast is a clean data object rather than alternating paragraphs.

Match the format to the data

One rule ties all of this together: pick the format that mirrors the shape of the information, and stop there.

Shape of the information Format HTML element
Steps in a sequence Ordered list <ol>
Parallel options or points Unordered list <ul>
Multi-attribute comparison or specs Data table <table> with <th>
A term and its meaning Definition list <dl>
A source's exact words Blockquote <blockquote>
A single flowing argument Paragraph <p>

Over-formatting is a real cost, not a free win. Chopping continuous reasoning into bullet fragments strips out the logic that made it an argument. Wrapping every stray number in a one-row table adds structural noise without adding a single labeled relationship. Both make the page harder to read. Government style manuals reached the same conclusion years before any of this involved AI: the Australian Government Style Manual tells writers to limit the number of lists, because "content with too many lists is hard to follow" and "the content should flow so people can read it easily". GEO-SFE reports a readability cost above its structure threshold, but no study compares over-formatted pages with pages whose format was matched to the information, so treat "more structure is better" as unsupported rather than disproved, and format for the information you actually have. Let a paragraph be a paragraph.

Frequently asked questions

Do tables and lists really get cited more by AI?

Less directly than the advice suggests. GEO-SFE measured a 17.3% average citation-rate gain across six engines from a package of structural changes, not from lists or tables on their own, and Vishwakarma et al. found formatting-only edits had little impact. The "tables are cited 2.5x" figure circulating in agency posts has no controlled measurement behind it.

Should I turn my prose into lists and tables everywhere?

No. Match the format to the data: a sequence wants an ordered list, a comparison wants a table, a flowing argument wants a paragraph. Over-formatting adds structural noise and hurts readability, and no study shows it earns anything back. There is an optimal amount of structure, not a maximum.

What makes a table extractable by AI?

Real markup. Use an HTML table with <th> header cells; without headers, which label belongs to which axis is absent from the HTML and has to be inferred from position. Put units in the headers, keep cells plain text, avoid merged cells, and publish the table as a table rather than a screenshot, which leaves the row and column relationships out of the HTML.

Do I need definition lists and blockquotes?

Only where the content calls for them. A definition list is ideal for a real glossary, and a blockquote for a genuine quotation, but nothing published suggests a page is disadvantaged for having neither. Do not contort ordinary prose into these formats to tick a structural box.

Does a list of bullet points typed into a paragraph count?

No. If the bullets are typed glyphs or line-broken sentences without real list markup, nothing in the HTML marks the items as items. Use real <ul> or <ol> markup so list membership is encoded rather than only drawn.

See if AI can read, trust, and cite your site

Add to Chrome

Free · No signup · Every issue links back to a guide like this one