Writing for AI citation: answer first, stand alone

Published

Ask ChatGPT, Perplexity or Google's AI Overviews a question and you get a short written answer with links, not a reproduction of any source page. What each link is doing there is standing in for a specific claim, which is usually traceable to a paragraph or a sentence rather than to a whole page. That changes the unit of work for anyone who wants to be cited. The unit is the passage, not the article.

Writing extractable passages is real craft, but it is a smaller and more concrete job than the mega-checklists circulating in this field suggest. It comes down to four habits: open with the answer, keep every passage self-contained, define your terms crisply, and cover the sub-questions a query fans out into as focused passages, not as an encyclopedia.

How engines lift passages

Two vendors document passage-level machinery in their own products. Google runs passage ranking in Search: "Passage ranking is an AI system we use to identify individual sections or 'passages' of a web page to better understand how relevant a page is to a search." (Google ranking systems guide). Anthropic documents finer granularity in a different setting. Its Citations API chunks documents a developer supplies into sentences, so Claude can cite a single sentence or a longer passage from them: the feature "lets Claude ground its answers in source documents. Claude can now provide detailed references to the exact sentences and passages it uses to generate responses, leading to more verifiable, trustworthy outputs." (Anthropic Citations announcement). These examples show passage-level selection in two different settings, but neither vendor documents how public-web answer tools choose cited passages. ChatGPT search, Perplexity and Copilot present claim-level citations in their interfaces, but none of the three publishes how a source passage is selected, so treat the resemblance as an observation of the interface rather than a documented mechanism.

Write each paragraph so it can work on its own, because retrieval can match a question to a passage. The relevance and quality of the surrounding page still matter. The craft in the rest of this article therefore works at paragraph level. This paragraph-level advice draws on the retrieval mechanics above. It also aligns with one measured result: the GEO study (Aggarwal et al., 2024) found that improving fluency and readability produced "a significant visibility boost of 15-30%" in generative engine responses. The precise capsule numbers that circulate alongside this advice, such as percentages of citations supposedly won by openings, are industry rules of thumb, not measured findings, and this article does without them.

Passage craft also sits on top of page-level basics covered elsewhere in this series: heading structure, content length, FAQ sections, and facts and citations. None of that is repeated here. The subject is the passage itself.

Open with the answer

State the conclusion first, then the context. Journalists call this the inverted pyramid; the military calls it BLUF, bottom line up front. The principle is the same either way: a reader who stops after the first sentence still leaves with the answer, and a model that lifts only your opening paragraph lifts something complete. It applies twice over. The page's first paragraph should carry the bottom line of the whole piece, and each section should answer its own heading before it elaborates. AutoGEO (Wu et al., 2025), a study that mined citation rules from generative engine behavior, points the same way: the rules it extracted favor answers stated up front and explanations that go into causes and mechanisms rather than restating the surface.

A strong opening is a real, self-contained paragraph, not a tagline or a hero line. Three openings fail reliably. The buried lede: two or three sentences of throat-clearing before the actual point arrives, the "in today's fast-moving landscape" pattern. The dependent opening: a first paragraph that only makes sense after you have read something further down the page. And the empty opening: a one-line fragment that looks fine on the page and says nothing when quoted on its own.

The common advice that the opening paragraph specifically is what engines lift as a page's summary comes from years of featured-snippet observation; it is industry consensus, not a controlled result. The durable principle does not depend on it: front-load the answer everywhere, because any passage on the page may be the one that gets lifted.

Write self-contained passages

A paragraph may be read with very little of its surroundings, so each one has to survive being read on its own. The concrete test is clear referents: who, what, and when should resolve inside the passage itself. Composition teaching has a name for the habit, the topic sentence: a paragraph that opens by naming its own subject. A paragraph that opens with as mentioned above, or that hangs on an unresolved it, this, or the tool, breaks the moment it is separated from the paragraphs that gave those words meaning.

Fails when lifted:
  As mentioned above, it handles this automatically on most plans.

Stands alone:
  Acme's scheduling feature converts time zones automatically
  on every plan except the free tier.

The fix is mechanical: name the subject again. Repeating a product name or a full name where a pronoun would flow more smoothly feels redundant to the author, and it is exactly what lets the paragraph read correctly anywhere it lands.

Two further failure modes make a passage a poor candidate for retrieval. Vague passages: generic prose that gestures at an area without saying anything specific. And bare claims: assertions with nothing behind them. How to supply the backing, with statistics, quotations, and named sources, is its own topic, covered by the facts and citations article in this series.

Vagueness also has a subtler form: the semantic near-miss. Retrieval systems match the meaning of the question against the meaning of the passage, and content that is merely about the topic loses to content that answers the question. A paragraph broadly discussing AI crawler behavior will not be retrieved for "does GPTBot render JavaScript"; the paragraph that states specifically what GPTBot does will. Write the specific answer, not a gesture at the area.

Define your terms crisply

A clean definition near the top of a page is one of the most liftable passages you can write. "What is X" is one of the most common question shapes, and a plain declarative definition answers it in a single self-contained passage. Put it where the term first appears, and state it in the unglamorous form that extraction favors: X is a Y that ....

What makes a definition well formed was settled long before search engines existed, and the classical rules translate directly. Non-circular: the definition must not use the term to define itself. Genus and differentia: name the category the thing belongs to, then the feature that separates it from everything else in that category. Self-contained: the definition must not lean on references resolved elsewhere on the page. Non-obscure: never explain one piece of jargon with more jargon that the same reader would also have to look up.

Circular:  An answer engine is an engine that answers questions.

Crisp:     An answer engine is a search tool that responds with
           a written answer assembled from sources, rather than
           a list of links.

These rules govern form, not truth. A definition can be perfectly shaped and still wrong; checking the facts is a separate exercise from checking the form. But form is what makes a definition usable on its own. A vague or circular definition gives a reader, or anything excerpting the page, nothing usable, however sound the facts behind it are.

Cover the sub-questions, as focused passages

Answer engines do not treat a question as a single query. Google documents this for its own AI surfaces: "Both AI Overviews and AI Mode may use a 'query fan-out' technique, issuing multiple related searches across subtopics and data sources, to develop a response." (Google AI features documentation). A question like "is X worth it for a small team" fans out into cost, setup time, alternatives, and risks, and the engine goes looking for a passage to resolve each branch.

The writing consequence follows from everything above: identify the sub-questions your readers actually bring, and answer multiple related questions on one page, each in its own self-contained passage. A model resolving the cost branch of a fan-out needs a paragraph about cost that stands alone, not a cost figure scattered across a narrative. The same goes for the risks branch, the alternatives branch, and whichever branches are native to your topic.

Fan-out is regularly read as a mandate to make every page encyclopedic. It is not. Coverage is a property of the retrieved set, not of any one page: the engine stitches its answer from passages across many sources, and a focused page that answers one sub-question well is a candidate for that branch of the answer. Passage ranking points the same way: Google says it weighs individual sections in addition to the relevance of the overall page, so a strong passage can be found even when it sits deep in a longer piece. So add the adjacent sub-questions your readers genuinely expect, and stop there. Padding a page with everything tangentially related buries the passages that actually answer something.

Frequently asked questions

Do AI engines quote whole pages or passages?

Passage-level selection is the common model, but only Google documents a passage-ranking system, and Claude's sentence-level citation is a feature for documents a developer supplies. The other engines do not publish how they select from a web page. A strong, relevant paragraph may be surfaced even when the rest of the page covers a broader topic, which is why the paragraph must make sense on its own.

How long should an answer passage be?

There is no verified magic length. The capsule word counts that circulate, 40 to 60 words, the answer inside the first 100, are industry rules of thumb, not measured thresholds. Write a passage that answers one question completely and stands on its own; the right length follows from that.

Does covering more subtopics get me cited more?

Not by breadth alone. Query fan-out rewards a focused passage that nails a specific sub-question, and coverage is assembled across many sources rather than demanded of one page. Add the adjacent questions readers expect, but padding a page with everything tangential dilutes the quotable passages you already have.

What makes a passage self-contained?

Clear referents: who, what, and when, stated inside the passage itself. A paragraph that opens with "as mentioned above" or hangs on an unresolved "it" or "this" breaks when an engine lifts it out of context. Name the subject so the passage reads correctly on its own.

What makes a definition quotable?

Form as much as facts. A quotable definition is non-circular, names the category and the distinguishing feature (genus and differentia), stands on its own, and does not explain jargon with more jargon. State it plainly, as "X is a Y that ...", close to where the term first appears.

See if AI can read, trust, and cite your site

Add to Chrome

Free · No signup · Every issue links back to a guide like this one