Content freshness and dates

Published

"Fresh content ranks better in AI search" is the advice on every optimization checklist, and it quietly merges two different questions. The first: are the facts on your page still true? The second: does your page carry a current, trustworthy date stamp? The first is the half with direct evidence about answer quality. The second has been measured too, but what the measurement shows is a model bias that a fake date triggers as readily as a real one.

Google documents freshness systems that fire only when a query calls for recency, measured citation age varies from product to product, and the controlled experiments on generative engine optimization found no freshness lever at all. The goal, then, is honest currency: pages whose claims are correct and whose dates mean something, not a treadmill of cosmetic date changes.

How AI answer engines use recency

ChatGPT, Perplexity, Gemini, Copilot and Google's AI Overviews do not answer from a frozen snapshot. For many questions they retrieve live web pages and compose an answer from what they retrieve. Only two of these products document their weighting, and in both cases recency is part of it. Google says so plainly in its ranking systems guide: "We have various "query deserves freshness" systems designed to show fresher content for queries where it would be expected." Google's own illustration of the idea: someone searching about a just-released movie or an earthquake that happened this morning expects fresh pages, while plenty of other searches are served perfectly well by older material. For queries that do not call for recency, no vendor publishes what decides the ranking instead.

Microsoft describes the same pressure from the infrastructure side. The Microsoft Web IQ announcement says the grounding corpus behind its answers "must be global, fresh, honor publisher preferences by default, and continuously evolving," and that its grounding quality metric, GDSAT, "captures whether the grounding truly meets user intent across completeness, freshness, and authority." A companion Microsoft engineering post explains why staleness is treated as a defect: "missing or stale context propagates into reasoning errors rather than degrading gracefully." The other products publish nothing comparable.

Notice what is actually established here. The claim-level problem, stale facts producing wrong answers, is measured and documented. The page-level pattern, that recently dated pages get cited more, is industry correlation: Ahrefs, for example, reported that pages cited by AI assistants were on average about 25.7% fresher than ordinary top-ranking pages across roughly 17 million citations, a correlation that does not show date stamps cause citations. The same analysis found the pattern differs by product rather than pointing one way: ChatGPT cited the newest pages, Google's AI Overviews the oldest, with Gemini, Copilot and Perplexity in between. That is a measured spread in the age of cited pages, not evidence that any engine weights recency more heavily. And the GEO paper, the controlled benchmark that tested nine content optimizations for generative engines, includes no freshness or recency method among them. Neither does the later AutoGEO work. Four practices follow from that split: keep your claims true, show one clear date, maintain pages whose subjects move, and never fake the timestamp.

Keep your claims true

This is the half with real measurement behind it. The FreshLLMs paper (Vu et al., 2024) built FreshQA, a benchmark that sorts factual questions by how their answers age: "never-changing, in which the answer almost never changes; slow-changing, in which the answer typically changes over the course of several years; fast-changing, in which the answer typically changes within a year or less; and false-premise". The useful move is to apply that taxonomy to your own pages. A paragraph about the boiling point of water never expires. A paragraph about a framework's current major version can expire within months.

Two failure modes are worth telling apart. A stale claim is a specific that is no longer current: "the current version is 14" when the current version is 17, "as of 2021 the limit is 10 MB," a statistic presented as today's number but anchored to an old survey. These are visibly wrong to anyone who checks, and an engine that retrieves your page alongside a fresher source has every reason to prefer the one whose specifics hold up. An aging claim is different: relative wording with no date anchor, recently, currently, "just launched." It is correct on the day you publish it and rots silently, because nothing in the sentence says when it was true.

The fixes differ accordingly. Wrong specifics need correcting, since a page caught carrying one gives a retrieval system a concrete reason to route around it. Relative wording is not wrong yet, and no engine documents a penalty for it; pinning it to an explicit date ("as of March 2026") is simply good style that keeps the sentence true no matter when it is read.

Show a clear, consistent date

The second practice is not "have a date," it is "have a date the machine can trust." A date on a web page lives in several places at once: the visible text a reader sees, the Open Graph article:published_time meta tag, and the JSON-LD datePublished and dateModified fields. The trustworthy configuration is all of them present and all of them agreeing. A visible published date, plus a "last updated" line where the page has actually been revised, corroborated by matching structured data, is the gold standard: a human can verify it, and a crawler can confirm the human-facing page says the same thing the markup does.

For dated editorial content, three configurations can undermine that trust. A page with no date at all reads as evergreen-of-unknown-age, which is fine for some pages and a real cost for editorial content, where an engine that wants to caption a citation with "Published" or "Updated" has nothing to work with. A markup-only date is a machine signal with zero human-facing reinforcement, exactly the shape a template default produces. And dates that contradict each other, a visible "January 2024" over markup that says 2022, look like a stale CMS field or a template bug, and a signal that disagrees with itself is one an engine has reason to discount entirely.

The scope here is honest: not every page needs a date. A product page or a landing page without one is normal, and adding a date to it solves nothing. This practice is for dated editorial content, articles, guides, comparisons, anything where "when was this written?" is a question a reader or an engine would reasonably ask.

Keep the page genuinely maintained

The third practice concerns the page's life after publication. Two dates tell that story: if dateModified meaningfully postdates datePublished, the page has a visible maintenance history. If the two are identical, or dateModified is missing, then a page you revisit every quarter and a page you abandoned in 2022 look exactly the same to a crawler. Nothing distinguishes "still true because we checked" from "untouched since launch."

Let the subject's rate of change determine the maintenance schedule. For time-sensitive material, pricing, product comparisons, "best X" roundups, anything news-adjacent, keeping current genuinely matters. For evergreen reference content, an old page that is still accurate is fine. A 2023 "best tools of the year" page has expired by design; a 2023 explainer of a stable protocol has not, and a crawler cannot tell the two apart from the date alone. The right cadence is a periodic review that validates the facts, updates what changed, and leaves alone what did not. Rewriting a page whose facts have not moved, just to advance the date, is motion rather than value.

When you do make a real update, several signals carry it, and they should agree with each other and with reality: a visible "last updated" line on the page, dateModified in the JSON-LD, lastmod in the XML sitemap, the Last-Modified HTTP header, and any year references inside the prose itself. A page that says "updated 2026" above text that discusses "the upcoming 2024 release" contradicts itself in exactly the way the previous section warns about.

Do not fake the date

The tempting shortcut is obvious: bump dateModified, or add an "Updated 2026" line, and change nothing else. It does not hold up. Crawlers fetch pages repeatedly, and a date that advances while the content between crawls stays byte-identical is a detectable pattern; an engine comparing snapshots can discount the new date because nothing corroborates it.

The research adds a sharper point. A controlled study presented at SIGIR-AP 2025, "Do Large Language Models Favor Recent Content?" (Fang et al.), found that prepending artificial publication dates to otherwise unchanged passages shifted how language models ranked them, promoting the "fresh"-looking text. Read carefully, that is not evidence that fresh content is better. It is evidence that the recency preference inside these models is a shallow bias operating on surface cues, one that can be triggered by a fake date as easily as a real one. A signal that gameable is a fragile thing to build on. The study's own conclusion is that the bias needs mitigating, which is a reason to expect the behavior to change, not a promise that it has.

The durable version of the same instinct is a real update: statistics refreshed against current sources, recommendations revised where the landscape moved, wrong specifics corrected, sections that no longer earn their place removed. Then the date advances because the content did, and every signal from the visible line to the sitemap confirms that the page is genuinely fresher.

Frequently asked questions

Does updating a page's date help it get cited by AI?

Only when the query deserves recency. For evergreen questions, no vendor documents what takes its place. The controlled GEO experiments did not test freshness at all, so they establish nothing either way, and a date change alone guarantees nothing.

How often should I update my content for AI?

As often as the facts actually change, not on a fixed calendar. Time-sensitive pages such as pricing, comparisons and news benefit from frequent real updates; an accurate evergreen reference can sit untouched for a long time.

Is it enough to change the dateModified timestamp?

No. Any system that keeps page snapshots could spot a date that moved while the content did not, though no engine documents doing so. The stronger reason to avoid it is that the date would simply be false. One study found fake dates can shift model rankings, but that reflects a shallow, gameable bias rather than a real gain. Make a meaningful update first, then let the date record it.

Where should the date go, on the page or in the markup?

Both, and they must agree. A visible published date a reader can see, corroborated by Open Graph and JSON-LD values that match it, is the trustworthy signal.

My content is evergreen. Does freshness still matter?

Less, but accuracy always does. An old reference page that is still correct can keep earning citations. What hurts is a stale specific: an outdated version number, price, or "current" statistic. Validate the facts periodically rather than resetting the date on an unchanged page.

See if AI can read, trust, and cite your site

Add to Chrome

Free · No signup · Every issue links back to a guide like this one