How often should I update my content for AI search? The 25.7% freshness number everyone quotes comes from one 2025 study — and the same study says cited pages average 2.9 years old
There is a genre of post that answers this question with a table: product pages weekly, blog posts quarterly, everything annually. The tables disagree with each other, none of them cite a source for the cadence, and several state as fact that freshness is a "top three ranking factor for AI search" — a phrase which, as far as we can determine, nobody has published a method for.
Underneath the noise there are two real studies. This post is about what they actually say, what they cannot say, and what you can do in the meantime that does not depend on either of them being right.
Grading
| Mark | Meaning |
|---|---|
| ◐ | Our own first-party measurement. One brand, one category, one week, no control group, and we sell a product in this space. |
| △▲ | Observational study, no control group, published by a company that sells a related product |
| ○▲ | Preprint with a statistical result, author affiliated with a vendor in this space |
| ✗ | Quoted widely, no method we could locate |
1. The number everyone is quoting, traced
The figure that has propagated furthest is some version of "AI search relies more on recent content than Google does — by about a quarter."
It is traceable. Ahrefs, by Ryan Law and Xibeijia Guan, published 28 July 2025. Method as stated: 16.975 million cited URLs collected via their Brand Radar product from ChatGPT, Perplexity, Gemini, Copilot and AI Overviews, plus organic Google SERPs for comparison, with both publication date and last-updated date recorded. Grade △▲ — observational, no control, and Ahrefs sells the tool the data came from.
The headline:
The average age of URLs cited by AI assistants is 1064 days, compared to 1432 days for URLs in organic SERPs — 25.7% "fresher."
That is a real comparison with a stated denominator, and it is the strongest thing in this area. Note precisely what it is: a difference between two average ages. It is not a measurement of what happens to a page when you update it.
⚠ If you have seen this quoted as "AI search relies 26% more on recent information" — that is a paraphrase, and the paraphrase changes the claim. "The pages it cites are on average 25.7% younger" and "it depends 26% more on recency" are different sentences. The second implies a weighting inside a system nobody outside Google or OpenAI can see.
2. The same post contains the correction
This is why it is worth reading rather than quoting. The authors state, in their own words:
- "The average age of cited pages is still 2.9 years." Like traditional search, AI assistants still prefer citing long-lived content.
- "Google is the least influenced by content freshness" — and Google is still where the majority of searches happen.
- "Content freshness is one among many factors. Low-quality, irrelevant content that's updated every day will not have a magic positive effect."
- They explicitly warn against updating publish dates without changing the content, citing Google's John Mueller.
So the honest one-sentence summary of the largest study in this area is: pages cited by AI assistants skew younger than organic results, by about four months on a base of nearly three years. That is a real effect. It is not a mandate to republish everything every 30 days.
3. A second study, different method, compatible picture
Seer Interactive, published 2026, grade △▲ (a marketing agency that sells work in this area). Method as stated: 7,683 pages carrying 47,097 citations, March–June 2026, across ChatGPT, Gemini and Perplexity, in four industries — pet retail, vacation rentals, retail energy and commercial banking. Two-thirds of pages had readable last-modified dates from schema, sitemaps and headers. Stated limitation: the verticals are not the same size, so they report percentages rather than raw counts.
| Finding (Seer △▲) | |
|---|---|
| Cited pages updated within the past year | 75% |
| Updated within two years | 88% |
| Gemini / ChatGPT / Perplexity, updated ≤ 1 year | 78% / 73% / 65% |
| Pages with both dates: fresh by update date | 72% |
| The same pages: fresh by publish date | 42% |
| Share of "fresh" content that is refreshed older pages | 27–28% |
★ The last three rows are the most useful lines published on this topic. The gap between 72% and 42% is the entire practical question. If you look at publish dates, most cited content looks old. If you look at update dates, most of it looks new. These are the same pages.
And it points at something the cadence tables get right for the wrong reason: roughly a quarter to a third of everything that looks "fresh" in the citation set is old content that was updated, not new content that was written. Refreshing is not a lesser substitute for publishing in this data. It is most of the freshness.
⚠ Also note the engine spread — 78% to 65% — which is the same shape as everything else in AI search. The surfaces are not one thing.
4. The third study, and the most methodologically serious of the three
Scrunch AI with Stacker, published March 2026. Method as reported: over 3 million citation events, eight industries, six AI platforms, tracked across 26 weeks, analysed with cohort-based survival analysis and 200 bootstrap simulations. Headline: the median citation half-life for non-distributed, single-domain content is 4.5 weeks. Grade △▲ — Scrunch sells AI visibility software and Stacker sells content distribution, and the finding is flattering to precisely what Stacker sells.
★ Survival analysis is the right instrument for this question and, as far as we can tell, nobody else in this area is using it. Average age — what the other two studies report — describes the stock of cited pages at a moment. A half-life describes the rate at which an individual page falls out, which is the thing a refresh cadence is actually trying to counter. Those are different questions and only the second one has an operational answer in it.
⚠ We have not opened the original. We read it through secondary write-ups and are reporting the method as they describe it. Until we have read the primary we would not put weight on 4.5 weeks as a value — only on the observation that somebody has finally measured decay rather than stock.
4-A. And then there is the 13-week number
You will meet the claim that 50% of AI citations come from content under 13 weeks old. It is the single most repeated figure on this topic and the clearest illustration of why we grade sources at all.
Tracing it, we found it credited to four different origins by different republishers — Seer Interactive, Ahrefs, Scrunch/Stacker, and an analysis by Lily Ray presented at Tech SEO Connect 2026.
🔴 We read the Ahrefs post directly and it contains no 13-week figure. It reports average ages of 1064 and 1432 days. So at least one of those four attributions is wrong, and a reader has no way of telling which one without repeating the work we just did.
This is not us calling the number false. It may be perfectly good, and one of those four may well have produced it. The problem is structural: a figure whose provenance changes depending on who is repeating it cannot be checked, which means it cannot support a decision you will be held to.
How we handle this internally: a claim gets a source, a date, a method and a conflict mark before it can be used in a recommendation. If it cannot carry all four, it does not get to influence what we tell anyone to do. That rule is the reason this section exists rather than a tidier one.
4-B. Still untraced after looking
| Claim | Status |
|---|---|
| "Content has a 1-year half-life; each year of age cuts visibility 40–60%" | ✗ — no method located, and note how awkwardly it sits beside Scrunch's 4.5 weeks |
| "Freshness is a top-3 ranking factor for AI search" | ✗ — a statement about internal system weights. We do not see how it could be measured from outside at all |
| "76.4% of ChatGPT's top-cited pages were updated in the last 30 days" | ✗ — no denominator located |
| "Queries of 8+ words trigger AI Overviews 7x more" | ✗ — and worse: the same number appears stated two incompatible ways by different republishers, once as growth over time and once as a likelihood ratio |
5. 🔴 The confound none of them control for
Here is the objection we would raise if any of these studies were submitted to us, and it is the reason we are not treating the 25.7% as an answer.
AI citations and organic results are not drawn from the same population of page types. So comparing their average ages compares two mixtures, and a difference in the mixtures produces a difference in average age with no freshness preference existing anywhere in the system.
Our own citation data makes this concrete rather than theoretical ◐ — one brand, one category, July 2026:
| Where the citations went ◐ | Google AI Overviews | ChatGPT (head-term set) | Gemini (head-term set) |
|---|---|---|---|
| Vendor product / about pages | 52.1% | 66% | 23% |
| Vendor-published "best tools" listicles | (included above) | 3% | 38% |
| YouTube | 26.6% | 0% | 26% |
| Review sites | — | 19% | 0% |
| Community | 8.5% | — | — |
Now consider how age behaves per genre, before anyone optimises anything:
- A "best tools 2026" listicle carries the year in its title and gets re-dated annually as a matter of routine. It is fresh by construction.
- A YouTube page has a publish date and effectively never gets updated — but the ones that surface for a live query tend to be recent uploads.
- A vendor product page may be edited constantly and carry no visible date at all.
- A community thread is dated by its newest reply.
If AI citations over-select listicles and video relative to organic results, the citation set will look younger than the organic set even if freshness is not a factor at all. Sorting by age would recover the genre mix, not a preference.
🔴 We have not proven this confound is what is happening. To prove it you would have to hold genre constant and compare ages within it — compare listicles to listicles, product pages to product pages. Neither published study reports doing that, and we have not done it either. What we can say is that we hold data showing the genre mixes differ sharply by surface, which is the precondition for the confound to bite.
6. Why a naive before-and-after will not settle it either
Suppose you skip the studies and just test it: update thirty pages, re-measure, see if citations rise.
That design will produce a number, and the number will be close to meaningless, because the citation set churns on its own. Published measurement puts the overlap between consecutive days' cited sources at a Jaccard index of roughly 0.336–0.423 — around 65% of cited sources change overnight without anybody touching anything (Schulte et al. 2026, arXiv 2604.07585, grade ○▲, authors affiliated with a vendor in this space). At a single sample, the 95% interval on a brand-detection rate is about ±72 percentage points.
So a before-and-after on a handful of pages, measured once at each end, is drawing two numbers from a distribution wide enough to contain almost any story you want to tell.
7. The test that would actually settle it
We would want to see this, and would run it if we had the pages:
- Matched pairs, not a single group. Pair pages by genre, topic and current appearance rate. Update one of each pair; leave the other alone. The untouched half is the whole experiment — it is what absorbs the overnight churn from section 6.
- Repeated sampling at both ends. A rate per question per engine, from multiple samples, never a single observation. Never pooled across engines, given the 78/73/65 spread in Seer's data and the far wider spread in ours.
- Substantive updates only, and log what you changed. Seer's 72%-vs-42% gap means "date changed" and "content changed" are separable in the data. If your treatment is a date bump, you have tested date bumps, which is the thing Ahrefs' authors explicitly warn against.
- Hold genre constant. Do not update your listicles and compare against untouched product pages. That reproduces the section 5 confound inside your own test.
- Fix the observation window in advance. Nobody has published how long an intervention takes to reach these surfaces. We do not know either. Choosing the window after seeing the data is how a null result becomes a positive one.
- Publish the null result if you get one. This entire field currently has a publication filter on it: the studies that exist were run by companies that sell adjacent products, and a null result is worth nothing to them commercially.
8. What we have not done, stated plainly
We would rather be the ones to say this than have you notice it.
- We have not run the section 7 test. We do not currently have enough published pages of matched genre to pair.
- We cannot compare our own historical numbers. Our earlier measurements of our own exposure were taken with retrieval switched off, which measures what a model remembers rather than what AI search retrieves. Those figures are not comparable to anything we measure now, so we discarded the series and restarted rather than quietly continuing it. We are at the start of a baseline, not in the middle of a trend.
- We have no periodic runs yet. A freshness study is a time-series study by definition. Without a scheduler there is no series.
- We have not held genre constant in our own citation data. Section 5 is a stated confound, not a result.
9. So: a cadence you can defend
Given all of the above, here is the answer we would actually give, and note that none of it depends on the freshness studies being right:
Update a page when a claim on it has expired, not when a calendar says so. Every page has a decay rate you already know:
| Page contains | Decays |
|---|---|
| Pricing, plan limits, model or version numbers, integration lists | Fast — weeks to months. These go wrong on their own |
| Competitor names, category rosters, "best tools" line-ups | Fast, and visibly. A 2024 roster reads as abandoned |
| A method, a measurement, a result with conditions attached | Slowly. Ahrefs' own finding is that cited pages average 2.9 years old |
| Definitions, mechanisms, how something works | Barely |
This gives you a cadence without a study, and it survives either outcome of the section 7 test:
- If freshness is causal, you have updated the pages that had reason to change, which is what the Ahrefs authors describe as the version that works.
- If the 25.7% is a genre artefact, you have still corrected content that had become wrong, which was worth doing regardless.
⚠ What we would not do: bulk date-bumping, scheduled republishing with no substantive edit, or rewriting an evergreen explainer on a quarterly timer because a table on a vendor blog said to. The one thing multiple sources here agree on — including the authors of the largest study — is that changing the date without changing the page is the version that does not work.
Method
External figures are attributed inline with publisher, date, method as stated by the publisher, and a conflict-of-interest mark. We read the Ahrefs and Seer Interactive write-ups directly; we did not receive or inspect either underlying dataset, and neither is public as far as we can tell. Figures marked ✗ are ones we searched for and could not trace to any stated method — we have listed rather than used them.
Figures marked ◐ come from our own stored measurements on one brand in the AI video-editing category, July 2026: Google AI Overviews retrieved through a SERP API across 13 queries, and engine responses retrieved through the OpenAI and Google APIs with retrieval enabled, producing 154 citations in the head-term query set. Raw responses retained. We do not crawl third-party sites.
I build peekr. The reason section 8 exists is that a company selling measurement should be the first to say which measurements it does not have.
If you have run a matched-pair freshness test — in either direction — I would like to read it, and will link it here.