Skip to main content
geo

AI Has a Long Memory for Everything Except Your Own Content

Your pages decay in weeks while forum threads about you persist for years. Why freshness rules apply asymmetrically, and where the budget usually goes wrong.

ยท By Veljko Plavsic ยท 6 min read

Semrush's analysis of 248,000 Reddit posts found the median post cited by AI engines is roughly 900 days old. Profound found 4% of cited posts date from 2019 or earlier. Meanwhile, the guidance for your own pages is to refresh them every two months.

Both of those things are true, and almost nobody puts them next to each other.

Freshness rules apply asymmetrically. Your content decays in weeks. What other people wrote about you persists for years, and it is frequently the thing answering the question.

Key Takeaways

  • Freshness cuts one way: your pages need constant updating; third-party content about you does not expire.
  • The median cited forum post is years old: roughly 900 days, per a 248,000-post analysis.
  • Complaints are cited slightly more than praise: 6.1% against 5% in one measurement.
  • Most budget targets the wrong side: refresh cycles improve the input with less influence.
  • Old sources cannot be refreshed, only diluted or corrected: which is a different discipline entirely.

Why Does AI Cite Old Content About Brands but Prefer Fresh Content From Them?

The two behaviours serve different purposes. When a model needs current facts, it favours recently updated sources. When it needs to characterise something, it draws on accumulated discussion, and accumulated discussion is by definition old.

A question like "what changed in this regulation" pulls recent pages. A question like "is this product any good" pulls whatever people have said over several years. Most brand-relevant queries fall into the second category.

What the measurements show

Profound's analysis found the average Reddit post cited by AI models in 2025 was written roughly a year earlier, with 4% dating from 2019 or before. Semrush's larger sample of 248,000 posts puts the median at around 900 days.

A complaint about a product version you discontinued two years ago can still be the current answer. There is no page two in an AI answer, so when a single old thread is the cited source, that thread is the entire response.

The Budget Consequence

Most GEO programmes spend heavily on refreshing owned content and almost nothing on the third-party sources that carry more weight and never expire on their own.

Refresh cycles are attractive because they are controllable and measurable. You can schedule them, assign them, and report on them. Auditing what a forum said about you in 2022 is none of those things.

The controllable work is not the influential work. Muck Rack's analysis of 25 million cited links found earned media accounts for 84% of AI citations, against 0.3% for paid placements. Owned refresh cycles improve a smaller share of the input.

This is the practical counterpart to what our AI Citation Freshness report measures on the owned side. Freshness governs your pages. It does not govern the archive.

AI Citation Freshness: The 30-Day and 13-Week Windows That Decide Who Gets Cited

Fresh vs Persistent: Two Different Problems

Your own contentThird-party content about you
Typical age when citedRecentOften 1-3 years
Decay speedFastVery slow
ControlDirectIndirect at best
FixUpdate the pageCorrect, respond, or dilute
Share of citationsSmallerLarger
Usual budget shareMostLittle

Why Negative Content Persists Particularly Well

Complaints are cited at a slightly higher rate than praise. One measurement put negative brand sentiment on Reddit at a 6.1% citation rate against 5% for positive sentiment.

The reason is structural rather than adversarial. Complaints tend to be specific. They name features, describe circumstances, and explain what went wrong. Positive comments are frequently short and generic. Specificity is what makes a passage useful to a retrieval system, so detailed criticism outcompetes brief approval.

This also explains why resolved issues resurface. A thread documenting a support failure from three years ago contains more citable specificity than a hundred satisfied customers who never wrote anything down.

How much this affects you depends on where your buyers ask. Community platforms weigh heavily on some engines and barely register on others, which is one more consequence of the fact that platforms cite largely different sources.

Which Domains AI Engines Trust Most: 86% of Top Sources Are Not Shared Across Platforms

What Actually Works on Old Sources

You cannot refresh someone else's post. Three approaches remain, in descending order of reliability.

Correction at the source

Where the content is factually wrong rather than merely unflattering, a factual correction added to the original thread or page becomes part of what models retrieve. A dated, specific reply stating what changed and when frequently outperforms attempting removal, which rarely succeeds and sometimes amplifies.

Displacement through specificity

Old content persists because it is specific. It gets displaced by content that is more specific and more recent about the same thing. Vague positive content does not compete with a detailed complaint; a detailed account of what changed does.

Dilution across the same sources

For content mentioning you in passing within a broader discussion, removal is not available, and correction may not be appropriate. Building genuine presence in the same communities over time shifts the balance of what is available to cite. This is slow, cannot be accelerated with spend, and is the only durable option in many cases.

How to Audit the Archive

Auditing what already exists is different from monitoring what appears next. Both are necessary; most teams only do the second.

  • Run your brand queries and read the sources, not just the answers. Note the publication date of every cited source. The distribution is usually surprising.
  • Flag anything describing a version of you that no longer exists. Discontinued products, old pricing, resolved service issues, former leadership.
  • Sort by fixability. Factually wrong content is correctable. Unflattering but accurate content is not, and needs a different response.
  • Check across platforms and locations. Source sets differ substantially, so an audit run from one engine in one country is incomplete.

This runs alongside rather than inside standard measurement, which is why tracking AI search visibility as a volume metric misses it entirely. Citation counts can hold perfectly steady while a five-year-old thread quietly defines you.

How to Track AI Search Visibility Across ChatGPT, Perplexity, Gemini, and Claude in 2026

Why This Gets Missed

Three reasons, all of them organisational rather than technical.

The work is unglamorous. Nobody presents a quarterly review slide about correcting a 2022 forum post.

It sits between functions. Owned content belongs to marketing, community threads belong to support or comms, and the archive belongs to nobody.

It produces no publishing metric. A refresh cycle generates output. An archive audit generates a list of things to leave alone, which is harder to justify and frequently the correct answer.

The Reframe

Freshness advice is not wrong. It is incomplete in a way that reliably misallocates budget.

Your pages need updating because models prefer current sources when they need current facts. But when someone asks whether you are any good, the model reaches for an archive you do not control and cannot refresh, where the most specific thing anyone ever wrote about you is likely to win. Knowing what is in that archive is cheaper than any content programme, and almost nobody has looked.


FAQs

1. How old is the content AI engines typically cite about brands?

Older than most people expect. Semrush's analysis of 248,000 Reddit posts found the median cited post is roughly 900 days old, and Profound found 4% of cited posts date from 2019 or earlier. Discussion-based sources are cited long after publication.

2. If freshness matters, why does AI cite years-old forum threads?

Because freshness preference depends on query type. Models favour recent sources for current facts and draw on accumulated discussion when characterising something. Questions about whether a product is good fall into the second category, where old content is the available evidence.

3. Why does negative content about my brand get cited more?

Complaints tend to be more specific than praise, naming features and describing circumstances in detail. Specificity makes a passage more useful to a retrieval system. One measurement put negative Reddit sentiment at a 6.1% citation rate against 5% for positive.

4. Can I remove old negative content that AI cites?

Rarely, and attempting removal often fails or amplifies the content. More reliable options are correcting factually wrong information at the source, publishing more specific and recent material on the same subject, and building genuine presence in the same communities over time.

5. Should I stop refreshing my own content?

No. Refresh cycles work for the queries where models prefer current sources. The point is that owned refresh is not sufficient on its own, because it does not touch the older third-party sources that carry a larger share of brand-characterising citations.


Figures reference published analyses by Semrush, Profound, and Muck Rack. Citation age and sentiment patterns vary by platform, sector, and query type, and shift as models are updated.

Updated on Aug 11, 2026