For years, one of the more sensible rules of enterprise data management has been fairly simple. If information is useful, keep it accessible. If it might be needed later, archive it. If there’s no longer a good reason to have it, delete it. AI is making the middle of that equation much harder to judge. 

Information that hasn’t been touched in years can suddenly have a new use. An old technical document might give a retrieval-augmented generation (RAG) system the context it needs to answer a question. Historical customer records can support long-term analysis. Past reports, conversations and project files may contain knowledge that was difficult to find when humans had to search for it manually. 

em360tech image

Organisations are already responding. IDC research sponsored by Western Digital found that 74.3% of 763 surveyed IT and business decision-makers are retaining data for longer because of AI and generative AI. Another 75.9% are bringing increasing amounts of archived data back online to support AI workloads. 

At the same time, enterprises aren't starting with neat, carefully curated archives. CTERA analysed 16 petabytes of live production data across 856 enterprise file-share scans and found that only 4.4% of stored capacity was actively used. More than 83% hadn't been modified in over a year. Put those findings together and an uncomfortable question appears. 

What happens when organisations already holding huge amounts of inactive data acquire a powerful new reason not to delete any of it? The answer can't simply be to keep everything in case AI finds a use for it someday. AI hasn't made all old data valuable. It's made deciding which old data still has value considerably harder.

AI Is Changing The Enterprise Data Lifecycle

Traditional data lifecycle management has always involved trade-offs. Some information needs to remain readily available because people and applications use it every day. Some has to be retained for legal or regulatory reasons. Other information can move into cheaper archive storage, while data that has reached the end of its useful life can eventually be deleted. 

AI adds something new to that calculation: a use that might not exist yet. The significance becomes clearer when we look at what organisations are doing with their archives. The Western Digital-sponsored IDC research found that 96% of respondents expect they'll need faster archive retrieval to support AI inference and RAG. 

The traditional distinction between active and archived information is already becoming less clear. RAG is one reason why. Rather than expecting an AI model to know everything itself, RAG allows it to retrieve information from external sources when answering a question. That makes documents, reports and other enterprise information potentially useful even if nobody has opened them recently. 

The same principle extends beyond RAG. Historical information can support analytics, help organisations understand how something changed over time or provide context for AI applications that haven't been built yet. That changes what inactivity tells us. A file can stop being useful to today's employees without losing all possible business value. 

The point at which information leaves everyday use and the point at which it stops being worth keeping may no longer be the same.

Inactive Data And Valuable Data Aren't Opposites

There's an obvious response to that uncertainty: keep more. Except that enterprises already have a lot of information they're doing very little with. CTERA found that 97.5% of the individual files in the environments it analysed hadn't been modified for more than a year, while 88% hadn't even been accessed during that period. 

Some of those files may be valuable precisely because they're old. Historical contracts, engineering records, research, incident reports and other documents can preserve knowledge that can't simply be recreated later. Lack of activity doesn't tell us whether that knowledge will eventually be useful. But it doesn't tell us the opposite either. 

A ten-year-old document might contain unique institutional knowledge. It could also be a duplicate, an abandoned draft or a record whose value disappeared years ago. Both files look inactive when measured by access frequency. Their value is completely different. This is where AI complicates enterprise data retention without providing an easy replacement for the old rules. 

Recent activity is an imperfect measure of value, but "AI might need it someday" isn't a meaningful data retention strategy either. The better question is what knowledge the information contains, whether that knowledge is unique and whether there's a credible reason the organisation may need it again. 

And even then, keeping something doesn't necessarily mean an AI system should be able to retrieve it.

More Data Can Make AI Less Useful

It's easy to imagine an AI knowledge base improving as more information is added to it. More documents mean more knowledge, which should mean better answers. The problem is that enterprise information doesn't accumulate in such an orderly way. Policies change. Product specifications are updated. Processes are replaced. 

Someone creates a new version of a document without deleting the old one. Another team saves its own copy. Years later, all of those files can still contain information that looks relevant to the same question. Research published at the 2025 Annual Meeting of the Association for Computational Linguistics tested what happens when outdated and current information coexist in RAG knowledge bases. 

The researchers found that outdated information could reduce response accuracy by distracting models from correct information. It could also lead to potentially harmful outputs even when current information was available. That creates an important distinction for enterprise data strategy. 

Historical information can be worth preserving without being appropriate for unrestricted AI retrieval. A previous policy may be useful when someone wants to know what the rules were five years ago. It becomes much less useful when an employee asks what they're allowed to do today and an AI system can't reliably distinguish the old version from the current one. 

So the question isn't only whether information deserves to survive. Enterprises also need to know whether AI should be able to use it, for which purposes and with enough context to understand what it's looking at.

Retention Still Comes With A Cost

Potential future value can be difficult to measure. The costs and risks of keeping data are usually much more immediate. Storage is one part of that equation. The same IDC research found that 98.2% of respondents consider total cost of ownership per terabyte important or very important when making storage decisions. 

But retained information also has to be managed, protected and governed for as long as the organisation keeps it. Security adds another concern. An old file may no longer have operational value, but it can still contain sensitive information. 

CTERA's analysis warns that inactive data can remain in network-accessible environments under the same permissions as active information, leaving more data available than employees and applications need for everyday work. Then there's privacy. 

The UK's Information Commissioner's Office expects organisations using personal information in AI systems to review whether the data they're retaining remains relevant and justified. Its current AI audit guidance also calls for documented retention schedules and the deletion of personal information when it's no longer required. 

That creates a limit around speculative AI value. An organisation can't assume that a possible future AI application automatically justifies retaining personal information indefinitely. The real cost of retention therefore goes well beyond capacity. Enterprises are accepting an ongoing management, security, privacy and governance burden in exchange for whatever future value they believe the information may provide. 

That makes retention a decision about individual data assets rather than a blanket policy for everything an organisation owns.

The Better Question Is What The Data Deserves

Once AI makes inactivity a weaker indicator of value, enterprise data retention needs a more deliberate middle ground. The choice isn't simply between deleting information and keeping everything readily available. 

A more useful approach is to decide what each category of information deserves based on what the organisation actually knows about its value, usability and risk.

Keep data with continuing value

Some information has a clear reason to remain. It may still support the business, contain unique institutional knowledge, preserve useful historical context or need to be retained for legal and regulatory purposes. The important distinction is that its value can be explained. 

The organisation isn't keeping it because somebody imagines AI could eventually find something interesting. There's a defensible reason for retaining it and an appropriate way to govern it.

Archive data that may matter later

Other information may have plausible future value without needing to sit alongside everyday operational data. Archiving allows enterprises to preserve that knowledge while recognising that possible future usefulness doesn't justify treating it as active information. 

The archive still needs to be discoverable and managed properly, particularly as AI makes historical information easier to reuse. But discoverable doesn't have to mean universally available.

Enrich data before AI can use it

Are you enjoying the content so far?

Some information sits somewhere between valuable and usable. The underlying knowledge may be worth keeping, but an AI system needs more context before it can use that knowledge reliably. Metadata can help explain where information came from, when it applied, who owns it and whether a newer version exists. 

Provenance provides evidence about its origin and history. Together, those details can help machines distinguish a useful historical record from something they should treat as current. This creates another option beyond keeping or deleting data: retain the knowledge, but improve its context before allowing AI to depend on it.

Delete data whose cost and risk outweigh its value

AI shouldn't eliminate deletion from the enterprise data lifecycle. Some information is duplicated, obsolete, legally unnecessary or simply no longer useful enough to justify the cost and risk of keeping it. In those cases, the possibility that an unknown future AI application could theoretically use it isn't much of a business case. 

Defensible deletion still requires organisations to understand what they're removing and why. The difference is that AI becomes another consideration in that decision, rather than a reason to avoid making it. The goal isn't maximum retention. It's being able to explain why different information deserves different treatment.

AI Access Should Become A Separate Decision

That still leaves one more question. Once an organisation decides information is worth retaining, what should its AI systems actually be allowed to see? The answer doesn't have to be everything. An enterprise data estate can contain information that remains valuable for legal, historical, analytical or operational reasons without placing all of it inside the same AI knowledge base. 

Access can depend on the application, the user, the sensitivity of the information and whether historical or current knowledge is required. We're already seeing infrastructure move in this direction. 

In June 2026, Rubrik introduced an approach that scans and catalogues distributed unstructured data where it already lives, then allows organisations to identify specific subsets for downstream AI workloads rather than copying the entire estate into another environment. It's one vendor's architecture rather than an industry standard, but the principle is useful. 

It separates knowing what data exists from giving AI direct access to all of it. That distinction could become increasingly important as archived information returns to use. Organisations may need historical knowledge to remain discoverable without treating every old document as equally trustworthy, relevant or appropriate for every AI application. 

Retention can therefore become one decision, discoverability another and AI access a third. Separating them gives enterprises more room to preserve potentially useful knowledge without turning their entire history into one enormous retrieval pool.

Final Thoughts: AI Needs Better Retention Decisions, Not More Data

Enterprise data retention used to become easier as information aged. Usage fell, business relevance became clearer and organisations could gradually decide what belonged in an archive and what no longer deserved to exist. AI has disturbed that neat progression. Historical information can become useful again because machines can find, combine and work with knowledge that people rarely touched. 

But the possibility of future use doesn't turn every forgotten file into an asset. The better enterprise data strategy isn't to keep as much information as possible. It's to understand what deserves to remain, what needs more context, what belongs in an archive, what should disappear and which parts of that retained history AI should actually be allowed to use. 

As AI capabilities continue changing what organisations can do with their existing information, those distinctions are likely to become more important, not less. EM360Tech will continue following how AI is reshaping the data architectures, governance decisions and infrastructure choices enterprises build around that information.