Enterprise data has a habit of multiplying. A business starts with information inside the systems that create it. Then analytics needs the same information somewhere else, so it gets extracted, transformed and loaded into a warehouse.
Another team needs a different view, another platform needs its own format, and before long the same underlying information exists in several places, each with its own pipeline and rules for keeping it current. AI is giving enterprises new reasons to do the same thing.
Data can be copied into specialist stores, turned into embeddings, indexed for search or moved through pipelines built specifically for AI applications and agents. Sometimes there’s a good technical reason for doing this. But adding AI to a workload doesn’t automatically mean it needs another copy of the data.
That assumption is becoming easier to challenge as AI data architecture changes. Databases are adding native vector capabilities, data platforms are making distributed information available without moving it first, and vendors are building ways for AI agents to work closer to the systems where enterprise data already lives.
So perhaps the better place to start isn’t asking where the data needs to move. It’s asking why it needs to move at all.
AI Is Creating A New Data Duplication Problem
Creating another copy of data rarely means creating only another copy. There’s usually a process behind it. Information has to be extracted from its source, transformed into the format the new system expects and kept up to date as the original changes. Someone also needs to manage access, monitor quality and know which version should be trusted when two systems disagree.
AI can add another layer to this already complicated picture. A retrieval-augmented generation (RAG) system, for example, may create embeddings, which are numerical representations of information that allow an AI system to search for content based on meaning rather than exact words.
Those embeddings might then live in a vector database or specialist index alongside another copy of the source content. None of this is necessarily a problem. The extra representation may be exactly what the workload needs. The problem begins when creating it becomes an automatic part of making data "AI-ready", without asking what additional value it provides.
The data estate many organisations are starting with is already difficult to understand. BARC's 2026 Harnessing Unstructured Data for AI Innovation study, based on 225 responses from data, analytics and AI leaders, found that only 29 per cent of respondents fully knew where their AI-relevant unstructured data was located.
Meanwhile, 70 per cent said less than half of their unstructured data was discoverable and usable for analytics or AI. Adding more representations doesn't fix that underlying problem. In some cases, it can give the organisation more places where data needs to be discovered, governed and kept consistent. This becomes harder to ignore as AI moves into production.
Informatica's CDO Insights 2026 study found that 57 per cent of the 600 data leaders surveyed saw data reliability as a key barrier to moving AI projects from pilot to production. Half also identified data quality as the leading challenge when deploying agentic AI.
Before adding another store or pipeline, then, there’s a more basic question to answer: what does the AI workload gain from having the data somewhere else?
Zero-Copy AI Doesn't Mean Nothing Gets Created
"Zero-copy AI" sounds like the obvious alternative. Leave the data where it is, connect the AI to it and the duplication problem disappears. The reality is more complicated. The term can describe several different approaches.
Inside Data-First AI Leadership
Stathopoulos links data architecture, AI readiness and workforce skills into one operating model for enterprise decision making.
NetApp, for example, is developing what it calls "in-place, zero-copy data activation", using metadata to help organisations discover, understand and govern information across distributed storage rather than treating AI readiness as a separate data-copy project. Google Cloud's borderless Lakehouse similarly allows analytics engines and AI agents to query data across clouds, operational systems and SaaS environments without first moving the physical files.
Elsewhere, the change is happening inside the database. Amazon DynamoDB introduced native vector search in August 2026, allowing organisations to store vector embeddings alongside operational data and search them without replicating that information into a separate vector database.
AWS specifically positions this as a way to remove the additional synchronisation pipeline previously required between the two systems. But notice what hasn't disappeared. DynamoDB still needs embeddings and a vector index to perform similarity searches. Google's architecture still needs shared catalogue information so different systems can understand distributed data.
Keeping authoritative data in place doesn't mean AI suddenly works without indexes, metadata, semantic models, caches or other derived forms of that information. That's why zero-copy AI is more useful as an architectural direction than a literal rule. The aim isn't necessarily to create nothing new.
It's to avoid duplicating authoritative data when a smaller or more specialised representation can provide what the workload actually needs. Once that distinction is clear, the question stops being whether copies are good or bad. It becomes much more practical.
When Does AI Actually Need Another Copy?
There isn't one correct architecture for enterprise AI because AI workloads don't all ask the same things of enterprise data. A customer-facing assistant retrieving product documentation has very different requirements from an autonomous agent making decisions against live inventory.
A semantic search service behaves differently from a model being trained against a fixed historical dataset. Even two applications using the same source information can need different levels of freshness, performance and isolation. Architecture should follow those requirements, rather than the presence of AI itself.
When another representation earns its place
Inside AI’s Hidden Plumbing
AI fails without resilient data pipelines, governance and infrastructure; this piece reframes AI as a business utility, not an IT project.
Performance is one of the clearest reasons to create something new. An operational database may be excellent at processing transactions but poorly suited to thousands of semantic searches from an AI application. Forcing the source system to handle that workload could slow down the application it was actually built to support.
Isolation can be just as important. An organisation may not want unpredictable AI traffic hitting a critical production system directly, even if doing so is technically possible. A separate store, cache or index can protect the source while giving the AI workload the performance it needs.
Some use cases also require a different representation of the information. Semantic search needs embeddings somewhere because it compares the mathematical meaning of content rather than searching only for matching words. Historical analysis may deliberately need a snapshot that doesn't change when the live system does.
Other workloads may require data to be cleaned, combined or transformed before an AI system can use it reliably. In each case, the additional representation has a job. That's the distinction worth preserving. Another copy earns its place when it solves a requirement that governed access to the authoritative source can't solve well enough.
When governed access is the better fit
The balance changes when an AI system mainly needs trusted access to information that already exists in a usable form. Freshness is an obvious example. If the authoritative data changes constantly, every additional copy creates a delay between something happening in the source and that change reaching the AI system.
The pipeline might update in seconds, minutes or hours, but it still has another state to synchronise. That becomes increasingly important when AI systems can act on the information they retrieve. An IDC InfoBrief commissioned by Solace surveyed 623 senior technology decision-makers at enterprises with more than $1 billion in revenue.
When Bad Data Powers Smart AI
Explore how biased, misgoverned datasets silently undermine AI decisions in hiring, health and finance, and what governance must add.
Organisations identified as leaders in real-time data maturity were more than three times as likely to report measurable results across the majority of their AI projects. There are other reasons to leave authoritative data where it is. Sovereignty requirements may restrict where information can move. Different business units may retain ownership of their own data.
A source platform may already provide secure, efficient access without requiring another physical store. In those cases, moving the data can create another synchronisation and governance problem without solving a meaningful workload problem. But leaving information distributed introduces a different challenge.
The AI still needs to know where the right data lives, what it means and whether it's allowed to use it.
Keeping Data In Place Shifts The Architecture Problem
An AI system can't make much use of data it can't find or understand. This is where distributed data architecture becomes less about putting everything in the same place and more about creating a consistent way to understand information across different places. Metadata can describe what a dataset contains and where it lives.
A semantic layer can give different systems a shared understanding of business concepts. Identity and policy controls can determine who or what is allowed to access the information. Lineage adds another piece by showing where data came from and how it changed along the way.
Together, these capabilities help create logical consistency without requiring every dataset to share the same physical home. There’s evidence that enterprises are already moving in this direction. Accenture surveyed executives at 2,000 companies across 15 countries and nine industries for its 2026 AI-ready data research.
It found that only six per cent currently had a unified logical view of data across systems and ecosystems. Accenture argues that wholesale centralisation in one cloud isn't realistic for many organisations and recommends a federated model, where data can remain distributed while common standards govern security, quality and access.
When AI Becomes A Profit Engine
Shift AI from pilots and activity metrics to portfolios that evidence margin, growth, risk reduction and operational resilience.
Gartner describes a related shift as moving from "data gravity" towards "semantic gravity". Rather than relying primarily on centralising data, its August 2026 research argues that AI and analytics increasingly depend on maintaining consistent meaning and governance across hybrid and multicloud environments.
But federation isn't an automatic shortcut to AI readiness either. BARC's Data Fabric Survey 26, based on 776 participants worldwide, found that 69 per cent of users reported high or very high benefits from data-management tools for accessibility. Yet 36 per cent saw little or no benefit from those tools when it came to preparing for AI.
The physical location of the data is only one part of the problem. Whether data moves or stays where it is, AI still needs reliable ways to discover it, understand it and use it within the organisation's rules.
Make Every AI Copy Prove Its Value
Once AI data architecture is treated as a workload decision rather than a fixed pattern, the conversation changes. Instead of beginning with "Where should we copy this data?", data leaders can start with what the workload actually needs from it. A few questions can expose the difference:
- Freshness: How current does the information need to be, and what happens if the AI works from an older version?
- Performance: Can the authoritative source handle the volume and type of queries the AI will generate?
- Transformation: Does the workload need embeddings, restructuring or another representation that the source can't provide directly?
- Isolation: Would direct AI access put unacceptable pressure or risk on a production system?
- Authority: If another copy exists, which system wins when the two disagree?
- Governance: Can permissions, classifications and lineage be maintained when the data moves?
- Lifecycle: Who owns synchronisation, updates and deletion across every representation?
- Value: What capability does creating another copy provide that the existing architecture can't?
The answers won't always point towards keeping data in place. Nor should they. A vector database, materialised dataset, cache or specialist store can be the right architecture when its benefits justify the extra state the enterprise has to manage. The important change is making that justification explicit.
Every new representation brings something with it, whether that's better retrieval, faster performance or safer isolation. It also brings responsibility for keeping that representation accurate, secure and useful for as long as it exists.
Final Thoughts: AI Shouldn't Make Data Movement The Default Again
Enterprises have spent years learning what happens when the same information spreads across systems faster than the organisation can keep track of it. AI doesn't need to repeat that pattern simply because the new copies have different names. Embeddings, vector indexes, caches and specialist stores aren't unnecessary by definition.
For some AI workloads, they're exactly what makes the application practical. But as databases, lakehouses and data-management platforms become better at providing governed access to information where it already lives, moving or duplicating the underlying data is becoming one architectural option rather than the automatic starting point.
The stronger question is also the simpler one: what problem will another copy solve? If the answer is performance, isolation, transformation or another clear workload requirement, the additional complexity may be worth carrying. If the answer is simply that this is how the AI architecture was designed, it may be time to look at the data again.
As AI becomes a routine consumer of enterprise information, good data architecture will depend less on how much data organisations can move towards it and more on giving each workload what it needs without creating complexity it doesn't.
As AI continues to change how enterprises use and access their data, EM360Tech will keep following the architectural decisions shaping what comes next.
Comments ( 0 )