기사

The Data Foundation: Why AI at Scale Fails Without It

See why banks that can't fix their fragmented data foundation will always struggle to scale their AI programs beyond the initial pilot stage.

Simon Axon
Simon Axon
2026년 9월 3일 6 최소 읽기

Data fragmentation doesn’t cause AI to fail in banking. It causes AI to fail at the point where it appears to be working.

For years, fragmented data estates were tolerated. Various systems held different versions of the same information, definitions shifted by function, and duplication was common. It wasn’t elegant, but it was workable. Traditional analytics could absorb a degree of inconsistency. But AI is far less forgiving.

As institutions move beyond pilots, the data environment stops sitting quietly in the background and starts shaping outcomes. Models that have appeared stable begin to produce different results depending on where and how they are applied. Costs rise in ways that are difficult to forecast. Governance efforts grow faster than delivery.

At that point, what looked like an AI scaling issue reveals itself as something more fundamental. The data foundation is not built to support what the organization is asking of it.

Fragmentation only becomes visible when you try to scale

In most large banks, the data estate reflects years of layered investment. Core platforms, risk systems, finance warehouses, and newer cloud environments all coexist, each designed for a specific purpose. The complexity this creates is easy to underestimate because, for a long time, it doesn’t get in the way. In early AI work, teams can contain that complexity. They work with curated datasets, align definitions upfront, and limit integration points. Under those conditions, models perform well and confidence grows. That containment does not hold once scale begins.

As soon as a model is connected to multiple systems, it starts pulling from data sources and types that were never designed to work together. Customer information comes from one platform, transactions from another, risk indicators from a third. The inconsistencies that were previously manageable begin to surface in the outputs.

Take income as an example. One system records it as gross, another as net, and a third derives it through estimation. Each version makes sense in isolation. When combined, they introduce ambiguity that feeds directly into model behavior.

Across a single use case, that might be manageable. Across dozens of models spanning credit, fraud, servicing, and regulatory processes, it becomes a question of decision integrity. Fragmentation rarely causes a visible failure. It shows up more subtly, as drift. Over time, that drift erodes trust in the outputs, and once confidence is lost, scaling becomes much harder.

Every data copy creates a new version of the truth

Copying data is often treated as a practical way to move quickly. It solves an immediate problem, so it becomes a habit.

The longer-term effect is less benign. Each copy carries its own transformation logic. Definitions are adjusted to suit local needs, business rules are embedded, and over time, those versions begin to diverge. What started as a shared dataset becomes a collection of similar but not identical ones.

At AI scale, that divergence becomes difficult to manage. Governance teams spend more time tracing lineage and reconciling outputs. Engineering effort shifts toward maintaining consistency across environments. Costs increase as data is repeatedly stored and moved.

More importantly, control, to a certain extent, goes out the window. When there is no clear point of truth, answering a simple question about how a decision was made can require navigating multiple systems and interpretations.

AI systems amplify this because they do not respect organizational boundaries. They pull from wherever data is available. If consistency is not designed into the data layer, it will not appear in the outcomes.

What feels like a shortcut early on tends to create a slower, more complex path to scale.

Lineage breaks under pressure

At production scale, explainability moves from a compliance requirement to an operational necessity.

Many banks still rely on reconstructing lineage after the fact. Data flows are documented manually, transformations are pieced together, and control frameworks sit alongside the data rather than within it. This can work at low volume.

But this kind of manual effort doesn’t hold when AI is operating across millions of decisions.

The volume of lineage information grows quickly, and manual approaches struggle to keep pace. Gaps appear, investigations take longer, and confidence in the system begins to weaken. When something goes wrong, the ability to trace and respond becomes constrained.

This is where regulatory pressure and operational reality start to converge. Supervisors expect clear, timely explanations of model behavior, particularly in areas such as model risk and resilience. At the same time, the business needs to be able to act quickly when issues arise.

Both depend on the same thing.

Lineage must be captured as the data moves, not reconstructed later.

Governance overhead scales faster than models

One of the more predictable outcomes of scaling AI is the rate at which governance efforts grow.

Each new model introduces requirements around validation, approval, and documentation. As the number of models increases, so does the workload that surrounds them. In fragmented environments, this effort doesn’t scale cleanly.

The same data element may be reviewed multiple times in different contexts. Controls are repeated, documentation duplicated, and small variations in implementation lead to different interpretations of the same underlying data.

The effect is cumulative. Delivery slows, and friction builds between teams that are measured on different outcomes.

This is often described as a governance problem. More often, it reflects the structure of the data environment. When governance is applied on top of fragmentation, it becomes heavier with each new use case. When it is embedded into a consistent data foundation, it becomes part of how the system operates rather than an additional layer around it.

A “single source of truth” is not what most banks think it is

The idea of a “single source of truth” is widely used and rarely defined in practical terms. In reality, a trusted data foundation is less about centralization and more about consistency.

Three conditions tend to separate banks that can scale AI from those that cannot.

  1. Core data definitions are agreed once and reused. Concepts such as customer, transaction, and exposure mean the same thing across risk, finance, and commercial functions.
  2. Unnecessary duplication is reduced. Data is not copied into multiple environments unless there is a clear reason. Access is managed in a way that limits divergence.
  3. Governance is built into the data layer. Lineage, quality controls, and access policies are captured as data moves through the system, rather than applied afterward.

When these conditions are in place, the dynamic changes. New AI use cases no longer require rebuilding data pipelines from scratch, models can be trained and deployed against consistent inputs, and governance doesn’t need to be recreated each time.

Scale becomes a property of the environment, not a separate effort applied to each initiative.

What leaders should do differently

Most AI investment in banking still gravitates toward visible use cases. They are easier to justify and easier to measure.

The harder work sits beneath them. Leaders need a clear view of where data is being duplicated, where definitions diverge across the organization, and whether existing lineage approaches would hold under sustained demand. They also need to understand whether governance is being designed once or repeated across every model.

These questions are less visible, but they are what determines whether AI programs have the ability to truly scale.

This requires a shift in trade-offs. Investment in data architecture takes longer to show results and is harder to tie to a single outcome. But this is what allows multiple outcomes to build on one another.

There is also a discipline required to avoid layering new solutions onto an already fragmented estate. Managing fragmentation more efficiently is not the same as reducing it.

The foundation determines the outcome

As AI becomes more embedded in banking, the conversation is changing. The question is no longer whether models can perform, but whether institutions can support them in a consistent and controlled way. That capability sits in the data layer. Where the foundation is coherent, AI becomes easier to deploy, govern, and extend. Costs are easier to anticipate, and new use cases build on what already exists. Where fragmentation persists, each new initiative carries additional overhead, and the effort required to maintain control grows with it.

This difference is not obvious at pilot stage. It only becomes clear as institutions attempt to scale. At that point, the quality of the data foundation defines the boundary of what is possible.

For a deeper perspective on how leading banks are addressing these challenges, and what separates those that scale from those that stall, read the full whitepaper: From AI ambition to AI at scale.

Tags

약 Simon Axon

Simon’s primary focus is to help Teradata customers drive more business value from their data by understanding the impact of integrated data, advanced analytics and AI. With a background that includes leadership roles in Data Science, Business Analysis and Industry Consultancy across Europe, Middle East & Asia-Pacific, Simon applies his diverse experience to understand customers’ needs and identify opportunities to put data and analytics to work – achieving high-impact business outcomes.

Having worked for the Sainsbury’s Group and CACI Limited prior to joining Teradata in 2015, Simon is now the Global Financial Services Industry Strategist for Teradata.

모든 게시물 보기Simon Axon
알고 있어

테라데이트의 블로그를 구독하여 주간 통찰력을 얻을 수 있습니다



I consent that Teradata Corporation, as provider of this website, may occasionally send me Teradata Marketing Communications emails with information regarding products, data analytics, and event and webinar invitations. I understand that I may unsubscribe at any time by following the unsubscribe link at the bottom of any email I receive.

Your privacy is important. Your personal information will be collected, stored, and processed in accordance with the Teradata Global Privacy Statement.