July 29, 2026 in Data Mangement
Enterprises Are Scaling AI on Foundations That Weren’t Built for It. Enter DataOps.
SHARE: PRINT ARTICLE:
https://doi.org/10.1287/LYTX.2026.03.01
In just the past couple years, AI has rapidly moved from experiment to enterprise imperative. Semarchy’s 2026 survey of 1,000 C-suite leaders across the United States, United Kingdom, and France found that 73% of executives now consider AI a top strategic priority, up from 25% just one year ago. Nearly all of those surveyed (97%) say they are actively investing in AI technologies. The mandates are real, the budgets are following, and the competitive pressure is relentless.
But speed is exposing a structural problem that no amount of ambition can paper over: the data foundations underneath most enterprise AI were not built for it. Organizations are deploying AI agents, LLMs, and automation tools on top of data infrastructure that can’t be versioned, tested, or promoted reliably. They’re trying to govern AI outputs without first resolving which “John Smith” in their CRM is the same “J. Smith” in their ERP. And they are discovering, often too late, that ungoverned data doesn’t just limit AI – it actively corrupts it.
The answer isn’t a multi-year data remediation program. It’s DataOps, and enterprises that treat it as an afterthought are building on sand.
You Can’t Wait for Perfect Data
The legacy approach to enterprise data management (comprehensive MDM programs, waterfall governance models, waiting for “golden” data before enabling any consumption) cannot survive contact with AI’s velocity. By the time a traditional program delivers its first governed domain, the AI landscape has moved three generations.
The temptation to boil the ocean on data quality before deploying AI is understandable. The problem is that your competitors aren’t waiting. Every quarter spent on a perfect data program is a quarter they use to deploy AI on governed-but-imperfect data and iterate toward better outcomes. Consider that only 57% of enterprises currently have DataOps capabilities, meaning that almost half are building AI on data foundations they cannot version, test, or deploy reliably. Perfection is the enemy of progress here, and the enterprises winning with AI right now are the ones that found the agile middle ground. That middle ground is DataOps.
DataOps Is Architectural Commitment
DataOps needs to be front and center in enterprise AI strategy, equal in importance to AI development and delivery. The core principle is straightforward: treat data like software by iterating, automating, and shipping continuously.
In practice, this means data products move through environments the same way software does: developing, testing, staging, and producing. There are gates at each transition: quality gates, test gates, dependency gates, governance gates. Every asset versions together, promotes together, and governs together as a coherent unit. This model enables enterprises to do something the traditional approach never allowed: start small, learn fast, and expand deliberately.
Start with one domain – master customer data – and get AI consuming it in weeks. Observe what’s working, including quality scores, API usage, and AI consumption patterns. Iterate by improving matching rules, enriching semantic models, and tightening quality thresholds. Then expand by adding product data and supplier data, each building on the same proven promotion model.
No serious software organization ships by waiting until the code is perfect. They ship iteratively, with CI/CD pipelines, automated testing, observability, and rollback capability. Enterprise data must work exactly the same way. The organizations that have internalized this are building durable AI infrastructure. The ones still waiting for perfect data are falling further behind with every sprint cycle.
A New Type of Consumer
Master data management (MDM) has historically been about producing clean, consolidated records for human-facing applications, including CRM screens, ERP dashboards, and compliance reports. The consumer was always a person who could spot an obvious duplicate, tolerate a little inconsistency, and flag something that looked wrong. AI is a fundamentally different kind of consumer, and MDM programs that haven’t adapted are creating serious downstream risk.
When a human user encounters duplicate customer records, they notice and work around it. When an AI agent encounters duplicates, it reasons about them as distinct entities, producing different lifetime value calculations, different risk scores, and different recommendations for the same customer. And it does this at scale, across thousands of records, without ever flagging the inconsistency. The AI doesn’t know it’s wrong; it’s confident it’s right.
This is why MDM is not a “nice to have” for enterprise AI. Rather, it is the foundational layer that ensures AI reasons about the right entities with the right relationships governed by the right policies. Without it, you don’t have an AI problem; you have a data problem that AI is amplifying at machine speed.
The Investment Disconnect
Semarchy’s survey revealed that although 51% of enterprises say that data management is their number-one AI challenge, only 40% are prioritizing it for investment. That 11-point gap between recognizing the problem and funding the solution is where AI initiatives stall.
Budget is flowing toward production AI, agentic use cases, and GenAI tooling, while the data infrastructure underneath is underfunded. It’s like investing in faster cars while ignoring the quality of the roads.
The enterprises that close this gap – those that treat DataOps as a first-class architectural commitment rather than an operational afterthought – are the ones that will be able to scale AI with confidence, iterate without chaos, and actually measure whether any of it is working. The ones that do not will keep deploying AI on foundations that weren’t built for it and wondering why their results don’t match their ambition.
Craig Gravina is chief technology officer at Semarchy, a leader in master data management, intelligence and integrated solutions for global enterprises. With deep expertise in AI, cloud and data technologies, Craig is recognized in the sector as a developer of disruptive, market-leading solutions. In his previous role as CTO at ObvioHealth, he led digital transformation in clinical trials using AI, machine learning and cloud technologies. Specializing in distributed architectures, SaaS/PaaS business models and AI/ML innovations, Craig is passionate about bridging business, product and technology to deliver scalable, high-impact solutions that empower organizations to unlock the full potential of their data.