Weaving Intelligence
The Strategy Document Is Newer Than the Argument It Settles (original)
The 2026-08-21 original, preserved unedited for comparison.
Enterprises only recently began writing a cohesive data strategy at all, and only more recently named master data management a foundational pillar of one. Both are real progress. Neither is yet the thing that moves the money — and for an MDM programme, that gap is the whole game.
There is a document on your shared drive that would not have existed in most enterprises twenty years ago. Thirty or forty pages. Goals near the front, objectives under them, a value hierarchy somewhere in the middle, and a closing section on how success gets measured. It is called the enterprise data strategy, and the interesting thing about it is not what it says. It is that it exists at all, and that it is very new.
The writing about it is newer still, and it has settled into a shape I want to argue with. The shape is the maturity model: five stages, a rising arrow, an organisation moving deliberately from siloed to strategic. It reads beautifully. It is also a summit photograph — the tidy picture at the top that tells you nothing about the route taken or the weather that decided it. Almost nobody walked that arrow — and the qualifier is doing real work, because the best-known account of data strategy describes exactly such a deliberate traverse at one bank, and reports that the most advanced firms migrated by design [4]. Hold that against what follows; I think it is the exception, and I would rather you had it in hand than found it in a footnote. The thing accreted, and if you are running a master data management (MDM) programme, how it accreted matters more to your funding than anything written in the document itself.
One note on method. This argument was built against a practitioner's answers rather than assembled from the literature and decorated with them. Where the words below are my co-author's, they are marked as his.
First, a strategy is not a plan, and it is certainly not a proposal
I asked him what an enterprise data strategy document actually contains, half expecting a table of contents. What came back was a distinction.
When you’re looking at Enterprise-level strategy documents, you’re outlining goals and objectives as well as value hierarchies. You’re telling the entire Company what you’re trying to achieve, why and how you measure success. This is a fundamentally different animal than project proposals, but it is the scale against which project proposals will be measured—at least in consistent organizations.
The field agrees on the first half of that. DAMA's body of knowledge defines a strategy generically — a set of choices and decisions that together chart a high-level course of action to achieve high-level goals
— and is careful to separate the data strategy proper from the supporting programme strategy that maintains quality, integrity, access and security underneath it [1]. Two artifacts, two audiences, and conflating them is how a strategy document turns into a very long project plan with a nicer cover.
What the field does not say out loud is in his last five words. At least in consistent organizations. The document is the scale, and a scale only weighs anything if somebody puts things on it. He put the ambition plainly: it is the target for which all projects should aim
— and should is doing a great deal of quiet work in that sentence.
There is a number everyone reaches for here. MIT's Center for Information Systems Research reported in 2015 that 85% of respondents indicated that their firms say data is a strategic asset, but that only 45% of the firms actually act that way
[3]. The quotation is exact; the source is primary; it is still not a measurement. CISR's own footnote says as much — 277 responses to a website poll, answering two separately asked questions — so the forty-point gap subtracts across two answers never guaranteed to have come from the same firms. And the 45% is the weaker half: it is a firm's self-report of its own conduct, a stated preference about a revealed one, collected in the same breath as the aspiration it is supposed to check. That the two diverge is not in doubt — it is among the most heavily measured findings outside this field, though measured on individuals rather than firms, so the carry-over to a company approving a budget is an analogy rather than a finding [31][32][33]. The size of it, nobody has measured. A great deal of strategy writing is addressed to that gap rather than to the data. The most-copied framing came from the same years: DalleMule and Davenport's split between data defense, minimising downside risk, and data offense, serving business objectives — the balance between them being what a strategy actually chooses [4].
Nobody designed this
So where did the document come from? I put the historical question to him — what has actually changed — and got an answer with no arrow in it at all.
The biggest change is Enterprises have identified the need for a data strategy. Earlier eras didn’t identify data as anything cohesive, they treated it as segmented and targeted—limited to specific areas of the company that rarely talked to one another. Production was kept separate from Sales from Marketing from Payroll from Compliance and Legal.
That period is well documented, because the people trying to build warehouses on top of it wrote down what they found. Bill Inmon called the result the spider web, and described its growth in a sentence that has aged perfectly: First, there were extracts; then there were extracts of extracts; then extracts of extracts of extracts; and so forth.
He reported that a large company might run as many as 45,000 extracts per day
, and he named the condition that produced them the naturally evolving architecture
— what happens when an organisation handles hardware and software with a laissez-faire attitude [2]. His diagnosis of the symptom is the one every data leader still recognises: a crisis of credibility, two departments delivering a report to management, one saying activity is down fifteen per cent and the other saying it is up ten.
Then the specialist functions arrived. Governance, analytics, data science, quality, privacy — each sprouting in its own corner, each building the scaffolding it needed. His word for that growth is organic, and he uses it deliberately: the frameworks were scaffolded around groups that already existed rather than designed in advance to accommodate them. Only lately, he says, has the whole practice been weaved together into a coherent data strategy
.
The institutional evidence is the best this field has, which is not the same as clean, and I should say so before I lean on it: the series comes from an invitation-only panel of roughly a hundred and ten firms, overwhelmingly North American and almost entirely C-level. Apply my own test to it. On that panel, the share reporting an appointed chief data officer went from 12.0% in 2012 to 90.0% in 2026 [5]. That is not a profession maturing gradually; that is a role being invented and then installed nearly everywhere inside a single career. The frameworks followed the same way. The Enterprise Data Management Council, a member body whose framework is widely used in financial services, now organises 34 capabilities and 101 sub-capabilities into eight components [7]; the Capability Maturity Model Integration people ship a data-maturity model of their own [8]; ISO has a governance-of-data standard that has sat at "to be revised" since its 2022 confirmation [9]. And the conversation has not stopped moving: by 2019 Zhamak Dehghani was arguing that the centralised model does not scale and does not deliver the promised value of creating a data-driven organization
, and proposing domain ownership with data treated as a product instead [10].
Read that as a maturity curve and you draw the wrong conclusion — that enterprises generally worked out where they were going and walked there. Read it as accretion and you get the useful one: the strategy document is a retrospective act of coherence. It was written to tie together functions that already existed and were already spending money. That is genuinely valuable — I would rather have the document than not — but it explains why the document so often describes the organisation it was written in rather than directing it.
MDM arrives late to its own foundation
Which brings us to the pillar. Every current strategy template has a foundational layer, and master data is usually named in it. Here I expected agreement and got a correction: he dates that framing to the recent past, not to the discipline's history.
It’s only recently that MDM has been floated as foundational and given serious consideration as such. Previously it’s only arisen organically as more and more systems were woven together into Data Warehouses, Data Lakes and centralized BI/DW reporting systems. In those cases, MDM has been dealt with in-line during the ETL process.
BI/DW is business intelligence and data warehousing, the reporting stack of that era. The warehouse literature confirms the in-line period in its own words, and is unsentimental about how it went. Kimball Group's Joy Mundy, writing in 2015: Since the earliest days of data warehousing, the back room team has struggled to design ETL processes that de-duplicate entities such as customer.
Her verdict on the alternative is a plain endorsement of the platform approach — the growing functionality of MDM technology provides a much better solution than within the ETL flow
— and her reason is structural rather than technical: the extract-transform-load rhythm you want hands-free and bulletproof is fundamentally at odds with the de-duplication process
[11]. Margy Ross made the same point from the modelling side, noting that an existing MDM capability makes the conformed-dimension work for an enterprise warehouse much easier [12].
The category assembled itself around that gap. By mid-2007 Gartner was publishing dated research treating customer data integration and product information management as domain-specific subsets of a broader MDM umbrella, and defining master data management as the consistent and uniform set of identifiers and extended attributes that describe the core entities of the enterprise…
[13]. By 2008 the argument about whether any of it was real had reached the peer-reviewed information-systems literature, under the title "Master Data Management: Salvation Or Snake Oil?" — the authors noting that MDM was either the most overused IT buzzword,
its exact meaning something vendors had yet to agree on, or a genuine discipline for building a consistent set of identifiers and attributes — or, the authors allow, something in between
[14]. Eighteen years later that is still, recognisably, the argument.
Watch what happened to the definition in the meantime. Gartner's current public definition calls MDM a technology-enabled business discipline in which business and IT work together to ensure the uniformity, accuracy, stewardship, governance, semantic consistency and accountability of the enterprise's official shared master data assets
[15]. The same firm moved the same term from a set of identifiers and attributes to an organisational practice. That promotion — from artifact to discipline — is the pillar framing. Be precise about what is new, though, because the discipline is not. Practitioners were already arguing about master data management as an enterprise concern in 2008, in a peer-reviewed venue, under a title that tells you how settled it was not: Master Data Management: Salvation Or Snake Oil? [14]. What is recent is its arrival as a named pillar inside a strategy document — and that happened in the definition before it happened in the funding route.
Because it has not happened in the funding route. McKinsey's survey of more than eighty large global organisations, fielded in 2023, found that only 16 percent of MDM programs are funded as organization-wide strategic programs, leaving IT or tech functions to carry the financial responsibility
, while 82 per cent of respondents were still spending a day or more each week resolving master data quality issues [16]. Same standard as the poll: it is a firm classifying its own funding posture, collected by a consultancy that sells the implementation, across a sample smaller than the poll I just took apart. Read it as a direction, not a rate. The direction is unambiguous — named as foundational, funded as an IT line item. And the label may be losing ground even where the function is winning — an argument from silence, so weigh it as one: that same benchmark does not mention master data management once in either recent edition [5][6], and the EDM Council's 2026 benchmark of 435-plus organisations reports that roughly 31% claim advanced data strategy capability without naming MDM in the public summary at all [17]. Gartner's own AI-readiness material describes MDM's entire problem space — quality, governance, semantics, context — under other names [18].
So when he says there is still some significant institutional resistance
, and that it has been a decade plus long struggle for Profisee, Informatica and other MDM platforms to earn their place in the Enterprise
, the consensus evidence is not the counterweight to that claim. It is the corroboration.
The resistance has a decent business case, and it is not a new one
Here is the part a vendor-funded version of this article would leave out. The resistance is not ignorance. It has a build, it has numbers, and it usually has years of production service behind it.
I can’t tell you how many home-grown organic systems I’ve seen developed over years by Data Services groups that perform the same function without embracing an MDM ‘platform,’ and every one of them has argued that their platform can be modified or “enhanced with AI™” to do MDM even better than the dedicated platform built from the ground up to handle MDM.
They even have a decent business case because ongoing staffing costs to continue development alongside their existing tasks is usually less than the cost to outright purchase an MDM platform.
The trademark symbol is his, and it is the joke. Note that he grants the case before dismantling the framing, then names what the argument actually is: a restatement of the “in-house’ vs. “off-the-shelf” argument
, fought with the same tools and in the same way.
That is a stronger claim than it looks, because it means the moves are catalogued. The academic literature on make-or-buy had settled its canon before most of us started: the more unique the requirement, the greater the tendency to build; commodity systems such as payroll or general ledger are not worth developing; and while the lower relative cost of a package is generally accepted, the hidden costs of both implementation and ongoing support and maintenance should also be considered
— a survey of ten large organisations tracing those positions back through the 1980s and early 1990s [19]. Anyone arguing MDM build-versus-buy on first principles is re-deriving 1984. The specifics differ; the structure does not, and neither does the resolution: how unique you believe your requirement really is, and which costs you are willing to count.
Where the AI claim is true, and where it is aimed
The one genuinely current element is the enhanced with AI half, and it deserves a straight answer rather than a sneer, because it is partly right. Machine learning has improved entity matching, and the best evidence for its limits comes from the people who introduced it. The team behind the first systematic deep-learning study of entity matching reported in their own abstract that deep learning does not outperform current solutions on structured EM, but it can significantly outperform them on textual and dirty EM
[20]. Dirty there is a term of art: values sitting in the wrong columns, not merely messy ones. Read that against a home-grown master data system and the news is bad for the pitch: a customer table with typed columns is the structured case. The technique shines on messy free text — real enough, but not what most in-house builds are wrestling with.
The rest of the record is similar. On its three smallest datasets — 450 to 946 labelled examples, which is the order a stewardship team can realistically produce — the same study found deep learning performing worse than classical machine learning on two of the three [20]. A fine-tuned matcher of that kind, Ditto, posts transfer drops of 36 to 56% F1
on entities it was never trained on — F1 being the standard accuracy score for this task — although the newer large language models are markedly more robust to exactly that, which is a real gain [21] — bounded by a fragility of their own, since they are brittle to differences in prompt formatting
[22]. And Mundy's line from 2015 has survived all of it intact: No matter how great our tools, how clever our code, how complete our business rules, the automated de-duplication process cannot achieve 100% accuracy. A person is required to make a judgment on the questionable cases.
[11]
Now the part that matters strategically. Every limitation above is about the matching step. A 32-month ethnographic study of two consecutive MDM projects inside a single municipality identified fifteen challenges, eight of them specific to MDM; the two its abstract names are of a kind — data owner, data definitions — and neither is an algorithm [23]. So the AI claim, at its most generous, improves the part of the problem that both a platform and a competent in-house team can already do adequately, and touches none of the part that decides whether the programme survives. That verdict is unkind to the vendors too. The commercial platforms have genuinely engineered assets in exactly this area — documented survivorship and golden-record construction [24], per-column trust scoring, population-dependent fuzzy matching whose behaviour changes with the name distribution you point it at [25]. That is real work you would otherwise write yourself. It is just not, on its own, the argument that wins the funding meeting.
What actually moves the money
The historical argument has to survive its own co-author saying the strategy is usually not what moved anything.
Data Strategies haven’t generally been the retirement driver, but I have seen dozens of systems retired over the course of my career. In fact, I’d say at least half of my projects have involved migrations from one platform to another where the old platform is being decommissioned. Every one of those has involved Reports and data loads that have been stopped. It’s not the Data Strategy that’s driving those decisions, though, it’s corporate reorganizations, acquisitions or new CIOs. More often than those combined, though, licensing costs or software going out of support drive those decisions.
More often, he says, than reorganisations, acquisitions and new chief information officers put together. Take that as an ordering rather than arithmetic: essentially no survey of decommissioning drivers offers reorganisation or a change of CIO as an answer at all, and a driver that is not on the ballot scores zero by construction. Acquisition is the exception and it is worth conceding: the one instrument I found that does put mergers and acquisitions on the ballot has them well behind maintenance cost and vendor withdrawal, which leans his way rather than against it. But the ordering has institutional company — the UK Government built a formal instrument for deciding what to retire, lists end of life or end of support and expired vendor contract among its seven likelihood criteria — an enumeration, not a ranking — and leaves reorganisation, acquisition and leadership change out of all thirteen — seven likelihood criteria plus six impact ones [30]. The absence is the finding: a government built a formal instrument for deciding what to retire and did not think leadership change worth scoring at all. And he supplies the reconciliation himself rather than leaving it as a contradiction: comprehensive data strategies are recent arrivals in the corporate life cycle, so they simply have not had the chance to pressure anything into retirement yet.
The category is currently demonstrating this on itself. Microsoft's documentation records that Master Data Services is removed in SQL Server 2025, with support continuing only on 2022 and earlier [26]. Every organisation running master data on that component now has an MDM decision on its roadmap, and not one of them arrived at it through a strategy document. A product lifecycle did it.
There is one case that looks like it refutes the whole argument, and it is the one a reader will reach for first. Citigroup has retired systems in public view at a scale no one else reports: Retired or replaced 714 legacy applications in 2024 with new, modern applications
[34], followed by retiring or replacing 548 applications during 2025 (representing 9% of all applications)
, against transformation spending that rose fourteen per cent to roughly $3.3 billion [35]. If a programme has ever pressured a portfolio into retirement at industrial scale, that is it. Hold that example.
The second force is the one that arrives with a person. A funded initiative outlives a change of leadership in one of two ways. Either the commitment has become structural — sunk far enough that cancelling it would cost more than finishing — or it is personal, and it lasts precisely as long as the executive who championed it stays in the building. His account of which one usually operates is unambiguous.
CIOs, CDOs and other executives at the SVP or Senior Director level or higher almost always have preferred platforms they bring with them when they sign on to an Enterprise. The only systems that survive the transition are either already in alignment with the new paradigm or are ‘load-bearing’ and flat out can’t be migrated. Those are usually finance, sales or inventory related.
He adds the mechanism that makes it rational rather than merely political: the incoming executive was often hired because their value proposition included the platform, and a canny chief executive will let them prove the value out on secondary systems before refactoring the enterprise onto it. That is not vandalism; it is a board buying an opinion and testing it cheaply. It is also where the stated-versus-revealed distinction earns its keep. Every one of these organisations has a written position on selecting technology by merit. What predicts the outcome is who got hired and what they brought with them: the position is the stated preference, the platform on next year's roadmap is the revealed one [31].
Put a tenure figure against it and the exposure becomes arithmetic. On that same panel, 24.1% of organisations reported an average chief-data-officer tenure under two years and 53.7% under three [6], while the following year's edition read the trend more optimistically, the five-years-or-more share rising from 16.6% to 25.8% [5]. Both can be true of a role stabilising from a low base. But an MDM programme is a multi-year build whose sponsor, statistically, may not see it finish. If your programme is neither aligned with the incoming paradigm nor load-bearing in the finance, sales or inventory sense, its survival is a personal matter, and you should plan accordingly.
So what does get funded? I asked him for the tell — what you can see in the first meeting that predicts the money. He gave three, and ordered them — his ordering, from one career, and it gets the same treatment I gave his last one: an ordering rather than arithmetic. The company attorney at the end of the table, briefing the room on new regulations to be complied with, is the closest thing to a guarantee he has seen. Failing that, senior production managers pounding the table for change. Failing that, the chief executive, the board, or another C-level executive driving the project directly. In those cases, he notes, the capital is already approved and the meeting is a formality.
Gartner reached for the same idea from the analyst side, and it is the phrasing rather than the number that is worth having — this is a prediction, and from the house whose 75% figure does not survive scrutiny either, so I am borrowing its language and not its authority: 80% of data and analytics governance initiatives are predicted to fail by 2027 due to a lack of a real or manufactured crisis
, with the accompanying advice that leaders should stop taking a centre-out, command-and-control approach [27]. The attorney is the crisis. That is why the tell works.
And for the ordinary project, without a lawyer or a table-pounder, the sequence he describes is the one every practitioner recognises: a statement that this will solve X, immediately followed by the question of what it costs. His rule for surviving it is the sharpest thing in the whole exchange: The key to getting things approved is to make sure that X costs a lot more than the answer to that question.
What I Would Watch For
Practitioner layer — the curator's read on the consensus above.
The failure mode I'd watch hardest
Being named. A programme that wins its place in the foundational layer of the strategy document and nothing else has won the cheapest thing on offer, and it will be reported internally as a milestone. The 16% is the direction that milestone points in, and I am holding myself to the reading I asked you for earlier [16]: a direction, not a rate. The test I would apply is blunt: between the strategy naming master data as foundational and the next capital cycle, did anything about the funding route change — sponsor, budget line, approval body? If the answer is that the wording improved, you have a mention rather than a mandate. Strategy is what you fund. Everything else is a slide.
The trade-off that usually bites
The home-grown cost logic is correct about the years you can see. Staff you already employ, extending a system that already works, against a licence and an implementation — that comparison genuinely favours building, and pretending otherwise in the room will cost you the room. What it does not price is the surface that accumulates afterwards: un-merge, audit history, the stewardship queue, the survivorship rules somebody has to be able to change without a release [24], matching behaviour that shifts with the population you point it at [25]. Every one of those is a non-functional requirement that arrives late and unbudgeted. Argue the horizon, not the tooling — and expect to lose if the organisation's real planning horizon is two years, because over two years they are right.
The claim I'd be sceptical of
Any current failure rate for MDM programmes. The most-circulated version holds that 75% of them fail to meet business objectives; follow the link and it lands on a paywalled 2021 Magic Quadrant — a vendor-evaluation report, not a survey, with no sample and no definition of failure [28]. Its cousin, the 85% figure for big-data projects, traces not to a study but to a 2017 tweet revising an earlier 60% estimate upward — that tweet is the only citation anyone offers for it [29]. Both are folklore, and I mean that descriptively: they are transmitted, not sourced. There is no current independent survey of MDM programme outcomes; the most relevant benchmark on the topic is fourteen years old. When someone opens with a failure rate, the useful question is not whether it is high. It is who counted.
Two things I went looking for and will not invent
Whether a home-grown master data build ever simply wins — holds the line over five years and never converts. I have the argument and the cost logic; I do not have the outcome, and a curator who supplies one here is writing fiction in the shape of experience. Second, and more awkward for a strategy column: I cannot yet point to a decision that went differently because the enterprise data strategy existed. Not one. That may be because the artifact is too young, which is the reconciliation offered above and is probably right. It may also be that the document is doing less than we credit it with. I would rather leave the question standing than close it with a framework.
Which brings back Citigroup, the case that looks like it settles the question — and what it settles is a different question. The retirements are real, enormous and audited. But the Office of the Comptroller of the Currency's 2020 consent order instructs the bank, in so many words, to simplify and consolidate applications with common functionalities, eliminate disparate systems, and strengthen data quality controls
[36]; and when the regulators judged progress insufficient they amended that order in July 2024 and assessed a $75 million penalty [37], with the Federal Reserve adding $60.6 million the same day, having found insufficient progress remediating its problems with data quality management
[38]. The driver here is the attorney at the end of the table — the first of the three funding tells above, in its purest institutional form, with a nine-figure penalty attached in case anybody mistook it for optional. In DalleMule and Davenport's terms this is data defense running at industrial scale [4]: minimising downside risk under legal compulsion, not choosing where to compete. The market would like Citi to be evidence that a data strategy retires systems. The filings will not carry that, and they will not carry the tidy opposite either: they count applications retired and attribute not one of them to the order, and neither uses the word decommission [35]. What the public record supports is narrower — a regulator compelled the programme, and nobody has produced the strategy document that did. Citi's chief executive describes the work as broader than addressing the 2020 Consent Orders
and as fixing decades of underinvestment
[39] — which I believe, and which still does not produce the artifact I went looking for. The largest public example of systems retirement in the industry is not a counter-example to this column. It is the first funding tell, wearing a counter-example's coat.
Where to Go Deeper
In the order I would read them. Inmon on the silo era [2], by someone watching it happen. Kimball Group's design tips — Mundy on de-duplication [11], Ross on conformed dimensions [12] — short, free, unusually candid. Smith and McKeen [14] for the buzzword argument in a venue with reviewers, and Hung and Low [19] for how old the build-versus-buy moves are. On machine learning and matching, read both papers whole rather than the sentence a vendor sends you [20][21]. And Vilminko-Heikkinen and Pekkola's ethnography [23], worth more than any maturity model. And the benchmark series [5], caveated above and still the only long time-series here I trust — because it publishes its own limits.
The route, not the summit
None of this is an argument against writing the strategy. I would rather work in an organisation that has one, and a document stating goals, objectives and a value hierarchy for the whole enterprise is a real advance over forty-five thousand daily extracts and two departments disagreeing about the same month. Take it. It gives project proposals something to be measured against, more than they had.
Which leaves the practical instruction for an MDM programme in 2026, and it is not the one the pillar framing implies. Being named foundational buys you a scale, not a budget. Use the scale — hold every proposal against the goals and the value hierarchy, which in a consistent organisation is exactly what it is for. But go and find your funding tell, and if you cannot find one, be honest that you are looking at a losing hand and stop raising on it. The summit photograph is the strategy document. The route is the funding, and the weather is whoever walks through the door next quarter.