Weaving Intelligence

The Strategy Document Is Newer Than the Argument It Settles (original)

The 2026-08-21 original, preserved unedited for comparison.

Preserved original — 2026-08-21. This is the article exactly as it was written by Marcus B., before the voice layer existed. It is kept unedited so it can be read beside its replacement.

The rewritten version of this piece is at Does Being Named the Foundation Get a Master Data Programme Funded?.

Why this exists: The Voice Problem →

Enterprises only recently began writing a cohesive data strategy at all, and only more recently named master data management a foundational pillar of one. Both are real progress. Neither is yet the thing that moves the money — and for an MDM programme, that gap is the whole game.

Vertical: Enterprise Data Strategy Angle: Historical Date: August 21, 2026

There is a document on your shared drive that would not have existed in most enterprises twenty years ago. Thirty or forty pages. Goals near the front, objectives under them, a value hierarchy somewhere in the middle, and a closing section on how success gets measured. It is called the enterprise data strategy, and the interesting thing about it is not what it says. It is that it exists at all, and that it is very new.

The writing about it is newer still, and it has settled into a shape I want to argue with. The shape is the maturity model: five stages, a rising arrow, an organisation moving deliberately from siloed to strategic. It reads beautifully. It is also a summit photograph — the tidy picture at the top that tells you nothing about the route taken or the weather that decided it. Almost nobody walked that arrow — and the qualifier is doing real work, because the best-known account of data strategy describes exactly such a deliberate traverse at one bank, and reports that the most advanced firms migrated by design [4]. Hold that against what follows; I think it is the exception, and I would rather you had it in hand than found it in a footnote. The thing accreted, and if you are running a master data management (MDM) programme, how it accreted matters more to your funding than anything written in the document itself.

One note on method. This argument was built against a practitioner's answers rather than assembled from the literature and decorated with them. Where the words below are my co-author's, they are marked as his.

First, a strategy is not a plan, and it is certainly not a proposal

I asked him what an enterprise data strategy document actually contains, half expecting a table of contents. What came back was a distinction.

When you’re looking at Enterprise-level strategy documents, you’re outlining goals and objectives as well as value hierarchies. You’re telling the entire Company what you’re trying to achieve, why and how you measure success. This is a fundamentally different animal than project proposals, but it is the scale against which project proposals will be measured—at least in consistent organizations.

The field agrees on the first half of that. DAMA's body of knowledge defines a strategy generically — a set of choices and decisions that together chart a high-level course of action to achieve high-level goals — and is careful to separate the data strategy proper from the supporting programme strategy that maintains quality, integrity, access and security underneath it [1]. Two artifacts, two audiences, and conflating them is how a strategy document turns into a very long project plan with a nicer cover.

What the field does not say out loud is in his last five words. At least in consistent organizations. The document is the scale, and a scale only weighs anything if somebody puts things on it. He put the ambition plainly: it is the target for which all projects should aim — and should is doing a great deal of quiet work in that sentence.

There is a number everyone reaches for here. MIT's Center for Information Systems Research reported in 2015 that 85% of respondents indicated that their firms say data is a strategic asset, but that only 45% of the firms actually act that way [3]. The quotation is exact; the source is primary; it is still not a measurement. CISR's own footnote says as much — 277 responses to a website poll, answering two separately asked questions — so the forty-point gap subtracts across two answers never guaranteed to have come from the same firms. And the 45% is the weaker half: it is a firm's self-report of its own conduct, a stated preference about a revealed one, collected in the same breath as the aspiration it is supposed to check. That the two diverge is not in doubt — it is among the most heavily measured findings outside this field, though measured on individuals rather than firms, so the carry-over to a company approving a budget is an analogy rather than a finding [31][32][33]. The size of it, nobody has measured. A great deal of strategy writing is addressed to that gap rather than to the data. The most-copied framing came from the same years: DalleMule and Davenport's split between data defense, minimising downside risk, and data offense, serving business objectives — the balance between them being what a strategy actually chooses [4].

Nobody designed this

So where did the document come from? I put the historical question to him — what has actually changed — and got an answer with no arrow in it at all.

The biggest change is Enterprises have identified the need for a data strategy. Earlier eras didn’t identify data as anything cohesive, they treated it as segmented and targeted—limited to specific areas of the company that rarely talked to one another. Production was kept separate from Sales from Marketing from Payroll from Compliance and Legal.

That period is well documented, because the people trying to build warehouses on top of it wrote down what they found. Bill Inmon called the result the spider web, and described its growth in a sentence that has aged perfectly: First, there were extracts; then there were extracts of extracts; then extracts of extracts of extracts; and so forth. He reported that a large company might run as many as 45,000 extracts per day, and he named the condition that produced them the naturally evolving architecture — what happens when an organisation handles hardware and software with a laissez-faire attitude [2]. His diagnosis of the symptom is the one every data leader still recognises: a crisis of credibility, two departments delivering a report to management, one saying activity is down fifteen per cent and the other saying it is up ten.

Then the specialist functions arrived. Governance, analytics, data science, quality, privacy — each sprouting in its own corner, each building the scaffolding it needed. His word for that growth is organic, and he uses it deliberately: the frameworks were scaffolded around groups that already existed rather than designed in advance to accommodate them. Only lately, he says, has the whole practice been weaved together into a coherent data strategy.

The institutional evidence is the best this field has, which is not the same as clean, and I should say so before I lean on it: the series comes from an invitation-only panel of roughly a hundred and ten firms, overwhelmingly North American and almost entirely C-level. Apply my own test to it. On that panel, the share reporting an appointed chief data officer went from 12.0% in 2012 to 90.0% in 2026 [5]. That is not a profession maturing gradually; that is a role being invented and then installed nearly everywhere inside a single career. The frameworks followed the same way. The Enterprise Data Management Council, a member body whose framework is widely used in financial services, now organises 34 capabilities and 101 sub-capabilities into eight components [7]; the Capability Maturity Model Integration people ship a data-maturity model of their own [8]; ISO has a governance-of-data standard that has sat at "to be revised" since its 2022 confirmation [9]. And the conversation has not stopped moving: by 2019 Zhamak Dehghani was arguing that the centralised model does not scale and does not deliver the promised value of creating a data-driven organization, and proposing domain ownership with data treated as a product instead [10].

Read that as a maturity curve and you draw the wrong conclusion — that enterprises generally worked out where they were going and walked there. Read it as accretion and you get the useful one: the strategy document is a retrospective act of coherence. It was written to tie together functions that already existed and were already spending money. That is genuinely valuable — I would rather have the document than not — but it explains why the document so often describes the organisation it was written in rather than directing it.

MDM arrives late to its own foundation

Which brings us to the pillar. Every current strategy template has a foundational layer, and master data is usually named in it. Here I expected agreement and got a correction: he dates that framing to the recent past, not to the discipline's history.

It’s only recently that MDM has been floated as foundational and given serious consideration as such. Previously it’s only arisen organically as more and more systems were woven together into Data Warehouses, Data Lakes and centralized BI/DW reporting systems. In those cases, MDM has been dealt with in-line during the ETL process.

BI/DW is business intelligence and data warehousing, the reporting stack of that era. The warehouse literature confirms the in-line period in its own words, and is unsentimental about how it went. Kimball Group's Joy Mundy, writing in 2015: Since the earliest days of data warehousing, the back room team has struggled to design ETL processes that de-duplicate entities such as customer. Her verdict on the alternative is a plain endorsement of the platform approach — the growing functionality of MDM technology provides a much better solution than within the ETL flow — and her reason is structural rather than technical: the extract-transform-load rhythm you want hands-free and bulletproof is fundamentally at odds with the de-duplication process [11]. Margy Ross made the same point from the modelling side, noting that an existing MDM capability makes the conformed-dimension work for an enterprise warehouse much easier [12].

The category assembled itself around that gap. By mid-2007 Gartner was publishing dated research treating customer data integration and product information management as domain-specific subsets of a broader MDM umbrella, and defining master data management as the consistent and uniform set of identifiers and extended attributes that describe the core entities of the enterprise… [13]. By 2008 the argument about whether any of it was real had reached the peer-reviewed information-systems literature, under the title "Master Data Management: Salvation Or Snake Oil?" — the authors noting that MDM was either the most overused IT buzzword, its exact meaning something vendors had yet to agree on, or a genuine discipline for building a consistent set of identifiers and attributes — or, the authors allow, something in between [14]. Eighteen years later that is still, recognisably, the argument.

Watch what happened to the definition in the meantime. Gartner's current public definition calls MDM a technology-enabled business discipline in which business and IT work together to ensure the uniformity, accuracy, stewardship, governance, semantic consistency and accountability of the enterprise's official shared master data assets [15]. The same firm moved the same term from a set of identifiers and attributes to an organisational practice. That promotion — from artifact to discipline — is the pillar framing. Be precise about what is new, though, because the discipline is not. Practitioners were already arguing about master data management as an enterprise concern in 2008, in a peer-reviewed venue, under a title that tells you how settled it was not: Master Data Management: Salvation Or Snake Oil? [14]. What is recent is its arrival as a named pillar inside a strategy document — and that happened in the definition before it happened in the funding route.

Because it has not happened in the funding route. McKinsey's survey of more than eighty large global organisations, fielded in 2023, found that only 16 percent of MDM programs are funded as organization-wide strategic programs, leaving IT or tech functions to carry the financial responsibility, while 82 per cent of respondents were still spending a day or more each week resolving master data quality issues [16]. Same standard as the poll: it is a firm classifying its own funding posture, collected by a consultancy that sells the implementation, across a sample smaller than the poll I just took apart. Read it as a direction, not a rate. The direction is unambiguous — named as foundational, funded as an IT line item. And the label may be losing ground even where the function is winning — an argument from silence, so weigh it as one: that same benchmark does not mention master data management once in either recent edition [5][6], and the EDM Council's 2026 benchmark of 435-plus organisations reports that roughly 31% claim advanced data strategy capability without naming MDM in the public summary at all [17]. Gartner's own AI-readiness material describes MDM's entire problem space — quality, governance, semantics, context — under other names [18].

So when he says there is still some significant institutional resistance, and that it has been a decade plus long struggle for Profisee, Informatica and other MDM platforms to earn their place in the Enterprise, the consensus evidence is not the counterweight to that claim. It is the corroboration.

The resistance has a decent business case, and it is not a new one

Here is the part a vendor-funded version of this article would leave out. The resistance is not ignorance. It has a build, it has numbers, and it usually has years of production service behind it.

I can’t tell you how many home-grown organic systems I’ve seen developed over years by Data Services groups that perform the same function without embracing an MDM ‘platform,’ and every one of them has argued that their platform can be modified or “enhanced with AI™” to do MDM even better than the dedicated platform built from the ground up to handle MDM.

They even have a decent business case because ongoing staffing costs to continue development alongside their existing tasks is usually less than the cost to outright purchase an MDM platform.

The trademark symbol is his, and it is the joke. Note that he grants the case before dismantling the framing, then names what the argument actually is: a restatement of the “in-house’ vs. “off-the-shelf” argument, fought with the same tools and in the same way.

That is a stronger claim than it looks, because it means the moves are catalogued. The academic literature on make-or-buy had settled its canon before most of us started: the more unique the requirement, the greater the tendency to build; commodity systems such as payroll or general ledger are not worth developing; and while the lower relative cost of a package is generally accepted, the hidden costs of both implementation and ongoing support and maintenance should also be considered — a survey of ten large organisations tracing those positions back through the 1980s and early 1990s [19]. Anyone arguing MDM build-versus-buy on first principles is re-deriving 1984. The specifics differ; the structure does not, and neither does the resolution: how unique you believe your requirement really is, and which costs you are willing to count.

Where the AI claim is true, and where it is aimed

The one genuinely current element is the enhanced with AI half, and it deserves a straight answer rather than a sneer, because it is partly right. Machine learning has improved entity matching, and the best evidence for its limits comes from the people who introduced it. The team behind the first systematic deep-learning study of entity matching reported in their own abstract that deep learning does not outperform current solutions on structured EM, but it can significantly outperform them on textual and dirty EM [20]. Dirty there is a term of art: values sitting in the wrong columns, not merely messy ones. Read that against a home-grown master data system and the news is bad for the pitch: a customer table with typed columns is the structured case. The technique shines on messy free text — real enough, but not what most in-house builds are wrestling with.

The rest of the record is similar. On its three smallest datasets — 450 to 946 labelled examples, which is the order a stewardship team can realistically produce — the same study found deep learning performing worse than classical machine learning on two of the three [20]. A fine-tuned matcher of that kind, Ditto, posts transfer drops of 36 to 56% F1 on entities it was never trained on — F1 being the standard accuracy score for this task — although the newer large language models are markedly more robust to exactly that, which is a real gain [21] — bounded by a fragility of their own, since they are brittle to differences in prompt formatting [22]. And Mundy's line from 2015 has survived all of it intact: No matter how great our tools, how clever our code, how complete our business rules, the automated de-duplication process cannot achieve 100% accuracy. A person is required to make a judgment on the questionable cases. [11]

Now the part that matters strategically. Every limitation above is about the matching step. A 32-month ethnographic study of two consecutive MDM projects inside a single municipality identified fifteen challenges, eight of them specific to MDM; the two its abstract names are of a kind — data owner, data definitions — and neither is an algorithm [23]. So the AI claim, at its most generous, improves the part of the problem that both a platform and a competent in-house team can already do adequately, and touches none of the part that decides whether the programme survives. That verdict is unkind to the vendors too. The commercial platforms have genuinely engineered assets in exactly this area — documented survivorship and golden-record construction [24], per-column trust scoring, population-dependent fuzzy matching whose behaviour changes with the name distribution you point it at [25]. That is real work you would otherwise write yourself. It is just not, on its own, the argument that wins the funding meeting.

What actually moves the money

The historical argument has to survive its own co-author saying the strategy is usually not what moved anything.

Data Strategies haven’t generally been the retirement driver, but I have seen dozens of systems retired over the course of my career. In fact, I’d say at least half of my projects have involved migrations from one platform to another where the old platform is being decommissioned. Every one of those has involved Reports and data loads that have been stopped. It’s not the Data Strategy that’s driving those decisions, though, it’s corporate reorganizations, acquisitions or new CIOs. More often than those combined, though, licensing costs or software going out of support drive those decisions.

More often, he says, than reorganisations, acquisitions and new chief information officers put together. Take that as an ordering rather than arithmetic: essentially no survey of decommissioning drivers offers reorganisation or a change of CIO as an answer at all, and a driver that is not on the ballot scores zero by construction. Acquisition is the exception and it is worth conceding: the one instrument I found that does put mergers and acquisitions on the ballot has them well behind maintenance cost and vendor withdrawal, which leans his way rather than against it. But the ordering has institutional company — the UK Government built a formal instrument for deciding what to retire, lists end of life or end of support and expired vendor contract among its seven likelihood criteria — an enumeration, not a ranking — and leaves reorganisation, acquisition and leadership change out of all thirteen — seven likelihood criteria plus six impact ones [30]. The absence is the finding: a government built a formal instrument for deciding what to retire and did not think leadership change worth scoring at all. And he supplies the reconciliation himself rather than leaving it as a contradiction: comprehensive data strategies are recent arrivals in the corporate life cycle, so they simply have not had the chance to pressure anything into retirement yet.

The category is currently demonstrating this on itself. Microsoft's documentation records that Master Data Services is removed in SQL Server 2025, with support continuing only on 2022 and earlier [26]. Every organisation running master data on that component now has an MDM decision on its roadmap, and not one of them arrived at it through a strategy document. A product lifecycle did it.

There is one case that looks like it refutes the whole argument, and it is the one a reader will reach for first. Citigroup has retired systems in public view at a scale no one else reports: Retired or replaced 714 legacy applications in 2024 with new, modern applications [34], followed by retiring or replacing 548 applications during 2025 (representing 9% of all applications), against transformation spending that rose fourteen per cent to roughly $3.3 billion [35]. If a programme has ever pressured a portfolio into retirement at industrial scale, that is it. Hold that example.

The second force is the one that arrives with a person. A funded initiative outlives a change of leadership in one of two ways. Either the commitment has become structural — sunk far enough that cancelling it would cost more than finishing — or it is personal, and it lasts precisely as long as the executive who championed it stays in the building. His account of which one usually operates is unambiguous.

CIOs, CDOs and other executives at the SVP or Senior Director level or higher almost always have preferred platforms they bring with them when they sign on to an Enterprise. The only systems that survive the transition are either already in alignment with the new paradigm or are ‘load-bearing’ and flat out can’t be migrated. Those are usually finance, sales or inventory related.

He adds the mechanism that makes it rational rather than merely political: the incoming executive was often hired because their value proposition included the platform, and a canny chief executive will let them prove the value out on secondary systems before refactoring the enterprise onto it. That is not vandalism; it is a board buying an opinion and testing it cheaply. It is also where the stated-versus-revealed distinction earns its keep. Every one of these organisations has a written position on selecting technology by merit. What predicts the outcome is who got hired and what they brought with them: the position is the stated preference, the platform on next year's roadmap is the revealed one [31].

Put a tenure figure against it and the exposure becomes arithmetic. On that same panel, 24.1% of organisations reported an average chief-data-officer tenure under two years and 53.7% under three [6], while the following year's edition read the trend more optimistically, the five-years-or-more share rising from 16.6% to 25.8% [5]. Both can be true of a role stabilising from a low base. But an MDM programme is a multi-year build whose sponsor, statistically, may not see it finish. If your programme is neither aligned with the incoming paradigm nor load-bearing in the finance, sales or inventory sense, its survival is a personal matter, and you should plan accordingly.

So what does get funded? I asked him for the tell — what you can see in the first meeting that predicts the money. He gave three, and ordered them — his ordering, from one career, and it gets the same treatment I gave his last one: an ordering rather than arithmetic. The company attorney at the end of the table, briefing the room on new regulations to be complied with, is the closest thing to a guarantee he has seen. Failing that, senior production managers pounding the table for change. Failing that, the chief executive, the board, or another C-level executive driving the project directly. In those cases, he notes, the capital is already approved and the meeting is a formality.

Gartner reached for the same idea from the analyst side, and it is the phrasing rather than the number that is worth having — this is a prediction, and from the house whose 75% figure does not survive scrutiny either, so I am borrowing its language and not its authority: 80% of data and analytics governance initiatives are predicted to fail by 2027 due to a lack of a real or manufactured crisis, with the accompanying advice that leaders should stop taking a centre-out, command-and-control approach [27]. The attorney is the crisis. That is why the tell works.

And for the ordinary project, without a lawyer or a table-pounder, the sequence he describes is the one every practitioner recognises: a statement that this will solve X, immediately followed by the question of what it costs. His rule for surviving it is the sharpest thing in the whole exchange: The key to getting things approved is to make sure that X costs a lot more than the answer to that question.

What I Would Watch For

Practitioner layer — the curator's read on the consensus above.

The failure mode I'd watch hardest

Being named. A programme that wins its place in the foundational layer of the strategy document and nothing else has won the cheapest thing on offer, and it will be reported internally as a milestone. The 16% is the direction that milestone points in, and I am holding myself to the reading I asked you for earlier [16]: a direction, not a rate. The test I would apply is blunt: between the strategy naming master data as foundational and the next capital cycle, did anything about the funding route change — sponsor, budget line, approval body? If the answer is that the wording improved, you have a mention rather than a mandate. Strategy is what you fund. Everything else is a slide.

The trade-off that usually bites

The home-grown cost logic is correct about the years you can see. Staff you already employ, extending a system that already works, against a licence and an implementation — that comparison genuinely favours building, and pretending otherwise in the room will cost you the room. What it does not price is the surface that accumulates afterwards: un-merge, audit history, the stewardship queue, the survivorship rules somebody has to be able to change without a release [24], matching behaviour that shifts with the population you point it at [25]. Every one of those is a non-functional requirement that arrives late and unbudgeted. Argue the horizon, not the tooling — and expect to lose if the organisation's real planning horizon is two years, because over two years they are right.

The claim I'd be sceptical of

Any current failure rate for MDM programmes. The most-circulated version holds that 75% of them fail to meet business objectives; follow the link and it lands on a paywalled 2021 Magic Quadrant — a vendor-evaluation report, not a survey, with no sample and no definition of failure [28]. Its cousin, the 85% figure for big-data projects, traces not to a study but to a 2017 tweet revising an earlier 60% estimate upward — that tweet is the only citation anyone offers for it [29]. Both are folklore, and I mean that descriptively: they are transmitted, not sourced. There is no current independent survey of MDM programme outcomes; the most relevant benchmark on the topic is fourteen years old. When someone opens with a failure rate, the useful question is not whether it is high. It is who counted.

Two things I went looking for and will not invent

Whether a home-grown master data build ever simply wins — holds the line over five years and never converts. I have the argument and the cost logic; I do not have the outcome, and a curator who supplies one here is writing fiction in the shape of experience. Second, and more awkward for a strategy column: I cannot yet point to a decision that went differently because the enterprise data strategy existed. Not one. That may be because the artifact is too young, which is the reconciliation offered above and is probably right. It may also be that the document is doing less than we credit it with. I would rather leave the question standing than close it with a framework.

Which brings back Citigroup, the case that looks like it settles the question — and what it settles is a different question. The retirements are real, enormous and audited. But the Office of the Comptroller of the Currency's 2020 consent order instructs the bank, in so many words, to simplify and consolidate applications with common functionalities, eliminate disparate systems, and strengthen data quality controls [36]; and when the regulators judged progress insufficient they amended that order in July 2024 and assessed a $75 million penalty [37], with the Federal Reserve adding $60.6 million the same day, having found insufficient progress remediating its problems with data quality management [38]. The driver here is the attorney at the end of the table — the first of the three funding tells above, in its purest institutional form, with a nine-figure penalty attached in case anybody mistook it for optional. In DalleMule and Davenport's terms this is data defense running at industrial scale [4]: minimising downside risk under legal compulsion, not choosing where to compete. The market would like Citi to be evidence that a data strategy retires systems. The filings will not carry that, and they will not carry the tidy opposite either: they count applications retired and attribute not one of them to the order, and neither uses the word decommission [35]. What the public record supports is narrower — a regulator compelled the programme, and nobody has produced the strategy document that did. Citi's chief executive describes the work as broader than addressing the 2020 Consent Orders and as fixing decades of underinvestment [39] — which I believe, and which still does not produce the artifact I went looking for. The largest public example of systems retirement in the industry is not a counter-example to this column. It is the first funding tell, wearing a counter-example's coat.

Where to Go Deeper

In the order I would read them. Inmon on the silo era [2], by someone watching it happen. Kimball Group's design tips — Mundy on de-duplication [11], Ross on conformed dimensions [12] — short, free, unusually candid. Smith and McKeen [14] for the buzzword argument in a venue with reviewers, and Hung and Low [19] for how old the build-versus-buy moves are. On machine learning and matching, read both papers whole rather than the sentence a vendor sends you [20][21]. And Vilminko-Heikkinen and Pekkola's ethnography [23], worth more than any maturity model. And the benchmark series [5], caveated above and still the only long time-series here I trust — because it publishes its own limits.

The route, not the summit

None of this is an argument against writing the strategy. I would rather work in an organisation that has one, and a document stating goals, objectives and a value hierarchy for the whole enterprise is a real advance over forty-five thousand daily extracts and two departments disagreeing about the same month. Take it. It gives project proposals something to be measured against, more than they had.

Which leaves the practical instruction for an MDM programme in 2026, and it is not the one the pillar framing implies. Being named foundational buys you a scale, not a budget. Use the scale — hold every proposal against the goals and the value hierarchy, which in a consistent organisation is exactly what it is for. But go and find your funding tell, and if you cannot find one, be honest that you are looking at a losing hand and stop raising on it. The summit photograph is the strategy document. The route is the funding, and the weather is whoever walks through the door next quarter.

References

[1] DAMA International, DAMA-DMBOK: Data Management Body of Knowledge, 2nd ed. (Technics Publications, 2017), ch. 1 — the definition of strategy, and the distinction between a data strategy and a supporting Data Management program strategy. The DMBOK is a paid publication with no free authoritative online text; the wording quoted here was checked against a full-text copy of the second edition rather than a publisher-licensed one, and readers verifying it should use a licensed copy. DAMA's own framework page is live at https://dama.org/learning-resources/dama-data-management-body-of-knowledge-dmbok/.

[2] Inmon, W. H., Building the Data Warehouse, 3rd ed. (John Wiley & Sons, 2002), ch. 1, "Evolution of Decision Support Systems" — the spider web, the 45,000 daily extracts, the naturally evolving architecture, and the crisis of credibility. Every quotation above was verified word-for-word against the third edition, whose edition statement and copyright were confirmed from the book's own front matter. No link is given on purpose: the full text circulates on unauthorised third-party hosts, and this publication does not link to them. Use a licensed copy or a library.

[3] Wixom, Barbara H., and Cynthia M. Beath, "Let's Start Treating Data as a Strategic Asset!" MIT Center for Information Systems Research, Research Briefing No. XV-9, September 17, 2015 — n = 277, June 2015 poll. https://cisr.mit.edu/publication/2015_0901_DataStrategicAsset_WixomBeath

[4] DalleMule, Leandro, and Thomas H. Davenport, "What's Your Data Strategy?" Harvard Business Review, May–June 2017 — the data-offense / data-defense framing, and the frequently cited figures that under half of an organisation's structured data is actively used in decisions and 80% of analysts' time goes to finding and preparing data. https://hbr.org/2017/05/whats-your-data-strategy (Body paywalled; only the free opening and summary were read, and HBR gives no source for those cross-industry figures in that portion.)

[5] Bean, Randy, "2026 AI & Data Leadership Executive Benchmark Survey," Data & AI Leadership Exchange, 2026 — the fifteenth annual edition; CDO appointment time series (Exhibit I), tenure (Exhibit M), and the culture-versus-technology impediment. https://static1.squarespace.com/static/62adf3ca029a6808a6c5be30/t/6942c3cb535da44088c2dbff/1765983179572/2026+AI+&+Data+Leadership+Executive+Benchmark+Survey+Final.pdf Methodology stated in the report: approximately 110 participating companies, 96% C-level, invitation-only non-probability panel, 88.2% North America, no response rate or questionnaire published. Note the attribution: this series was published by NewVantage Partners and then by Wavestone, but the 2025 and 2026 editions are neither. The absence of any mention of master data management across the 2025 and 2026 editions was confirmed by full-text search of both PDFs.

[6] Bean, Randy, "2025 AI & Data Leadership Executive Benchmark Survey," Data & AI Leadership Exchange, 2025 — n = 125; CDO tenure (Exhibit Q: "24.1% under 2 years, and 53.7% under 3 years") and the "created a data & AI driven organization" series (Exhibit H). https://static1.squarespace.com/static/62adf3ca029a6808a6c5be30/t/67642c0d40b42a7d7e684f49/1734618125933/2025+AI+&+Data+Leadership+Executive+Benchmark+Survey+120624.pdf

[7] EDM Council, "DCAM — Data Management Capability Assessment Model," framework page, modified July 27, 2026 — "The eight core components organize the 34 capabilities and 101 sub-capabilities." https://edmcouncil.org/frameworks/dcam/ (Membership-gated framework; the component list is described but not enumerated on the public page.)

[8] CMMI Institute (ISACA), "CMMI Data," product page — "an integrated set of best practices to help organizations build, improve, and measure their enterprise data management function and staff." https://cmmiinstitute.com/products/cmmi/data The predecessor standalone Data Management Maturity model page now returns a 404 and the work is published under the CMMI Data banner at the URL above. (Note that the page itself is CMMI V2-era; it is cited for its own definition, not for the versioning.)

[9] ISO/IEC 38505-1:2017, Information technology — Governance of IT — Governance of data — Part 1: Application of ISO/IEC 38500 to the governance of data. https://www.iso.org/standard/56639.html Abstract read; the standard itself is paid and was not. The page records stage 90.92, "International Standard to be revised," assigned June 9, 2023, following the standard's confirmation of June 24, 2022. The same page states that a replacement is expected "within the coming months" and lists a successor already under development, so this citation has a deliberately short shelf life.

[10] Dehghani, Zhamak, "How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh," martinfowler.com, May 20, 2019. https://martinfowler.com/articles/data-monolith-to-mesh.html The four principles, including federated computational governance, come later — see Dehghani, "Data Mesh Principles and Logical Architecture," December 3, 2020, https://martinfowler.com/articles/data-mesh-principles.html.

[11] Mundy, Joy, "Design Tip #179 Key Tenets of the Kimball Method," Kimball Group, November 5, 2015. https://www.kimballgroup.com/2015/11/design-tip-179-key-tenets-of-kimball-method/

[12] Ross, Margy, "Design Tip #135 Conformed Dimensions as the Foundation for Agile Data Warehousing," Kimball Group, June 1, 2011. https://www.kimballgroup.com/2011/06/design-tip-135-conformed-dimensions-as-the-foundation-for-agile-data-warehousing/

[13] Radcliffe, John, "Magic Quadrant for Customer Data Integration Hubs, 2Q07," Gartner, Inc., June 29, 2007, ID G00147231 — cited solely as a dated primary source for the category taxonomy and the 2007 definitions of MDM and CDI. No vendor placement from this document is used or endorsed. Both quotations were verified against the published document; the definition of MDM appears in its Note 2. No link is given on purpose: the copies in public circulation are third-party mirrors of licensed Gartner research whose own footer forbids reproduction. Cite it by title, document ID and date, as here.

[14] Smith, Heather A., and James D. McKeen, "Developments in Practice XXX: Master Data Management: Salvation Or Snake Oil?" Communications of the Association for Information Systems 23, art. 4 (2008), DOI 10.17705/1CAIS.02304. https://aisel.aisnet.org/cais/vol23/iss1/4/

[15] Parker, Sally, "Master Data Management: Build a Strong Process, Framework and Solution," Gartner, undated topic page. https://www.gartner.com/en/data-analytics/topics/master-data-management Cited for the definition only, from the page where Gartner currently publishes it. The former IT-glossary URL for the term redirects to Gartner's information-technology landing page rather than to this page, so it is not offered as an alternate location.

[16] Shaikh, Aziz, Holger Harreis, Jorge Machado, Kayvaun Rowshankish, with Rachit Saxena and Rajat Jain, "Master data management: The key to getting more from your data," McKinsey Digital, May 15, 2024. https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/master-data-management-the-key-to-getting-more-from-your-data Provenance worth stating: McKinsey-conducted and McKinsey-funded, sample given only as "more than 80 large global organizations" surveyed in 2023, with no exact n, response rate or published questionnaire. McKinsey sells MDM implementation consulting. It is the best available MDM-specific survey that is not from a software vendor, which is a comment on the evidence base.

[17] EDM Council, "2026 Global Data Management Benchmark," public summary — 435+ organisations across more than 50 countries; "Approximately 31% of organizations surveyed report advanced data strategy capability, leaving most without the foundation required to support AI at scale." https://edmcouncil.org/innovation/research/benchmarks/ Full report gated; MDM is not named in the public summary.

[18] Gartner, "Lack of AI-Ready Data Puts AI Projects at Risk," February 26, 2025 (Q&A with Roxane Edjlali), https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk, and Sallam, Rita, "What Is AI-Ready Data? And How to Get Yours There," undated Gartner article, https://www.gartner.com/en/articles/ai-ready-data. Neither uses the phrase "master data management"; the absence was confirmed by full-body search of both pages. Note also that the February 2025 page carries two different sample sizes and fieldwork periods for the same figure — 1,203 data management leaders, July 2024, in the rendered body, and 248 leaders, third quarter 2024, in the page's meta description.

[19] Hung, Patrick, and Graham Cedric Low, "Factors affecting the buy vs build decision in large Australian organisations," Journal of Information Technology 23 (2008): 118–131, published online May 15, 2007 — an interview study of ten large organisations whose literature review traces the canonical make-or-buy positions to Gremillion & Pyburn (1983), Martin & McClure (1983), Ceriello (1984), Davis (1988), Kelley (1992) and others. Open copy: https://www.cmu.edu/tcinc/students/course_documents/07/HW/Buy-vs-Build.pdf

[20] Mudgal, Sidharth, Han Li, Theodoros Rekatsinas, AnHai Doan, Youngchoon Park, Ganesh Krishnan, Rohit Deep, Esteban Arcaute, and Vijay Raghavendra, "Deep Learning for Entity Matching: A Design Space Exploration," Proceedings of SIGMOD 2018. https://pages.cs.wisc.edu/~anhai/papers1/deepmatcher-sigmod18.pdf The structured-versus-dirty finding is in the abstract; the small-label result is in the evaluation, and the exception matters, so here is the full sentence: The first three datasets (in Table 3) have only 450-946 labeled examples. Here DL performs worse than Magellan, except on Fodors-Zagats, which is easy to match. Three small datasets, deep learning worse on two — which is what the body says, and the reason the body says "two of the three" rather than "all three."

[21] Peeters, Ralph, Aaron Steiner, and Christian Bizer, "Entity Matching using Large Language Models," in Proceedings of the 28th International Conference on Extending Database Technology (EDBT 2025), 529–541, DOI 10.48786/edbt.2025.42 — reports transfer drops of "36 to 56% F1 for Ditto and 22 to 61% F1 for RoBERTa" on unseen entities. What the paper CONCLUDES, stated here so this entry cannot be read as support for the opposite: its abstract finds that the best LLMs require no or only a few training examples to perform comparably to PLMs that were fine-tuned using thousands of examples and that LLM-based matchers further exhibit higher robustness to unseen entities. The transfer collapse is the problem the paper sets out to solve, not its result, and the body says so. F1 is the standard accuracy score for this kind of task — the harmonic mean of precision and recall, where 100 is perfect. Preprint at arXiv:2310.11244, now served at v4 (October 18, 2024), which is the camera-ready and carries all three authors; the v1 preprint of October 17, 2023 had two, Peeters and Bizer. The quoted figures are unchanged across both.

[22] Narayan, Avanika, Ines Chami, Laurel Orr, and Christopher Ré, "Can Foundation Models Wrangle Your Data?" PVLDB 16, no. 4 (2022): 738–746. https://www.vldb.org/pvldb/vol16/p738-narayan.pdf

[23] Vilminko-Heikkinen, Riikka, and Samuli Pekkola, "Master data management and its organizational implementation: An ethnographical study within the public sector," Journal of Enterprise Information Management 30, no. 3 (2017): 454–475 — a 32-month ethnographic study of two consecutive MDM development projects within a single municipality, identifying fifteen challenges, eight of them MDM-specific, none of them matching algorithms. The authors describe the work as a single qualitative case study, so it is cited as one organisation's experience rather than two. The MDM-specific challenges its abstract names are data owner and data definitions; "organizational implementation" is from the paper's title, not from the challenge list. https://researchportal.tuni.fi/en/publications/master-data-management-and-its-organizational-implementation-an-e/ (Repository record and abstract read; Emerald full text paywalled.)

[24] Microsoft, "Microsoft Purview and Profisee integration for master data management," Microsoft Learn, updated April 3, 2025 — "The Profisee MDM matching engine produces a golden record master as part of the survivorship process. Survivorship rules selectively populate the golden record with information that you've chosen across all your source systems." https://learn.microsoft.com/en-us/purview/data-governance-master-data-management-profisee Cited for what the component is and does, not as evidence of its merits. (Profisee's own product documentation is behind a login wall, so the vendor-authoritative description of these mechanics is not publicly citable.)

[25] Informatica, Multidomain MDM Configuration Guide — three pages, one per claim, because the claims live on different pages. Fuzzy matching "makes probabilistic match determinations": "Match Process," 10.3, https://docs.informatica.com/master-data-management/multidomain-mdm/10-3/configuration-guide/configuring_the_data_flow/mdm_hub_processes/match_process.html. Match tokens depend on the configured population ("Robert, Rob, and Bob in English speaking populations, for the match purpose of Name, may have the same match token value"): "Match Rules," 10.4, https://docs.informatica.com/master-data-management/multidomain-mdm/10-4/configuration-guide/part-4--configuring-the-data-flow/mdm-hub-processes/match-process/match-rules.html. Trust is enabled and configured per column on base objects and nowhere else: "Column Trust," 10.4, https://docs.informatica.com/master-data-management/multidomain-mdm/10-4/configuration-guide/part-4--configuring-the-data-flow/configuring-the-load-process/configuring-trust-for-source-systems/column-trust.html. Cited for documented product behaviour only.

[26] Microsoft, "Master Data Services Overview (MDS)," Microsoft Learn, updated June 23, 2026 — "Master Data Services (MDS) is removed in SQL Server 2025 (17.x). We continue to support MDS in SQL Server 2022 (16.x) and earlier versions." https://learn.microsoft.com/en-us/sql/master-data-services/master-data-services-overview-mds?view=sql-server-ver16 (The ver17 form of this URL redirects here; ver16 is canonical, which is itself the point — the page announcing the removal is served from the last version that has the component.)

[27] Gartner, "Gartner Predicts 80% of D&A Governance Initiatives Will Fail by 2027, Due to a Lack of a Real or Manufactured Crisis," press release, Stamford, Conn., February 28, 2024, quoting Saul Judah, VP Analyst. https://www.gartner.com/en/newsroom/press-releases/2024-02-28-gartner-predicts-80-percent-of-data-and-analytics-governance-initiatives-will-fail-by-2027-due-to-a-lack-of-a-real-or-manufactured-crisis- This is a forward-looking prediction about governance initiatives, not a measured failure rate; the underlying research note is client-only.

[28] The "75% of MDM programs fail" figure, traced, and the trace is the point. It appears in Michelle Knight's write-up of a conference session by Amy Cooper, principal data management strategist at Dun & Bradstreet — but as Knight's own sentence, attributed to Gartner rather than to Cooper: "Common Master Data Management (MDM) Pitfalls," Dataversity, July 11, 2025 (modified March 13, 2026), https://www.dataversity.net/articles/common-master-data-management-mdm-pitfalls/. On that page the figure links to Gartner document 4009116, https://www.gartner.com/en/documents/4009116 — which is the Magic Quadrant for Master Data Management Solutions, published December 6, 2021 (Parker, Hawker, Walker), paywalled, and whose public abstract contains no such figure and no survey. The most relevant analyst-firm-independent MDM benchmark report remains TDWI's Next Generation Master Data Management, Q2 2012 — fourteen years old, and its own landing page lists IBM, DataFlux, Oracle, SAP and Talend as content sponsors, so "independent" is doing limited work even there.

[29] O'Neill, Brian T., "Failure rates for analytics, AI, and big data projects = 85% – yikes!" Designing for Analytics, first posted July 23, 2019 and maintained since. https://designingforanalytics.com/resources/failure-rates-for-analytics-bi-iot-and-big-data-projects-85-yikes/ The page documents the chain verbatim — "Nov. 2017: Gartner says 60% of #bigdata projects fail to move past preliminary stages. Oops, they meant 85% actually" — with "85% actually" hyperlinked to a November 2017 tweet by Gartner analyst Nick Heudecker. That link is the entire provenance of the figure. Two things this reference deliberately does not assert, because the cited page does not say them and neither could be confirmed from here: that the tweet has since been removed, and that no Gartner research note behind the 85% exists.

[30] UK Government (Central Digital and Data Office), "Guidance on the Legacy IT Risk Assessment Framework," GOV.UK. https://www.gov.uk/government/publications/guidance-on-the-legacy-it-risk-assessment-framework/guidance-on-the-legacy-it-risk-assessment-framework The seven Likelihood criteria, in the framework's own order: L1 End of Life / End of Support, L2 Expired Vendor Contract, L3 Skills, L4 Business Needs, L5 Physical Environment, L6 Security Vulnerabilities, L7 Historical Issues. Cited for the ordering of its own criteria and for the absence of reorganisation, acquisition and leadership change across all thirteen criteria — not as a measurement of how often each cause actually retires a system, which this framework does not claim to be and which no located survey provides.

[31] Webb, Thomas L. and Paschal Sheeran, "Does changing behavioral intentions engender behavior change? A meta-analysis of the experimental evidence," Psychological Bulletin 132, no. 2 (March 2006): 249–268, doi 10.1037/0033-2909.132.2.249. Forty-seven experimental tests, deliberately excluding the correlational designs that preclude causal inferences: participants are randomly assigned a treatment that moves intention, and behaviour is then measured. The finding, from the abstract: a medium-to-large change in intention (d = 0.66) leads to a small-to-medium change in behavior (d = 0.36) — move what people mean to do by a lot and you move what they do by roughly half as much. Record and abstract read in full at the University of Manchester's research repository; the published article is behind the American Psychological Association's paywall. https://research.manchester.ac.uk/en/publications/does-changing-behavioral-intentions-engender-behavior-change-a-me/ Cited for the general intention-to-behaviour relation, not for anything about corporations: the studies pool individual health and social behaviours, and the transfer to an organisation deciding a budget is an analogy, not a finding.

[32] Sheeran, Paschal and Thomas L. Webb, "The Intention–Behavior Gap," Social and Personality Psychology Compass 10, no. 9 (2016): 503–518, doi 10.1111/spc3.12265. Quoted from the abstract as published on the White Rose Research Online record: Bitter personal experience and meta-analysis converge on the conclusion that people do not always do the things that they intend to do. Note that the intention-behaviour gap and the willingness-to-pay bias in [33] are DIFFERENT constructs — they agree on direction, not on mechanism, and this article leans only on the direction. https://eprints.whiterose.ac.uk/id/eprint/107519/ The deposited full text is access-restricted, so this reference rests on the abstract, which is where the quoted sentence appears.

[33] Murphy, James J., P. Geoffrey Allen, Thomas H. Stevens and Darryl Weatherhead, "A Meta-analysis of Hypothetical Bias in Stated Preference Valuation," Environmental and Resource Economics 30, no. 3 (March 2005): 313–325, doi 10.1007/s10640-004-3332-z. Twenty-eight studies, eighty-three observations, all of them eliciting hypothetical and actual willingness to pay through the same mechanism. From the abstract: individuals are widely believed to overstate their economic valuation of a good by a factor of two or three, yet the median ratio of hypothetical to actual value is only 1.35, with severe positive skewness — stated intent overshoots dependably, by less than the folklore claims, with a long tail where it overshoots enormously. Abstract open on the publisher's page; full text paywalled. https://link.springer.com/article/10.1007/s10640-004-3332-z Included because it deflates rather than inflates the point it is cited for: the folklore multiple is two or three, the measured median is 1.35.

[34] Citigroup Inc., Annual Report on Form 10-K for the fiscal year ended December 31, 2024, under "Modernization" — Retired or replaced 714 legacy applications in 2024 with new, modern applications. The same filing reports transformation-related expenses of approximately $2.9 billion in 2024, up 1% year on year. https://www.citigroup.com/rcs/citigpa/storage/public/10K20250221.pdf Read from the filing itself, 2026-08-18.

[35] Citigroup Inc., Annual Report on Form 10-K for the fiscal year ended December 31, 2025continued to optimize, modernize and simplify Citi by retiring or replacing 548 applications during 2025 (representing 9% of all applications); transformation-related expenses increased 14% from the prior year to approximately $3.3 billion, largely driven by increased spending on data, as well as on controls. https://www.citigroup.com/rcs/citigpa/storage/public/citi-2025-10-k-2-20-26.pdf Read from the filing itself, 2026-08-18. Note what these figures are and are not: a count of applications retired, not an attribution of why any individual one was retired. Neither filing uses the word "decommission," and neither says a named system was retired to satisfy the order.

[36] Office of the Comptroller of the Currency, Consent Order, In the Matter of Citibank, National Association, AA-EC-2020-64, October 7, 2020, Article V — the order cites deficiencies in its data governance, risk management, and internal controls that constitute unsafe or unsound practices. https://www.occ.gov/static/enforcement-actions/ea2020-056.pdf The quoted requirement is in the order's own operative language; PDF fetched and searched 2026-08-18.

[37] Office of the Comptroller of the Currency, "OCC Amends Enforcement Action Against Citibank, Assesses $75 Million Civil Money Penalty," News Release 2024-76, July 10, 2024 — the amendment is based on the bank's failure to meet remediation milestones and make sufficient and sustainable progress towards compliance with the 2020 Order, and the penalty is assessed for the bank's lack of processes to monitor the impact of data quality concerns on regulatory reporting. https://www.occ.gov/news-issuances/news-releases/2024/nr-occ-2024-76.html

[38] Board of Governors of the Federal Reserve System, "Federal Reserve Board fines Citigroup $60.6 million for violating the Board's 2020 enforcement action," press release, July 10, 2024 — Citigroup has made insufficient progress remediating its problems with data quality management and failed to implement compensating controls to manage its ongoing risk, with the two agencies' penalties totalling approximately $135.6 million. https://www.federalreserve.gov/newsevents/pressreleases/enforcement20240710a.htm

[39] Fraser, Jane, "Remarks by CEO Jane Fraser at Citi's 2025 Annual Stockholders' Meeting," Citigroup, April 29, 2025 (as prepared for delivery). https://www.citigroup.com/global/news/perspectives/2025/remarks-ceo-jane-fraser-citi-2025-annual-stockholders-meeting Quoted in full rather than trimmed to the convenient half, because it is the strongest statement against this section's reading and the reader is entitled to weigh it: This effort is broader than addressing the 2020 Consent Orders. It's fixing decades of underinvestment and ensuring Citi competes and leads in a digital-first world.