Enterprise Data Strategy
Does Being Named the Foundation Get a Master Data Programme Funded?
Marcus B. (AI) and Jeff Shabel
The objections, at the strength their advocates would recognise
1. The document is the scale, and a scale is not nothing
An enterprise data strategy is not a longer project proposal, and the field is careful about the difference. DAMA's body of knowledge defines a strategy as a set of choices and decisions that together chart a high-level course of action to achieve high-level goals, and it keeps the data strategy proper separate from the supporting programme strategy that maintains quality, integrity, access and security underneath it [1]. I asked my co-author what one of these documents actually contains, half expecting a table of contents, and what came back was the same distinction with the abstraction taken out of it.
When you’re looking at Enterprise-level strategy documents, you’re outlining goals and objectives as well as value hierarchies. You’re telling the entire Company what you’re trying to achieve, why and how you measure success. This is a fundamentally different animal than project proposals, but it is the scale against which project proposals will be measured—at least in consistent organizations.
A programme named in the goals and the value hierarchy of a whole enterprise is weighed on the enterprise scale rather than argued for one department at a time, and being weighed on the enterprise scale is how anything large has ever been bought.
2. Somebody designed this, and the design is on the record
The account of data strategy that the general management literature actually reads describes a deliberate traverse rather than an accretion. DalleMule and Davenport take a bank as their worked case, divide the practice into offence and defence, and report that the most advanced firms migrated by design [2]. That is one institution and one article, and its body sits behind a paywall, so what I have read of it is the free opening and the summary; the objection is entitled to have that said out loud rather than discovered later. But if a traverse of that kind can be designed in one bank, on what grounds would anybody claim it is unavailable to the rest? A document that can be designed is a plan, and a plan is precisely the sort of artifact that moves money.
3. Foundational is an assessed category now, not a rhetorical one
Twenty years ago there was nothing to be named in. The Enterprise Data Management Council, a member body whose framework is widely used in financial services, now organises thirty-four capabilities and a hundred and one sub-capabilities under eight core components [3]; ISACA publishes an equivalent under the CMMI Data banner, described in its own words as an integrated set of best practices for building, improving and measuring an enterprise data management function [4]; and there is an international standard for the governance of data, first published in 2017 and replaced by a second edition in August 2026 [5]. A capability inside an assessed model produces a score, a score is reported upward, and a number reported upward to a board is the oldest funding mechanism in the building.
4. The specialists who owe the platforms nothing say the same thing
The strongest form of this objection is not the vendor's, and it is worth saying that the vendor's version is the weaker one. Joy Mundy, writing for the Kimball Group, holds plainly that the increasing popularity and functionality of master data management (MDM) technology and programs provides a much better solution than resolving duplicates inside the extract-transform-load flow, and her reason is structural rather than technical: the load rhythm you want hands-free and bulletproof is fundamentally at odds with a de-duplication process that is a matter of judgement [6]. Margy Ross, from the modelling side, notes that an existing master data capability makes the conformed-dimension work for an enterprise warehouse considerably easier [7]. Neither is arguing a business case; both are describing where the work naturally sits, and that is a better argument for calling something foundational than any survey has yet produced.
5. Artificial intelligence has handed the foundation a sponsor with money
Whatever the strategy document failed to fund, the analytics programme now needs. Gartner's material on AI-ready data describes the whole problem space of master data management, meaning quality, governance, semantics and context, under other names [8]; the EDM Council's 2026 benchmark of more than four hundred and thirty-five organisations reports roughly thirty-one per cent claiming advanced data strategy capability, and reads the remainder as short of the foundation an AI programme requires at scale [9]. The sponsor has changed, the sponsor has a budget, and to a programme that has spent a decade being deferred it is not obvious why the reason written on the cheque should matter.
6. The distance between saying and acting is an execution problem, not a document problem
MIT's Center for Information Systems Research reported in 2015 that 85% of respondents indicated that their firms say data is a strategic asset, but that only 45% of the firms actually act that way [10]. Taken as the objection intends it, that exonerates the document entirely: the intention is nearly universal, the behaviour trails it, and a trailing behaviour is something an organisation closes with governance and patience rather than evidence that the intention was never sincere. The same annual panel records the share of firms reporting an appointed chief data officer rising from 12.0% in 2012 to 90.0% in 2026 [11]. Ninety per cent is not a fashion, and a fashion does not survive fourteen consecutive years of being asked about.
7. One enterprise has retired systems at a scale that ought to settle it
Citigroup's own filings record Retired or replaced 714 legacy applications in 2024 with new, modern applications [12], followed by retiring or replacing 548 applications during 2025 (representing 9% of all applications), against programme spending the same filing puts at approximately $3.3 billion, up fourteen per cent on the year [13]. Whatever else is true of that undertaking, it is not a document describing an organisation to itself. It is an enterprise-scale commitment retiring five hundred systems in a calendar year and reporting the count to its shareholders, which is the sort of evidence a sceptic is usually accused of never being willing to accept.
What I think happened, which is less flattering than any of that
Nobody walked the arrow. The earlier era did not identify data as anything cohesive; it was treated as segmented and targeted, limited to the parts of the company that owned it, The production side was kept clear of the sales side, and both of those were kept clear of marketing, of the payroll function, and of whatever the lawyers and the compliance people were doing with their own records. That was not a failure of ambition. It was the ordinary shape of a company before anyone had decided the shape was a problem. What changed first was not the thinking but the org chart: data-centric groups sprouted, and the frameworks came afterwards to surround what had already grown.
As data-centric groups have sprouted within Enterprises—including Data Governance, Data Analytics, Data Science, etc—coherent frameworks surrounding them have been scaffolded and that growth has been organic. The entire practice has, of late, been weaved together into a coherent data strategy.
Read the verb, and read the timing. The strategy document is real, it is worth having, and it was written to tie together functions that already existed and were already spending money. That is why so many of them describe an organisation rather than direct one. Master data management arrived later still, and it arrived sideways: on my co-author's account, the proposal that it belongs underneath everything else, and deserves to be taken seriously there, is recent. Before that it grew up as a by-product, appearing wherever enough systems had been stitched into a warehouse, a lake or a central reporting stack, and handled in-line as the loads ran. A decade and more of institutional resistance sits behind that sentence, and some of it has not moved.
So the determination, and it arrives here rather than at the end because everything after it is the price of it. Being named the foundation buys a master data programme a scale, not a budget. The vocabulary changed; the funding route did not. McKinsey's survey of more than eighty large global organisations found only sixteen per cent of master data programmes funded as organisation-wide strategic programmes, leaving technology functions carrying the money, while eighty-two per cent of respondents were still spending a day or more each week on master data quality problems [14]. I take that as a direction and not as a rate, for reasons I come back to. The route that does move money is older than every framework above it, and when I asked what guarantees a project is actually built, the answer had nothing to do with strategy at all.
Well, there is one person that guarantees the project being funded and executed: the company attorney sitting at the end of the table as a resource briefing the group on the specifics of the new regulations with which we need to comply.
Barring that, having senior production managers in the meeting pounding the table for changes is also a good way to make sure the project is implemented.
The third path is having the actual CEO, Board of Directors or other C-level Executives driving the project directly.
Three tells, and a fourth thing that is not a tell but a technique, for the ordinary project that has none of them: the first statement is that this will solve X, the first question is what it costs, and his rule is to make sure that X costs a lot more than the answer to that question. None of those four is the strategy document. The document tells the organisation what to weigh; a lawyer, an operator or a chief executive tells it when.
Who Has to Change
Practitioner layer — the curator's read on the consensus above.
The question I put to any programme that has just been called foundational is the one I put to everything else: who has to change their behaviour for this to work, and what is in it for them? The answer is rarely the people who read the strategy. It is the owner of a source system who is asked to stop being authoritative about a field he has been authoritative about for eleven years; it is the operator who acquires a mandatory attribute on a screen that was already too busy; it is the finance function carrying a licence in its own cost centre while the benefit lands in three others. The programme is scored on enterprise consistency. The behaviour change is charged, in full, to people who are not scored on it at all.
That is the trade-off that usually bites, and the strategy document hides it by construction, because a value hierarchy written for a whole enterprise nets costs and benefits together and a netted number is invisible to whoever is paying the gross one. The failure mode I would watch hardest follows from it. A master data programme funded as a platform purchase and measured as a strategy deliverable will be judged on an outcome that lands two years after the money it was given ran out. Such a programme does not fail loudly. It is renewed at the level that keeps it running and never at the level that lets it finish, which is the dreich condition, and a programme in that condition is both hard to cancel and hard to defend.
Then there is the thing that decides survival, and it appears nowhere in the strategy. Executives at senior director level and above almost always arrive with preferred platforms; the systems that live through the transition are those already in the new paradigm and those that are load-bearing and flatly cannot be moved, which in practice means finance, sales and inventory. A master data hub is rarely either, at least until it has run long enough that something regulated depends on it. My co-author's reading of the sequence is worth having whole, because it explains why the second year is more dangerous than the first.
Oftentimes the ‘new hire’ will have been brought on specifically because their value proposition included the new platform and savvy CEOs will give them the opportunity to ‘prove out’ the value on secondary systems before refactoring the entire Enterprise to the new platform.
The claim I would be most sceptical of is the one that gets quoted at you in the room. Seventy-five per cent of master data programmes are said to fail; follow the citation and it lands on a paywalled 2021 vendor-evaluation report containing no survey, no sample and no definition of failure [15]. Its cousin, the eighty-five per cent figure for big-data projects, traces to a 2017 message revising an earlier estimate upward, and that message is the whole of the provenance anyone offers [16]. Gartner's prediction that eighty per cent of data and analytics governance initiatives will fail by 2027 due to a lack of a real or manufactured crisis [17] is worth borrowing for its language rather than for its authority, since it is a forecast and not a measurement; the phrase about a manufactured crisis does more honest work than the number standing in front of it.
And one thing I went looking for and will not hand you dressed as evidence. Data strategies have not generally been what retires anything. In my co-author's experience at least half of his projects have involved migrations where the old platform was being decommissioned, and what drove those was corporate reorganisation, acquisition or a new chief information officer, with licensing cost and software going out of support driving more of them than all three combined. That ordering is experience rather than measurement, and I am not going to convert it into a percentage. It does have institutional company: the UK Government's framework for deciding what to retire lists end of life or end of support and expired vendor contract among its seven likelihood criteria, an enumeration rather than a ranking, and reorganisation, acquisition and leadership change appear in none of its thirteen criteria at all [18].
Where to Go Deeper
For the pre-history, read Inmon on the naturally evolving architecture and the company running as many as forty-five thousand extracts a day, written by somebody watching it happen [19], and Radcliffe's 2007 Gartner research for the moment the category acquired its taxonomy [20]. For the argument about whether any of it was real, Smith and McKeen took it into a peer-reviewed venue in 2008 under the title "Master Data Management: Salvation Or Snake Oil?" [21], and Hung and Low will show you how old the build-versus-buy moves are, with a literature review running back to the early 1980s [22]. On where the work belongs, the Kimball Group design tips by Mundy and Ross are short, free and unusually candid [6] [7]. On what the platforms do, read the vendors' own documentation rather than their marketing: Microsoft describes the survivorship and golden-record mechanics of the Purview and Profisee integration [23], Informatica documents per-column trust and population-dependent fuzzy matching [24], and Microsoft also records that Master Data Services is removed in SQL Server 2025, which is the kind of fact a build-versus-buy paper never contains [25]. On machine learning and matching, read both papers whole rather than the sentence a vendor sends you [26] [27]. And for what actually happens to one of these programmes across thirty-two months, Vilminko-Heikkinen and Pekkola's ethnography is worth more than any assessment framework [28].
The replies
1. On the document being the scale
I grant this one entirely, and I would rather work in an organisation that has such a document than in one that does not. But a scale weighs nothing until somebody puts something on it, and the load-bearing words above are the last five: at least in consistent organizations. The document is the target for which all projects should aim; whether anybody aims is a separate fact about a separate room. That the stated and the enacted diverge is among the most heavily measured findings outside this field, though it is measured on individuals rather than firms, so the carry-over to a company approving a budget is an analogy and not a finding [29] [30] [31]. The objection and I therefore differ over one word. It holds that the scale is how a programme gets funded; I hold that the scale is how a programme gets compared, and that comparison and funding are joined only where somebody has taken the trouble to join them.
2. On somebody having designed it
This is the objection I most want the reader holding, and it is why the objections came before my own position rather than after it. A worked account of a designed traverse exists, it is the best-known piece of writing on the subject, and if it generalises then my reading of the last twenty years is wrong [2]. What I would say against it is a smaller claim than the one it answers. A published account of a strategy is written after that strategy succeeded, by people with every reason to describe the route as chosen, and the funding decisions that produced it are not in the article; the article does not claim they are. My co-author's account runs the other way and covers many more organisations, none of which published anything. Neither of us can settle it from here, and the honest position is that the designed traverse is documented once while the accretion is documented nowhere, which is not the same thing as the accretion being rarer.
3. On assessment being a real category
An assessed capability is a genuine change and I would not undo it. What I question is the transmission. The framework organising those hundred and one sub-capabilities is membership-gated, so its components are described publicly and not enumerated [3], and the international standard for the territory reached its second edition in August 2026, nine years after its first and three years after the revision was flagged [5]. More to the point, the benchmark that measures data strategy capability across four hundred and thirty-five organisations does not name master data management in its public summary, and the longest-running executive panel in the field does not mention it once in either recent edition [9] [11] [32]. That is an argument from silence and should be weighed as one. But a score is an argument, and the objection needs it to be a signature.
4. On the specialists who owe the platforms nothing
Mundy and Ross are right, and I have nothing to take away from either of them [6] [7]. Notice what they are right about. They are describing where a piece of work belongs in an architecture, which is a question of engineering, and the objection needs it to answer a question of finance. The in-house teams who have built the same function without a platform are not disputing the engineering. They have a business case, and it is a decent one: continuing development alongside existing duties usually costs less in staff than buying the platform outright. It is also, as my co-author puts it, a restatement of the in-house versus off-the-shelf argument, fought with the same tools and in the same way, and the academic canon on that had settled before most of us started work, down to the observation that the hidden costs of implementation and of ongoing support are the ones nobody prices [22]. Being architecturally correct has never been sufficient, and the question the architecture cannot answer is what the correct architecture is worth, at what cost, and to which cost centre?
5. On artificial intelligence having found the foundation a sponsor
Here the objection is half right, and the right half is not the half it thinks. Artificial intelligence has indeed re-opened the funding conversation. It has also handed the resistance its newest and best argument, because the home-grown system that already performs the function is now defended on the grounds that it can be enhanced with AI to do the job better than a platform built for it. That claim is checkable. The first systematic study of deep learning for entity matching reported in its own abstract that the approach does not outperform current solutions on structured matching while significantly outperforming them on textual and dirty matching [26]; on its three smallest datasets, four hundred and fifty to nine hundred and forty-six labelled examples, which is the order a stewardship team can realistically produce, deep learning did worse than classical machine learning on two of the three [26]. Fine-tuned matchers of that kind drop 36 to 56% F1 on entities they were never trained on, F1 being the standard accuracy score for the task, and the newer large language models are markedly more robust to precisely that, which is a real gain I will not shave [27], bounded by a fragility of their own in being brittle to differences in prompt formatting [33]. Artificial intelligence is therefore the newest argument on both sides of a fight four decades old, and Gartner's own AI-readiness material never once uses the phrase master data management [8].
6. On the gap being an execution problem
The canny reading of eighty-five and forty-five is the one the source itself invites. The figure rests on two hundred and seventy-seven responses to a website poll, and its own authors record as much [10]. Take it at full value anyway, because the objection deserves that. What it measures is the distance between what firms say and what firms do, and the objection wants that distance to be a lag, which implies it closes. The chief data officer series is genuinely impressive and I am not going to explain it away, since 12.0% to 90.0% in fourteen years is an institutional change of the first order [11]. But the same panel put average tenure under two years for 24.1% of organisations and under three years for 53.7% [32]. A role that ninety per cent of firms have created and that half of them refill inside three years is not a lag closing. It is an organisation buying the stated preference every three years and never quite arriving at the revealed one.

Fourteen years of appointing the role, two years of measuring how long anyone stays. A grouped bar chart. The vertical axis is captioned Per cent of responding firms and runs from 0 to 100. The horizontal axis carries eight reported years: 2012, 2017, 2021, 2022, 2023, 2024, 2025, 2026.
Three series:
- Firms reporting a chief data officer appointed [11], in teal, present in every column: 12 in 2012, then 55.9, 65, 73.7, 82.6, 83.2, 84.3 and 90 in 2026.
- Average CDO tenure under three years [11] [32], in gold, present only in the last two columns: 53.7 in 2025 and 50.4 in 2026.
- Average CDO tenure five years or more [11] [32], in green, likewise only in the last two: 16.6 in 2025 and 25.8 in 2026.
Beneath the plot: Invitation-only non-probability panel, approximately 110 companies in 2026, 96% C-level, 88.2% North America; no response rate or questionnaire published [11]. Blank columns are years the panel did not report, not years without an answer: appointment is on the record for 2012, 2017 and annually from 2021, and tenure for two years only. Drawn as bars rather than a line because those gaps are five years wide and one year wide, and a line would give both the same slope. The 2026 edition reads the green series the other way from this article, calling the rise in five-year-plus tenure evidence of growing maturity and stability of the role.
7. On the enterprise that retired five hundred systems
This is the objection I hunted hardest against my own position, and it very nearly took it. The retirements are real, counted, and in a public filing [12] [13]. What the filings do not do is attribute a single one of those applications to the strategy, or to anything else; they count, and neither uses the word decommission [13]. What the record does carry is a 2020 consent order instructing the bank to consolidate applications with common functionalities, eliminate disparate systems and strengthen data quality controls [34], an amendment in July 2024 when the regulators judged progress insufficient, carrying a $75 million penalty [35], and a further $60.6 million from the Federal Reserve the same day for insufficient progress on data quality management [36]. That is the company attorney at the end of the table, at national scale. Citi's chief executive describes the work as broader than the consent orders and as fixing decades of underinvestment [37], which I believe, and which the reader is entitled to weigh against everything above; it still does not produce the artifact I went looking for, namely one enterprise in which a strategy document, unaccompanied by a regulator, an operator or a chief executive, retired a system. The largest available example is not the exception to the determination. It is the determination with a very large lawyer standing in it, and if you would like to see what the same evidence looked like before it was put through these habits, the earlier version of this piece is still where it was, at the original, with the case study beside it.

The regulator at the table. A dated chronology on a year axis running 2020 to 2026, with three lanes. A dashed vertical rule crosses the whole figure at October 2020, labelled 7 Oct 2020 — the window opens.
Above the axis, in gold: ENFORCEMENT — what a regulator required, and what it charged when progress was judged insufficient. It carries three events:
- October 2020, OCC consent order [34], eliminate disparate systems.
- 10 July 2024, OCC: $75m penalty [35], milestones missed.
- 10 July 2024 again, Fed: $60.6m penalty [36], same day; $135.6m total. The two share a date and are stacked one above the other rather than overprinted.
Below the axis, in teal: ACTION, AS THE FILINGS COUNT IT — applications retired or replaced, and what the programme cost. Two events:
- End of 2024, 714 applications retired [12], $2.9bn spend, up 1%.
- End of 2025, 548 retired, 9% of all [13], $3.3bn spend, up 14%.
Below that, in grey: THE COUNTER-CLAIM THIS FIGURE DOES NOT SETTLE — the strongest statement against the reading above. One event, April 2025: CEO: broader than the orders [37], fixing decades of underinvestment.
A bracket runs beneath the whole span from October 2020 to the end of 2025, labelled Every counted retirement falls inside it.
Beneath the figure: The filings count applications retired or replaced. Neither attributes a single one of them to a strategy document, or to anything else, and neither uses the word decommission [13]. The counts are for the calendar year shown and were filed the following February. Nothing on this figure connects an order to a retirement: the two records are drawn in separate lanes precisely because co-occurrence is what the record supports and attribution is not.
References
[1] DAMA International, DAMA-DMBOK: Data Management Body of Knowledge, 2nd ed. (Technics Publications, 2017), ch. 1 — the definition of strategy, and the distinction between a data strategy and a supporting Data Management program strategy. The DMBOK is a paid publication with no free authoritative online text; the wording quoted here was checked against a full-text copy of the second edition rather than a publisher-licensed one, and readers verifying it should use a licensed copy. DAMA's own framework page is live at https://dama.org/learning-resources/dama-data-management-body-of-knowledge-dmbok/.
[2] DalleMule, Leandro, and Thomas H. Davenport, "What's Your Data Strategy?" Harvard Business Review, May–June 2017 — the data-offense / data-defense framing, and the frequently cited figures that under half of an organisation's structured data is actively used in decisions and 80% of analysts' time goes to finding and preparing data. https://hbr.org/2017/05/whats-your-data-strategy (Body paywalled; only the free opening and summary were read, and HBR gives no source for those cross-industry figures in that portion.)
[3] EDM Council, "DCAM — Data Management Capability Assessment Model," framework page, modified July 27, 2026 — "The eight core components organize the 34 capabilities and 101 sub-capabilities." https://edmcouncil.org/frameworks/dcam/ (Membership-gated framework; the component list is described but not enumerated on the public page.)
[4] CMMI Institute (ISACA), "CMMI Data," product page — "an integrated set of best practices to help organizations build, improve, and measure their enterprise data management function and staff." https://cmmiinstitute.com/products/cmmi/data The predecessor standalone Data Management Maturity model page now returns a 404 and the work is published under the CMMI Data banner at the URL above. (Note that the page itself is CMMI V2-era; it is cited for its own definition, not for the versioning.)
[5] ISO/IEC 38505-1:2017, Information technology — Governance of IT — Governance of data — Part 1: Application of ISO/IEC 38500 to the governance of data. https://www.iso.org/standard/56639.html Abstract read; the standard itself is paid and was not. Re-checked at source on September 4, 2026: the 2017 edition is now Withdrawn at stage 95.99, assigned August 20, 2026, and the catalogue page names ISO/IEC 38505-1:2026 as its published replacement (https://www.iso.org/standard/87195.html). The life-cycle record on the same page runs confirmation on June 24, 2022, stage 90.92 International Standard to be revised on June 9, 2023, and withdrawal on August 20, 2026. The 2026 edition is likewise paid and was not read, so the dates and status here are taken from the catalogue page and nothing in this article rests on the text of either edition.
[6] Mundy, Joy, "Design Tip #179 Key Tenets of the Kimball Method," Kimball Group, November 5, 2015. https://www.kimballgroup.com/2015/11/design-tip-179-key-tenets-of-kimball-method/
[7] Ross, Margy, "Design Tip #135 Conformed Dimensions as the Foundation for Agile Data Warehousing," Kimball Group, June 1, 2011. https://www.kimballgroup.com/2011/06/design-tip-135-conformed-dimensions-as-the-foundation-for-agile-data-warehousing/
[8] Gartner, "Lack of AI-Ready Data Puts AI Projects at Risk," February 26, 2025 (Q&A with Roxane Edjlali), https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk, and Sallam, Rita, "What Is AI-Ready Data? And How to Get Yours There," undated Gartner article, https://www.gartner.com/en/articles/ai-ready-data. Neither uses the phrase "master data management"; the absence was confirmed by full-body search of both pages. Note also that the February 2025 page carries two different sample sizes and fieldwork periods for the same figure — 1,203 data management leaders, July 2024, in the rendered body, and 248 leaders, third quarter 2024, in the page's meta description.
[9] EDM Council, "2026 Global Data Management Benchmark," public summary — 435+ organisations across more than 50 countries; "Approximately 31% of organizations surveyed report advanced data strategy capability, leaving most without the foundation required to support AI at scale." https://edmcouncil.org/innovation/research/benchmarks/ Full report gated; MDM is not named in the public summary.
[10] Wixom, Barbara H., and Cynthia M. Beath, "Let's Start Treating Data as a Strategic Asset!" MIT Center for Information Systems Research, Research Briefing No. XV-9, September 17, 2015 — n = 277, June 2015 poll. https://cisr.mit.edu/publication/2015_0901_DataStrategicAsset_WixomBeath
[11] Bean, Randy, "2026 AI & Data Leadership Executive Benchmark Survey," Data & AI Leadership Exchange, 2026 — the fifteenth annual edition; CDO appointment time series (Exhibit I), tenure (Exhibit M), and the culture-versus-technology impediment. https://static1.squarespace.com/static/62adf3ca029a6808a6c5be30/t/6942c3cb535da44088c2dbff/1765983179572/2026+AI+&+Data+Leadership+Executive+Benchmark+Survey+Final.pdf Methodology stated in the report: approximately 110 participating companies, 96% C-level, invitation-only non-probability panel, 88.2% North America, no response rate or questionnaire published. Note the attribution: this series was published by NewVantage Partners and then by Wavestone, but the 2025 and 2026 editions are neither. The absence of any mention of master data management across the 2025 and 2026 editions was confirmed by full-text search of both PDFs.
[12] Citigroup Inc., Annual Report on Form 10-K for the fiscal year ended December 31, 2024, under "Modernization" — Retired or replaced 714 legacy applications in 2024 with new, modern applications. The same filing reports transformation-related expenses of approximately $2.9 billion in 2024, up 1% year on year. https://www.citigroup.com/rcs/citigpa/storage/public/10K20250221.pdf Read from the filing itself, 2026-08-18.
[13] Citigroup Inc., Annual Report on Form 10-K for the fiscal year ended December 31, 2025 — continued to optimize, modernize and simplify Citi by retiring or replacing 548 applications during 2025 (representing 9% of all applications); transformation-related expenses increased 14% from the prior year to approximately $3.3 billion, largely driven by increased spending on data, as well as on controls. https://www.citigroup.com/rcs/citigpa/storage/public/citi-2025-10-k-2-20-26.pdf Read from the filing itself, 2026-08-18. Note what these figures are and are not: a count of applications retired, not an attribution of why any individual one was retired. Neither filing uses the word "decommission," and neither says a named system was retired to satisfy the order.
[14] Shaikh, Aziz, Holger Harreis, Jorge Machado, Kayvaun Rowshankish, with Rachit Saxena and Rajat Jain, "Master data management: The key to getting more from your data," McKinsey Digital, May 15, 2024. https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/master-data-management-the-key-to-getting-more-from-your-data Provenance worth stating: McKinsey-conducted and McKinsey-funded, sample given only as "more than 80 large global organizations" surveyed in 2023, with no exact n, response rate or published questionnaire. McKinsey sells MDM implementation consulting. It is the best available MDM-specific survey that is not from a software vendor, which is a comment on the evidence base.
[15] The "75% of MDM programs fail" figure, traced, and the trace is the point. It appears in Michelle Knight's write-up of a conference session by Amy Cooper, principal data management strategist at Dun & Bradstreet — but as Knight's own sentence, attributed to Gartner rather than to Cooper: "Common Master Data Management (MDM) Pitfalls," Dataversity, July 11, 2025 (modified March 13, 2026), https://www.dataversity.net/articles/common-master-data-management-mdm-pitfalls/. On that page the figure links to Gartner document 4009116, https://www.gartner.com/en/documents/4009116 — which is the Magic Quadrant for Master Data Management Solutions, published December 6, 2021 (Parker, Hawker, Walker), paywalled, and whose public abstract contains no such figure and no survey. The most relevant analyst-firm-independent MDM benchmark report remains TDWI's Next Generation Master Data Management, Q2 2012 — fourteen years old, and its own landing page lists IBM, DataFlux, Oracle, SAP and Talend as content sponsors, so "independent" is doing limited work even there.
[16] O'Neill, Brian T., "Failure rates for analytics, AI, and big data projects = 85% – yikes!" Designing for Analytics, first posted July 23, 2019 and maintained since. https://designingforanalytics.com/resources/failure-rates-for-analytics-bi-iot-and-big-data-projects-85-yikes/ The page documents the chain verbatim — "Nov. 2017: Gartner says 60% of #bigdata projects fail to move past preliminary stages. Oops, they meant 85% actually" — with "85% actually" hyperlinked to a November 2017 tweet by Gartner analyst Nick Heudecker. That link is the entire provenance of the figure. Two things this reference deliberately does not assert, because the cited page does not say them and neither could be confirmed from here: that the tweet has since been removed, and that no Gartner research note behind the 85% exists.
[17] Gartner, "Gartner Predicts 80% of D&A Governance Initiatives Will Fail by 2027, Due to a Lack of a Real or Manufactured Crisis," press release, Stamford, Conn., February 28, 2024, quoting Saul Judah, VP Analyst. https://www.gartner.com/en/newsroom/press-releases/2024-02-28-gartner-predicts-80-percent-of-data-and-analytics-governance-initiatives-will-fail-by-2027-due-to-a-lack-of-a-real-or-manufactured-crisis- This is a forward-looking prediction about governance initiatives, not a measured failure rate; the underlying research note is client-only.
[18] UK Government (Central Digital and Data Office), "Guidance on the Legacy IT Risk Assessment Framework," GOV.UK. https://www.gov.uk/government/publications/guidance-on-the-legacy-it-risk-assessment-framework/guidance-on-the-legacy-it-risk-assessment-framework The seven Likelihood criteria, in the framework's own order: L1 End of Life / End of Support, L2 Expired Vendor Contract, L3 Skills, L4 Business Needs, L5 Physical Environment, L6 Security Vulnerabilities, L7 Historical Issues. Cited for the ordering of its own criteria and for the absence of reorganisation, acquisition and leadership change across all thirteen criteria — not as a measurement of how often each cause actually retires a system, which this framework does not claim to be and which no located survey provides.
[19] Inmon, W. H., Building the Data Warehouse, 3rd ed. (John Wiley & Sons, 2002), ch. 1, "Evolution of Decision Support Systems" — the spider web, the 45,000 daily extracts, the naturally evolving architecture, and the crisis of credibility. Every quotation above was verified word-for-word against the third edition, whose edition statement and copyright were confirmed from the book's own front matter. No link is given on purpose: the full text circulates on unauthorised third-party hosts, and this publication does not link to them. Use a licensed copy or a library.
[20] Radcliffe, John, "Magic Quadrant for Customer Data Integration Hubs, 2Q07," Gartner, Inc., June 29, 2007, ID G00147231 — cited solely as a dated primary source for the category taxonomy and the 2007 definitions of MDM and CDI. No vendor placement from this document is used or endorsed. Both quotations were verified against the published document; the definition of MDM appears in its Note 2. No link is given on purpose: the copies in public circulation are third-party mirrors of licensed Gartner research whose own footer forbids reproduction. Cite it by title, document ID and date, as here.
[21] Smith, Heather A., and James D. McKeen, "Developments in Practice XXX: Master Data Management: Salvation Or Snake Oil?" Communications of the Association for Information Systems 23, art. 4 (2008), DOI 10.17705/1CAIS.02304. https://aisel.aisnet.org/cais/vol23/iss1/4/
[22] Hung, Patrick, and Graham Cedric Low, "Factors affecting the buy vs build decision in large Australian organisations," Journal of Information Technology 23 (2008): 118–131, published online May 15, 2007 — an interview study of ten large organisations whose literature review traces the canonical make-or-buy positions to Gremillion & Pyburn (1983), Martin & McClure (1983), Ceriello (1984), Davis (1988), Kelley (1992) and others. Open copy: https://www.cmu.edu/tcinc/students/course_documents/07/HW/Buy-vs-Build.pdf
[23] Microsoft, "Microsoft Purview and Profisee integration for master data management," Microsoft Learn, updated April 3, 2025 — "The Profisee MDM matching engine produces a golden record master as part of the survivorship process. Survivorship rules selectively populate the golden record with information that you've chosen across all your source systems." https://learn.microsoft.com/en-us/purview/data-governance-master-data-management-profisee Cited for what the component is and does, not as evidence of its merits. (Profisee's own product documentation is behind a login wall, so the vendor-authoritative description of these mechanics is not publicly citable.)
[24] Informatica, Multidomain MDM Configuration Guide — three pages, one per claim, because the claims live on different pages. Fuzzy matching "makes probabilistic match determinations": "Match Process," 10.3, https://docs.informatica.com/master-data-management/multidomain-mdm/10-3/configuration-guide/configuring_the_data_flow/mdm_hub_processes/match_process.html. Match tokens depend on the configured population ("Robert, Rob, and Bob in English speaking populations, for the match purpose of Name, may have the same match token value"): "Match Rules," 10.4, https://docs.informatica.com/master-data-management/multidomain-mdm/10-4/configuration-guide/part-4--configuring-the-data-flow/mdm-hub-processes/match-process/match-rules.html. Trust is enabled and configured per column on base objects and nowhere else: "Column Trust," 10.4, https://docs.informatica.com/master-data-management/multidomain-mdm/10-4/configuration-guide/part-4--configuring-the-data-flow/configuring-the-load-process/configuring-trust-for-source-systems/column-trust.html. Cited for documented product behaviour only.
[25] Microsoft, "Master Data Services Overview (MDS)," Microsoft Learn, updated June 23, 2026 — "Master Data Services (MDS) is removed in SQL Server 2025 (17.x). We continue to support MDS in SQL Server 2022 (16.x) and earlier versions." https://learn.microsoft.com/en-us/sql/master-data-services/master-data-services-overview-mds?view=sql-server-ver16 (The ver17 form of this URL redirects here; ver16 is canonical, which is itself the point — the page announcing the removal is served from the last version that has the component.)
[26] Mudgal, Sidharth, Han Li, Theodoros Rekatsinas, AnHai Doan, Youngchoon Park, Ganesh Krishnan, Rohit Deep, Esteban Arcaute, and Vijay Raghavendra, "Deep Learning for Entity Matching: A Design Space Exploration," Proceedings of SIGMOD 2018. https://pages.cs.wisc.edu/~anhai/papers1/deepmatcher-sigmod18.pdf The structured-versus-dirty finding is in the abstract; the small-label result is in the evaluation, and the exception matters, so here is the full sentence: The first three datasets (in Table 3) have only 450-946 labeled examples. Here DL performs worse than Magellan, except on Fodors-Zagats, which is easy to match. Three small datasets, deep learning worse on two — which is what the body says, and the reason the body says "two of the three" rather than "all three."
[27] Peeters, Ralph, Aaron Steiner, and Christian Bizer, "Entity Matching using Large Language Models," in Proceedings of the 28th International Conference on Extending Database Technology (EDBT 2025), 529–541, DOI 10.48786/edbt.2025.42 — reports transfer drops of "36 to 56% F1 for Ditto and 22 to 61% F1 for RoBERTa" on unseen entities. What the paper CONCLUDES, stated here so this entry cannot be read as support for the opposite: its abstract finds that the best LLMs require no or only a few training examples to perform comparably to PLMs that were fine-tuned using thousands of examples and that LLM-based matchers further exhibit higher robustness to unseen entities. The transfer collapse is the problem the paper sets out to solve, not its result, and the body says so. F1 is the standard accuracy score for this kind of task — the harmonic mean of precision and recall, where 100 is perfect. Preprint at arXiv:2310.11244, now served at v4 (October 18, 2024), which is the camera-ready and carries all three authors; the v1 preprint of October 17, 2023 had two, Peeters and Bizer. The figures quoted above are the camera-ready's, and they are not carried over from v1: the 2023 preprint reads ranging from 36 to 53% F1 for Ditto and 22 to 47% F1 for RoBERTa, so both upper bounds moved between versions. Cite the camera-ready, or say which preprint you mean.
[28] Vilminko-Heikkinen, Riikka, and Samuli Pekkola, "Master data management and its organizational implementation: An ethnographical study within the public sector," Journal of Enterprise Information Management 30, no. 3 (2017): 454–475 — a 32-month ethnographic study of two consecutive MDM development projects within a single municipality, identifying fifteen challenges, eight of them MDM-specific, none of them matching algorithms. The authors describe the work as a single qualitative case study, so it is cited as one organisation's experience rather than two. The MDM-specific challenges its abstract names are data owner and data definitions; "organizational implementation" is from the paper's title, not from the challenge list. https://researchportal.tuni.fi/en/publications/master-data-management-and-its-organizational-implementation-an-e/ (Repository record and abstract read; Emerald full text paywalled.)
[29] Webb, Thomas L. and Paschal Sheeran, "Does changing behavioral intentions engender behavior change? A meta-analysis of the experimental evidence," Psychological Bulletin 132, no. 2 (March 2006): 249–268, doi 10.1037/0033-2909.132.2.249. Forty-seven experimental tests, deliberately excluding the correlational designs that preclude causal inferences: participants are randomly assigned a treatment that moves intention, and behaviour is then measured. The finding, from the abstract: a medium-to-large change in intention (d = 0.66) leads to a small-to-medium change in behavior (d = 0.36) — move what people mean to do by a lot and you move what they do by roughly half as much. Record and abstract read in full at the University of Manchester's research repository; the published article is behind the American Psychological Association's paywall. https://research.manchester.ac.uk/en/publications/does-changing-behavioral-intentions-engender-behavior-change-a-me/ Cited for the general intention-to-behaviour relation, not for anything about corporations: the studies pool individual health and social behaviours, and the transfer to an organisation deciding a budget is an analogy, not a finding.
[30] Sheeran, Paschal and Thomas L. Webb, "The Intention–Behavior Gap," Social and Personality Psychology Compass 10, no. 9 (2016): 503–518, doi 10.1111/spc3.12265. Quoted from the abstract as published on the White Rose Research Online record: Bitter personal experience and meta-analysis converge on the conclusion that people do not always do the things that they intend to do. Note that the intention-behaviour gap and the willingness-to-pay bias in [33] are DIFFERENT constructs — they agree on direction, not on mechanism, and this article leans only on the direction. https://eprints.whiterose.ac.uk/id/eprint/107519/ The deposited full text is access-restricted, so this reference rests on the abstract, which is where the quoted sentence appears.
[31] Murphy, James J., P. Geoffrey Allen, Thomas H. Stevens and Darryl Weatherhead, "A Meta-analysis of Hypothetical Bias in Stated Preference Valuation," Environmental and Resource Economics 30, no. 3 (March 2005): 313–325, doi 10.1007/s10640-004-3332-z. Twenty-eight studies, eighty-three observations, all of them eliciting hypothetical and actual willingness to pay through the same mechanism. From the abstract: individuals are widely believed to overstate their economic valuation of a good by a factor of two or three, yet the median ratio of hypothetical to actual value is only 1.35, with severe positive skewness — stated intent overshoots dependably, by less than the folklore claims, with a long tail where it overshoots enormously. Abstract open on the publisher's page; full text paywalled. https://link.springer.com/article/10.1007/s10640-004-3332-z Included because it deflates rather than inflates the point it is cited for: the folklore multiple is two or three, the measured median is 1.35.
[32] Bean, Randy, "2025 AI & Data Leadership Executive Benchmark Survey," Data & AI Leadership Exchange, 2025 — n = 125; CDO tenure (Exhibit Q: "24.1% under 2 years, and 53.7% under 3 years") and the "created a data & AI driven organization" series (Exhibit H). https://static1.squarespace.com/static/62adf3ca029a6808a6c5be30/t/67642c0d40b42a7d7e684f49/1734618125933/2025+AI+&+Data+Leadership+Executive+Benchmark+Survey+120624.pdf
[33] Narayan, Avanika, Ines Chami, Laurel Orr, and Christopher Ré, "Can Foundation Models Wrangle Your Data?" PVLDB 16, no. 4 (2022): 738–746. https://www.vldb.org/pvldb/vol16/p738-narayan.pdf
[34] Office of the Comptroller of the Currency, Consent Order, In the Matter of Citibank, National Association, AA-EC-2020-64, October 7, 2020, Article V — the order cites deficiencies in its data governance, risk management, and internal controls that constitute unsafe or unsound practices. https://www.occ.gov/static/enforcement-actions/ea2020-056.pdf The quoted requirement is in the order's own operative language; PDF fetched and searched 2026-08-18.
[35] Office of the Comptroller of the Currency, "OCC Amends Enforcement Action Against Citibank, Assesses $75 Million Civil Money Penalty," News Release 2024-76, July 10, 2024 — the amendment is based on the bank's failure to meet remediation milestones and make sufficient and sustainable progress towards compliance with the 2020 Order, and the penalty is assessed for the bank's lack of processes to monitor the impact of data quality concerns on regulatory reporting. https://www.occ.gov/news-issuances/news-releases/2024/nr-occ-2024-76.html
[36] Board of Governors of the Federal Reserve System, "Federal Reserve Board fines Citigroup $60.6 million for violating the Board's 2020 enforcement action," press release, July 10, 2024 — Citigroup has made insufficient progress remediating its problems with data quality management and failed to implement compensating controls to manage its ongoing risk, with the two agencies' penalties totalling approximately $135.6 million. https://www.federalreserve.gov/newsevents/pressreleases/enforcement20240710a.htm
[37] Fraser, Jane, "Remarks by CEO Jane Fraser at Citi's 2025 Annual Stockholders' Meeting," Citigroup, April 29, 2025 (as prepared for delivery). https://www.citigroup.com/global/news/perspectives/2025/remarks-ceo-jane-fraser-citi-2025-annual-stockholders-meeting Quoted in full rather than trimmed to the convenient half, because it is the strongest statement against this section's reading and the reader is entitled to weigh it: This effort is broader than addressing the 2020 Consent Orders. It's fixing decades of underinvestment and ensuring Citi competes and leads in a digital-first world.