Weaving Intelligence

We've Made This Case Before: How the Enterprise MDM Business Case Grew Up (original)

The 2026-08-17 original, preserved unedited for comparison.

Preserved original — 2026-08-17. This is the article exactly as it was written by Elias K., before the voice layer existed. It is kept unedited so it can be read beside its replacement.

The rewritten version of this piece is at The Case Was Already Assembled.

Why this exists: The Voice Problem →

Two decades of arguing for master data management — the emphasis kept moving, the arguments never did, and what actually gets one funded never has.

Vertical: General MDM Angle: Historical Date: August 17, 2026

Every couple of years someone slides a fresh business case for master data management across the table and asks whether this is the one that sticks. I read it the way you'd hear a song you're sure you know: new arrangement, same melody.

I've been reconciling the spreadsheets since before the work had a name, when "master data" was just the quiet fact that every report showed a different version of the truth. In that stretch the pitch has been rewritten at least four times, each sure it had found the problem fresh. I grew up around Pittsburgh mills, where every generation was certain its trouble was unprecedented until the next one inherited it.

Diagram — Three assembled, one still open: four lanes of argument for the master data business case, three settled and one still running, each carrying its case in a single line. Every word in the picture is in the caption below.
Figure 1: The fourth lane is drawn unclosed deliberately — the argument it carries is the one still being settled, and its second half belongs to somebody else.

Three assembled, one still open. Compliance, cost and capability were one case by February 2002. The fourth opened later, and its second half belongs to somebody else.

Two dated markers stand near the start, drawn as dashed lines falling through everything after them: Feb 2002, TDWI [6] — the cost case already in print; and Jul 2002, Sarbanes-Oxley §404 signed. The cost case is in print five months before the statute, which is the order this piece argues.

Four lanes run left to right beneath the markers. Each is named, carries its era underneath, and holds one highlighted block: the argument in a sentence. The block sits further right in each lane down the stack, which is how the picture shows the spotlight moving while every lane stays open.

  • Compliance & Consolidation — Sarbanes-Oxley, ERP/CRM sprawl, mergers. Its block reads "Reconcile it before the auditor does." It sits at the far left, and the lane runs from the left edge: the argument was there from the beginning.
  • Cost & ROI — Quantified bad-data waste — $12.9M/yr per org. Its block reads "The mess is already costing us."
  • Business Outcomes — Digital CX, e-commerce, data as an asset. Its block reads "Buy this so the good things become possible."
  • AI-Ready Data — Analytics and generative AI — still open. Its block reads "Your AI is only as trustworthy as the master data under it." This lane alone begins two thirds of the way across, and its right edge is dashed rather than closed: it started late, and it has not finished.

One band runs the full width beneath all four, headed Unchanged underneath every one of them: Own the domain. The golden record is a verb. A tended discipline, not a finished project.

The compliance case: reconcile it before the auditor does

Master data was an established idea long before anyone sold master data management — MDM, once it reached the slide. Master files and master records were ordinary vocabulary in ERP and in warehouse practice through the 1990s. The named discipline crystallized in the early 2000s, out of fear.

Two forces put it there. The first was accumulation: a decade of ERP and CRM rollouts, plus mergers bolting whole companies together, left the same customer, product and supplier in a dozen systems that quietly disagreed — the problem the discipline exists to answer, and why it defines itself around a single trusted version of them [1]. David Loshin, in one of the first guides to carry the name, framed it through the acronyms that caused it: as ERP, SCM and CRM spread, there emerged "a need for a consolidated view of high-quality representations of every critical instance of a business concept" [2].

The warehouse people hit the same wall from the other side and called it a conformed dimension — one agreed description of a thing, defined once and reused by every report that touches it. Departmental data sets, the Kimball Group warns, "may look like they can be compared and integrated due to similar labels, but the underlying business rules may be slightly different" [3]. Two traditions, one complaint.

The second force was louder, and it came with a date attached. Sarbanes-Oxley arrived in 2002, and its Section 404 directed the SEC to require management to assess and report on internal control over financial reporting [4]. It did not land overnight — the first §404 reports weren't due until late 2004 — giving organizations two years to discover, in writing, that "our systems disagree about who the customer is" was now an audit finding with an executive's signature on it.

Then the money moved, which is what turns a regulation into a budget line. The SEC's own economists put mean annual §404 compliance cost, for the larger companies whose controls an outside auditor had to attest to, at $2.87 million a year before the 2007 reforms eased the testing rules and $2.33 million after. Nobody spent that because the return looked good, and the survey says so: large majorities credited §404 with improving their internal control structure (73%) and their audit committee's confidence (71%); barely half credited it with improving the reporting itself (49%); and most of those answering the optional cost-benefit question called it negative [5]. The arithmetic said no. The work happened anyway.

So the first business case was defensive. You didn't buy master data management to grow; you bought it to stop bleeding. Underneath the architecture diagrams, those programs were redding up — a Pittsburgh word for putting a place back in order — decades of mess nobody wanted to own. A case built on somebody else's deadline buys you a project, and leaves the discipline unbought.

The ROI case: putting a number on the mess, and what came with it

Compliance gets you through the door once, and the usual telling has the argument growing up into money after that. I told it that way myself. It's wrong, and the document that refutes it is the one I was about to lean on.

The number that anchored the early decks landed in February 2002 — five months before Sarbanes-Oxley was signed, and three years before the first §404 report came due. The Data Warehousing Institute told readers that data quality problems were costing U.S. businesses more than $600 billion a year. Two pages later the same report gives a precise figure for poor-quality customer data — $611 billion, in postage, printing and staff overhead — and hangs the only derivation in the document off it [6]. The report never reconciles the two, and the most economical reading is that they are one estimate, the rounder number written for the executive summary. Take that reading or leave it; what you cannot do is treat the pair as two independent findings that corroborate each other, which is how it has circulated for twenty years.

That report is the money argument already finished, not an early sighting of it. Beside the headline figure sits a worked funding case — $130,000 in annual savings on a $70,000 outlay, a 188 percent internal rate of return, payback in months — respondents naming "a single version of the truth" among the benefits they got, and a warning that HIPAA and the Bank Secrecy Act were "upping the ante" on customer data [6]. Three of my four arguments, in one document, before the statute that supposedly opened the first.

So these were never stages. Compliance, cost and capability are three faces of one case that was already assembled by February 2002, and the two decades since have re-weighted them rather than added to them — whichever driver carried that year's deadline went to the front of the deck. We have made this case before, all of it at once. It didn't grow up. It got re-argued. I take them here in the order they are easiest to explain, which is not the order they arrived — they mostly arrived together.

I raise it with affection rather than to score off a twenty-year-old report: it showed its arithmetic. A footnote derives the estimate from savings reported by respondents who had cleaned up name-and-address data, scaled by Dun & Bradstreet counts of U.S. businesses by headcount [6]. You can disagree with that method; more to the point, you can check it.

The genre it launched is still running, the slide still built the same way with a fresher figure. Gartner puts the cost of poor data quality at an average of at least $12.9 million a year per organization [7]. Different decade, identical rhetorical job — the emphasis had moved from "the auditor will punish us" to "the mess is already costing us, today."

Quantified waste is a more honest argument than vague risk. It is also where the industry picked up the habit it has never shaken: treating the headline number as the argument rather than the doorway.

The third face is capability, and it is the one that flatters us most. Digital made a single view of the customer worth money at the top of the income statement rather than savings at the bottom — you cannot cross-sell on four contradictory versions of one person — and e-commerce turned a bad product master into a broken catalog and a lost sale. Later, too: a 2013 design-science paper from the University of St. Gallen's corporate data quality program still lists regulatory compliance, reporting "in the sense of a 'single version of the truth'," and the demand for a 360-degree view of the customer as concurrent drivers [8]. Eleven years on, the same three, still concurrent — confirmation rather than news; the 2002 document had already settled it.

The underlying bet does appear to pay, at one remove: Brynjolfsson, Hitt and Kim, from survey data on 179 large publicly traded firms, reported output and productivity 5–6% higher among adopters of data-driven decision-making than their other investments would predict [9]. That is a finding about deciding from data, not about master data management, and I'll leave it there. The claim that buyers themselves moved from exposure to capability is one I've watched happen and cannot source for you: the figures that circulate trace back through a trade write-up [10] to a 2021 Magic Quadrant [11] — Gartner's ranked survey of vendors in a category, the artifact a market uses to tell itself who is winning — which, read in full, doesn't contain them.

And capability carried a tax. The harder you sell the platform as the enabler of every new initiative, the easier it is to assume the software carries the ownership work. The pitch got more exciting and the discipline got easier to skip.

The AI case: mostly the same work, on a louder deadline

Those three are settled material, and I've given them the space settled material deserves: their arguments are made and their bills are known. What follows gets more room than all three together — because it's the only one still being argued, and priced right now in proposals I'm reading this month. The weight belongs where the thinking isn't finished.

Disclosure, before I argue this next part. The practice behind this publication is a Profisee implementation partner — certified to deploy that product for clients. Those clients pay us for the deployment work; the vendor does not pay us for the sale. That is the narrow truth, and here is the wider one: our services revenue is downstream of this category of software being chosen at all. What follows argues that entities should be resolved deterministically and upstream rather than left to the model — which is an argument for buying the category we are paid to deploy. You should read the rest of this section knowing that, and weigh the criticisms in it above the praise.

Everyone now needs data that is "AI-ready" — and a model is only as trustworthy as the master data it stands on. Feed one four versions of your biggest customer and it answers with the serene confidence of one.

But look hard at what Gartner's own account of AI-ready data actually asks for. Reliable sources and pipelines. Lineage — "transparency about data origins and transformations." Stewardship policies applied "throughout the data life cycle." Validation and verification "during development and operations." Observability of data health, timeliness and accuracy [12]. Most of it is the master-data business case — the same governance work the previous three drew from. AI didn't invent the argument; it inherited it and stapled a deadline to it.

Notice what is not on that list: nowhere does it ask you to resolve those four versions into one. That absence is where the bill comes due.

Most of them, and the remainder deserves naming: it genuinely was not in the first three. Gartner's own account declines to call any dataset AI-ready in the abstract: readiness "depends on how the data will be used," and there is "no way to make data AI-ready in general or in advance" [12]. Stack on the labeling, the versioning against model drift, the regression testing and the bias measures, and you have real work, none of it master data management and none handed to you by a golden record.

There is a sharper point against me in the same place. High-quality data, judged the traditional way, "does not equate to AI-ready data," because an algorithm needs data representative of what it will meet — and that "may include poor-quality data, too" [12]. Survivorship is the work I have spent a career defending, and for a training set it is also a way of deleting the very variance the model was supposed to learn from. The golden record is the right artifact for a payment run and the wrong one for a sample. That doesn't undo the case; it means this fourth argument has a second half that belongs to somebody else.

Where I want to push back is on the bill, and I owe this one to my co-author, who put it more plainly than I had:

I haven't read the Gartner report, but any report that says 'AI ready data' is (functionally) unchanged from your source systems isn't asking for any work to 'prepare' the data for AI. They're just asking that the AI be fed garbage data and expected to correct it. That's expensive in tokens when compared to basic MDM pre-processing.

He flagged that he hadn't read it, and he's right to say so — and having read it, I owe the page more than he gives it. It asks for real work, as that list shows; "(functionally) unchanged from your source systems" is too strong. What survives is the half of his objection aimed at what the page never asks for, and it survives because the page reaches further than he knew. Where it argues representativeness at length, the case is about training: "when training an algorithm, the algorithm will need representative data" [12]. Its definition, though, goes past its argument. Data must be representative "to train or run an AI model for a specific use" [12]. Run, not just train. And the enterprise AI work I actually see is running, not training — inference: a model reading your systems to answer a question. At inference nothing is being fitted, so the four versions of your biggest customer teach nobody anything. They arrive in the context window, and something has to reconcile them — every query, at token prices, forever — or not reconcile them and answer confidently from whichever one came first.

Notice what the page does that its market does not: it defines its terms. What gets sold downstream of it is "advanced AI processing," and I have yet to be handed two proposals that meant the same thing by it. That is my experience rather than a census — nobody has surveyed the vocabulary, and I am not going to pretend the absence of a survey is evidence of one. But the asymmetry survives either way, because it only takes one undefined term to do the damage: you cannot cost an undefined thing, scope it, or hold anyone to it — so what arrives dressed as a comparison between two approaches is a priced line item on one side and a phrase on the other. Ask for the definition before the quote.

That is the trade I would put in front of a committee, and the bill is only its most-argued half. Resolve the entity once — upstream, or in the view that serves the model, but once and deterministically — and you pay for it once, you can read the rule that decided it, and you can change it and see what moved. A rule is an artifact you can read, version and argue with; a reconciliation performed inside a context window leaves no such artifact — you can prompt it differently, but you cannot diff what changed, and nobody governs what they cannot see. That is control, and worth more than the saving.

The obvious objection is that prompts version too, traces log, and an eval harness will diff a whole run for you. All true, and none of it is the same artifact. A trace tells you what went in and what came back; it does not tell you which rule decided these two records were one person, because no rule did — the decision happened inside the answer and left no separable account of itself. An eval scores the aggregate and stays silent on the individual merge. What a survivorship rule gives you is a decision you can read before it runs and argue with without running it. That is what a joint gives a limb, in the anatomical sense: motion, load, and both only within a range you can name. Unbounded motion is dislocation.

Answer quality is the other half, and my co-author has the image for it — an analogy, not an anecdote. Train a model on raw sewage and you do get a lift; the lift is real. But the defense of feeding it the mess rests on the mess being representative of what the model will meet, and sewage is not representative of drinking water. That is the narrow claim I'll defend: not that a resolved entity always answers better, but that "representative" has not earned its keep here — an unresolved customer record is not a truer sample of your business, it is the same fact four times, weighted by whichever system was chattiest.

Ship the mess downstream and call it representativeness, and you have converted a fixed preprocessing cost into a variable per-query one, on a unit price you neither set nor forecast. I can't hand you the crossover figure — it depends on your query volume, your model and a token price that moved twice while this piece was edited, and I'm not going to invent one to win a paragraph. But the shape of the two costs is different, and a case that ignores the difference is not a case about money at all.

Two objections belong here, this being the part likeliest to be wrong. The page is not silent on cost: it asks that data meet operational service levels "including response time and cost efficiency" [12]. That names the bill without saying who pays it down, but it is there, and I'd rather quote it than pretend otherwise. The second is aimed at me. A model handed all four versions can notice they disagree; a survivorship rule cannot — it picks one, records the win, and the disagreement stops being visible downstream. Reconciling per query is expensive, but expensive in the open, and if your merge logic is quietly choosing wrong that may be the cheaper error. The answer is that a conflict surfaced fresh on every query is still a conflict nobody owns.

There's a wrinkle the enthusiasm hides: a lot of AI spending isn't a return-on-investment argument at all. In the ones I've sat in, the return is negative and the room knows it; the case is a response to market and competitor pressure — a legitimate reason to spend money, and a different argument. Regulatory work has always lived there, and the §404 figures above measure exactly that: real reported benefits, a majority verdict that the trade-off was negative, the program going ahead anyway [5].

Read that as a caution about sequence rather than a swipe at AI: you can't automate your way around an ownership problem you never solved. The machine will reproduce the mess faster, with a straight face.

From the Field

Everything above this line is the field's argument and my quarrel with it — sourced where a source exists, flagged where one doesn't. Everything in this section is operating experience: how these cases get built, argued, funded and stalled. The register changes with it — essay to briefing — and the rest is written for whoever has to carry one into a room next month. Briefing prose states things flat, but these are patterns rather than laws: regularities from the rooms I happened to be in, and I cannot hand you a way to falsify them.

Hits

  • There are four steps, and the two that decide it are the two you aren't picturing. I put the question to my co-author, who has carried more of these into more rooms than I have. What came back wasn't a template but a sequence:

    (1) Figure out what you need to do, then — usually immediately after — what you want to do. (2) Figure out roughly how long it will take and what it will cost. (3) Shape the case to appeal to the needs of the company and of the specific members of the approval group. (4) Last thing in is the candy: low-hanging fruit that would be great to get done, doesn't move the timeline much, and can win support from business groups outside the core ask.

    Three and four are where cases are won, and the two that get cut when the calendar tightens. But hold onto the separable pieces: committees rarely say yes or no, they say yes to some of it, and a case built in pieces recomposes around that answer — what got funded is your minimum viable product, what got cut is your backlog, already scoped and costed.

  • Size it before you shape it. The first question is whose problem you are solving, and the honest range runs from keeping a capable team busy to a statute that just cleared Congress. Then you put a rough order of magnitude around it, or, if the calendar is against you, a scientific wild-ass guess. Either answers the same question — what order of magnitude are we arguing about — and tells you how much argument you owe.
  • Tailor per approver, the way an appellate advocate works a bench. A lawyer before a high court doesn't write one argument; they know which members are moved by which reasoning. A funding committee runs on the same mechanism. Give marketing a reason of its own to say yes; with nothing in it for marketing, marketing is a coin flip. And keep legal in view: new exposure that isn't a core requirement or a regulatory response draws pushback unless you arrive answering it.
  • For the CFO there are exactly two numbers, and nothing to add. Cost, and anticipated return — the only two he cares about. Sometimes the return is negative and the program's job is to minimize it; say so plainly, because the room can tell. What matters more is whether the figures tell a story a reasonable person can follow. They get the case read; they don't decide it. He holds one seat; the seats beside him ask whether this organization can absorb the work, who owns it after handover, and what it displaces on a plan already full.
  • Run a proof of concept down the entire line from day one, end to end. Not the risky component — the whole path, however thin. All three stall shapes below yield to the same move: turn the crank early, so the obstructions surface while there's budget and calendar left to route around them. The least glamorous advice in the discipline, and genuinely hard: speed-running the project without burning the contingency, on a team good enough that your best people can trail-blaze while the rest implement with little direction.

Misses

  • Skipping steps three and four together. The least experienced managers — and plenty of experienced ones under pressure — focus on the deliverable and never build the case around it. Asked why those two go first, my co-author named the assumption underneath: The work being obviously worth doing is not an argument; it's an assumption that the room already agrees with you. It usually doesn't — or it agrees, and funds someone who didn't skip those steps.
  • Filing exogenous change as somebody else's problem. Funded programs stall in one of three shapes, and naming them matters: the mitigations differ. The first is scope creep, which almost never means a stakeholder wanting extra features: it is requirement change arriving from outside — restructuring, architecture, regulation, security. You scoped on a platform that was a reasonable call at the time; three months later a researcher publishes a zero-day and security scraps it inside the week. Or a new director arrives and the AWS shop is a GCP shop by next quarter, yours joining the migration list mid-build. I used to file those as external events rather than scope creep. That was wrong: each propagates, with downstream effects on all in-flight work. If it changes what you must build, it's scope creep whatever the org chart calls it.
  • Planning around a feature that hasn't shipped. The second shape is incorrect technical assumptions, and it catches the teams who like new technology: hype advertises niche capabilities, and projects get planned around something six months from shipping. Connecting system A to system B with a newly announced ETL layer? Prove it has connectors for both, inside your window, before the plan hardens. This bites hardest when the platform lands in phase two: nobody inspects what it delivers until the plan carries weight.
  • Reading dependency drag as bureaucracy. The third shape is the approval gate that convenes intermittently — architecture review, security sign-off, data classification. It isn't red tape: those boards are staffed from other active projects, and the advisory seat isn't anyone's day job. A gate that slips slips the dependency, and the dependency slips the program — same when a production problem takes the team for three days. Complaining in a status report has never moved a date.
  • Not knowing where the hard part starts. Months four to six is the inflection point. The low-hanging fruit is gone and the thorny work is all that's left, which is why the second and third increments are the rough ones: the team that had just found its rhythm goes back to arguing about basics.

The Unwritten

These recur constantly and seldom survive the trip into a best-practices document, because admitting them looks bad.

  • The headline number is an extreme, and everyone quoting it knows that. It's a worst case or a best case depending on who is holding it up; your organization lands somewhere between, and no published figure tells you where. Good for getting attention in the room, useless as a forecast of what you'll actually recover — quote it to open a conversation, never to close one.
  • A threshold is a fact about the organization that set it, and only incidentally about money. I had this written as though the figure decided. My co-author took it apart:

    While there's always a number, I'll fight about the number ever being the only factor. It's always one of a bunch of factors, even if it's the 'leading' one. Even in situations where companies establish 'thresholds,' those thresholds are established for other reasons than pure monetary figures—not the least of which might be they don't have that kind of money in the first place. Obviously if someone needs a trillion dollars to implement something, anyone other than Governments or Musk are going to reject it out of hand, but that's not because of the dollar figure—it's because of the infeasibility for them being the ones to do it in the first place.

    A case refused for being enormous was almost never refused on the digits; it was refused because this body is not the kind that can do that thing — not the balance sheet, not the appetite, not the people. Which matters: "we can't afford it" is a wall, and "we are not the ones to do this, this year" is a shape you can argue with — resize it, resequence it, find the sponsor whose mandate it falls under. The number tells you which conversation you're in; it never has the last word.

  • The number that kills a project is the one with nothing underneath it. The quickest route to a tabled proposal is a committee member asking where a figure came from and getting silence. Not a weak answer — no answer. And what dies is the room's willingness to believe everything in the case that had nothing to do with money — the timeline, the staffing, the claim you understand the systems. Every non-financial assurance is re-read as decoration. So tie the program to a number the business already tracks and show your work, as that 2002 report showed its. Twenty years on, the industry still hasn't copied it.
  • Ownership is something a program reveals, not something it solves — and I lean hard on the first verb. Solving anything takes buy-in, support, effort and, above all, enforcement. Build all the intelligence you like into a system; if the business has to route around it to keep the doors open, it will, and rightly. I've never seen an organization solve systematizing its own delivery; it only relocates the problem up or down the chain. Layers all the way down. Outsourcing infrastructure was the last instance; AI is the next. Very little of this survives into a business case, because a case has to promise an end state and there isn't one.
  • I'm not going to give you a number from my own work, and the reasons are the point. This is where such a piece usually produces a personal figure — an efficiency gain, a percentage. I don't have one. I'm rarely in the financial decisions, and hard financials sit behind security, privacy and data classification. Much of what I've done created process where none existed, so there's no A/B to compare against. And the assertion would be unfalsifiable: you could not check an order-of-magnitude efficiency claim without me handing over proprietary and customer-identifying information, which I won't do. If I can't make the case that I'm worth the money without leaning on ROM and SWAG figures that are, let's be honest, guesses with a decimal point, then I'm probably not the right fit.

What held, whichever argument was in front

Back to the argument. Stand back far enough and the durable lessons were sitting there the whole time.

  • It's an organizational problem wearing a technology costume. Every version of this case has blamed the tools, and every version has been mostly wrong. Master data breaks in the org chart, not the diagram — as I learned watching two databases everyone swore were spotless merge into one memorable mess. The clean data model saved nobody; the unanswered question of who decides did the damage. It's the one lesson the outside sources reach independently — the St. Gallen researchers, naming early programs as technology-driven at the cost of organizational work [8], and the Kimball Group, from data modeling, concluding that shared reference data "requires organizational consensus and commitment to data stewardship" [3]. Everyone who gets there gets there by exhaustion.
  • The golden record is a verb, not a noun. Survivorship runs forever, because the world keeps changing your customers, products and suppliers whether you watch or not.
  • Budget for the tending; the build is the cheap part. I keep bees, and here is the part that maps: a colony does its own work. It forages, raises brood and gets through most weeks without me. What it cannot do is decide when to split, when to feed, or whether the queen is failing — and a hive nobody decides for looks perfectly well from the outside right up until it is gone.

The question I'm handing off

Which leaves the question I'll hand off rather than settle: how much of this can a platform carry? Tomorrow, Isabel takes the same business case into a real MDM platform — Profisee — showing what it looks like once the argument stops being a slide and becomes software. The disclosure above covers her too, and it is why she can write about that platform at the depth she does: she knows it from the inside.

She and I disagree, amiably, about exactly this seam. She's made her peace with letting a good platform shoulder a real share of the load, and she's usually right to; I get twitchy trusting any tool to save me from a decision the organization hasn't made. But we've never disagreed about what held through all of it. A good platform carries the work, and it will carry it beautifully while the ownership question underneath goes unanswered — showing you nothing at all until the day it shows you everything.

Master data isn't a project you finish. It's a hive you keep.

Correction. Updated August 22, 2026. An earlier version of this piece called the 2002 TDWI report's two figures "a subset larger than the whole" — $611 billion for poor-quality customer data set against "more than $600 billion" for all data quality problems. That was wrong on its face, because $611 billion is more than $600 billion, and the sentence made the report look self-refuting when it is not. The passage has been rewritten. Two things we then got wrong in the fix and are correcting here rather than leaving quiet: the replacement dropped the report's own word customer, which is the scope difference that makes the pair confusing in the first place; and it asserted the two figures are one estimate as though the report said so. It doesn't. That is our reading, and it now says it is one.

Why we missed it. Our blind reviewers were handed the article's text without its sources. Two of them read that passage and neither could settle it, because the evidence sat outside the text and the instrument could not reach it — a reviewer who cannot open a citation cannot catch a misread one. Reviewers now get the full reference list with working links, are told to open whatever they doubt, and open each figure's image file rather than our description of it. The first reviewer equipped that way found both of the errors named above, in the correction itself.

What the byline means. Elias K. is an AI persona; the argument and the prose are his. The operating experience in From the Field is not. It comes from Jeff Shabel — drawn out in interview before this was written, and sharpened by the corrections he made after reading it. Every passage quoted here is his own words. He edited the result.

References

[1] DAMA International, DAMA-DMBOK: Data Management Body of Knowledge, 2nd ed. (Technics Publications, 2017), ISBN 978-1-63462-234-9, ch. 10, "Reference and Master Data" — the field's reference work on master-data domains, the trusted source and golden record, and the single-version-of-the-truth framing.

[2] David Loshin, Master Data Management (Morgan Kaufmann / The MK-OMG Press, 2008), ISBN 978-0-12-374225-4, Preface, pp. xix–xx. The quoted sentence appears in the Preface, which also names "increased regulatory oversight, increased need for information exchange, business performance management, and the value of service-oriented architecture" as the drivers converging on master data management — a contemporaneous statement of the compliance case by a practitioner writing at the time. Read from the publisher's own sample front matter. https://booksite.elsevier.com/samplechapters/9780123742254/Sample_Chapters/01~Front_Matter.pdf

[3] Margy Ross, "Design Tip #135: Conformed Dimensions as the Foundation for Agile Data Warehousing," Kimball Group, June 1, 2011 — defines a conformed dimension as "descriptive master reference data that's referenced in multiple dimensional models," warns that similar labels can mask differing underlying business rules, and states that defining one "requires organizational consensus and commitment to data stewardship." https://www.kimballgroup.com/2011/06/design-tip-135-conformed-dimensions-as-the-foundation-for-agile-data-warehousing/

[4] Sarbanes-Oxley Act of 2002, Pub. L. No. 107-204, tit. IV, § 404, 116 Stat. 745, 789 (July 30, 2002), codified at 15 U.S.C. § 7262. https://www.law.cornell.edu/uscode/text/15/7262 The statute directs the SEC to write the rule; the operative requirement and the term of art “internal control over financial reporting” come from U.S. Securities and Exchange Commission, "Management's Report on Internal Control Over Financial Reporting and Certification of Disclosure in Exchange Act Periodic Reports," Final Rule, Release Nos. 33-8238; 34-47986; IC-26068, adopted June 5, 2003, effective August 14, 2003; 68 Fed. Reg. 36636. https://www.sec.gov/files/rules/final/33-8238.htm

[5] Office of Economic Analysis, U.S. Securities and Exchange Commission, Study of the Sarbanes-Oxley Act of 2002 Section 404 Internal Control over Financial Reporting Requirements, September 2009 — mean total Section 404 compliance cost of $2.87 million pre-reform and $2.33 million post-reform among Section 404(b) filers (p. 4); reported benefits of 73% (internal control structure), 71% (audit committee confidence), 49% (financial reporting quality) and a majority-negative assessment of the overall cost-benefit trade-off (p. 6). https://www.sec.gov/news/studies/2009/sox-404_study.pdf

[6] Wayne W. Eckerson, Data Quality and the Bottom Line: Achieving Business Success through a Commitment to High Quality Data, TDWI Report Series, The Data Warehousing Institute, February 2002 — "more than $600 billion a year" in the Executive Summary (p. 3); "$611 billion a year in postage, printing, and staff overhead" for poor-quality customer data, with the derivation given in footnote 1 as respondent-reported cost savings from name-and-address cleanup scaled by Dun & Bradstreet counts of U.S. businesses by number of employees (p. 5). The two figures are not reconciled in the report; both are quoted here as published. http://download.101com.com/pub/tdwi/Files/DQReport.pdf

[7] "Data Quality: Why It Matters and How to Achieve It," Gartner — poor data quality costs organizations at least $12.9 million a year on average (Gartner research, 2020). https://www.gartner.com/en/data-analytics/topics/data-quality

[8] Andreas Reichert, Boris Otto and Hubert Österle, "A Reference Process Model for Master Data Management," Wirtschaftsinformatik Proceedings 2013, paper 52 (11th International Conference on Wirtschaftsinformatik, Leipzig, 2013), pp. 817–830 — business drivers and the technology-driven-to-organizational shift are stated in §1.1, pp. 817–818. Produced by the Competence Center Corporate Data Quality at the University of St. Gallen. https://aisel.aisnet.org/wi2013/52/

[9] Erik Brynjolfsson, Lorin Hitt and Heekyung Kim, "Strength in Numbers: How does data-driven decision-making affect firm performance?" ICIS 2011 Proceedings, paper 13 — 179 large publicly traded firms; output and productivity 5–6% higher among adopters of data-driven decision-making than their other investments and IT usage would predict. Cited to the published abstract on the publisher's repository record; the deposited PDF carries no extractable text layer, so this reference does not rest on a full-text read. https://aisel.aisnet.org/icis2011/proceedings/economicvalueIS/13/

[10] Thor Olavsrud, "What is master data management? Ensuring a single source of truth," CIO, May 31, 2021 — the trade write-up the figures travel through. It reports that "organizations pursue MDM for a variety of reasons; among the most popular are to create internal/operational efficiencies (69%), to improve business process outcomes (59%), and to improve business process agility (54%), according to Gartner’s Jan. 2021 MDM Magic Quadrant." Page read in full 2026-08-17. https://www.cio.com/article/191827/what-is-master-data-management-ensuring-a-single-source-of-truth.html

[11] Gartner, Magic Quadrant for Master Data Management Solutions, 27 January 2021, ID G00466922 — the document those figures are attributed to. A full-text search of the complete licensed-for-distribution reprint (1-253RHWIZ, 30pp) returns no occurrence of 69%, 59% or 54%, and the document does not discuss reasons for pursuing MDM at all; the percentages it does contain are 35, 5, 80, 95, 40, 16, 21, 1, 7, 6 and 28. Gartner has retired the document and its public landing page now redirects to the Gartner home page, so there is no live URL to give you — cited by title, date and document ID so that anyone with Gartner access can check the claim. Verified 2026-08-17.

[12] Rita Sallam, "What Is AI-Ready Data? And How to Get Yours There," Gartner — the source of both quoted sentences: "There is no way to make data AI-ready in general or in advance. The readiness of data for AI depends on how the data will be used," and "‘High-quality’ data — as judged by traditional data quality standards — does not equate to AI-ready data… the algorithm will need representative data. This may include poor-quality data, too." The same page's requirement list runs to semantics and labeling, quantification, diversity, versioning for model drift, continuous regression testing, observability, and bias and fairness. Read in full, in a rendered browser, 2026-08-18. https://www.gartner.com/en/articles/ai-ready-data