Weaving Intelligence

Before It Had a Name: How Master Data Ended Up at the Bottom of Every Governance Framework (original)

The 2026-08-28 original, preserved unedited for comparison.

Preserved original — 2026-08-28. This is the article exactly as it was written by Catherine D., before the voice layer existed. It is kept unedited so it can be read beside its replacement.

The rewritten version of this piece is at Before It Had a Name: How Master Data Ended Up at the Bottom of Every Governance Framework.

Why this exists: The Voice Problem →

The received history says a 2002 accounting statute created the pressure, and that master data management followed. My co-author was building the thing at the time, and he declines that story. Here are his words, the dates, and what happened when I went looking for evidence he was wrong.

Vertical: Enterprise Data Governance Angle: Historical Date: August 28, 2026

Open any current governance framework and master data is in it, usually near the floor, usually described as foundational. That placement looks obvious now. It was not always there, it did not arrive because a committee voted it in, and the story of how it got there is less tidy than the version that circulates.

This is a history piece, so a word about where it comes from. The spine is one practitioner's account — my co-author's, given in answer to questions about origins, and quoted as his where it is his. Around that I have put public sources: the statute, the vendor documentation, the academic framework, the trade press of the period. The sources are scaffolding. They are here so you can check the dates yourself, and in one place they made me correct him. Where that happened I have said so rather than smoothing it over.

The line I can't draw, and why that matters here

You cannot trace a relationship without saying what the relationship is, so let me get the awkward part out of the way. Ask my co-author where governance stops and master data management (MDM) begins and the first thing he does is concede that he blurs it himself, because he doesn't believe there is a hard line. Pressed, he draws one anyway: Data Governance manages the metadata surrounding Master Data Management while Master Data Management handles the underlying data. Governing, in that description, is the work of categorizing, classifying, writing the rules, and modeling the structures. And then the clean formulation: Governance is the rules and the framework within which MDM operates.

The academic version lands close enough to be worth putting beside it. Khatri and Brown, writing in Communications of the ACM in 2010, separate governance from management this way: governance is about what decisions must be made and who holds the right to make them, while management is making and implementing those decisions [3]. Their worked example is data quality — governance establishes who decides the standard, management picks the actual metric. That is the same cut, made in a different vocabulary.

Here is the part that matters for a history. Their framework has five decision domains: data principles, data quality, metadata, data access, and data lifecycle [3]. Master data is not one of them. Metadata is — which is precisely where my co-author puts governance's job. DAMA International's body of knowledge, the field's other durable framework, organizes data management into eleven core knowledge areas and does list reference and master data management, as area eight, sitting alongside data governance rather than underneath it [4].

So the modern picture — governance framework on top, master data as its foundational domain — is not what either canonical framework actually says. It is something practice arrived at. Which raises the reasonable question of where practice arrived at it, and when.

The practice arrived before the vocabulary

Ask my co-author where this work lived before anybody called it governance, and the answer is that it lived in three places at once. Data got cleaned during extract, transform and load — ETL, the movement of data out of source systems and into the warehouse. It got validated and translated inside the warehouse itself. And it got shaped again into specialized views and hierarchies in the reporting layer. Not a program. Three places where the problem showed up, each handled where it showed up.

Then this, which I am quoting rather than summarizing, because the summary loses it:

We were doing medallion builds and multi-system MDM integration with approval gates before we knew that’s what we were doing because it was the best way to meet the needs of the business.

He supplied the setting afterwards, and it is worth having, because before we knew that's what we were doing is a claim about a date. The engagement began in 1997 and ran ten years, at a publicly traded company. The build was a web-based document management system for the Engineering department that integrated with HR, Accounting, Sales and other departments' systems to produce current reports on factory upgrades, builds and installations. Read the nouns rather than the technology: cross-domain integration, an owning department that is neither the business nor IT, and reporting that has to reconcile across all of it. That is the whole shape of the discipline, five years before the statute that is supposed to have created it.

Two things to flag before that sentence can be used. First, medallion is an anachronism and a deliberate one. The bronze-silver-gold layering it refers to is lakehouse-era vocabulary — it describes raw ingestion, then cleaning and validation, then dimensional modeling and aggregation for business consumers [8]. That term did not exist when the work described here was done. Applying it backwards is useful, because it tells a reader in one word what the shape was. It is only honest if somebody says out loud that the label was fitted afterward, so: it was.

Second, the build had a name, and the name was architectural rather than disciplinary — a Kimball data warehouse, after Ralph Kimball's dimensional-modeling approach to warehouse design. And then the detail I would not have predicted: the Engineering department owned it, not the business and not IT—we were our own special animal. Two decades of governance literature is built on the question of whether the business or IT should own the data. His origin case answers neither, and he mentions it in passing, as a fact about the org chart rather than a position.

Now the obvious objection, because it is a good one. Saying nobody knew what they were doing is a retroactive-labeling claim, and the parts of that work were emphatically not nameless. Kimball had defined conformed dimensions precisely and publicly by 2003 — two dimensions conform when the fields used as common row headers are drawn from the same domain, and without that, cross-fact-table queries return garbage [5]. Cleansing had a name. Reference data management had a name. The vocabulary existed.

So the strong claim doesn't hold, and I am not going to defend it. The narrower one does, and it still forbids something. What had no name was the assembly: the layered, multi-system, gated pattern treated as one discipline, with its own tooling, its own roles, and its own line in a budget. The parts were named. The whole was not, and nobody was selling it as a whole. Kimball's own column makes that visible from the other side — conformed dimensions appear there as an argument about data warehouse architecture, inside a discussion of drilling across fact tables, not as a governance practice [5]. The idea was in the warehouse literature. The discipline had not been extracted from it.

Who named it, and the story I am not going to tell you

Somebody did extract it. Ask him who, and the answer is not an institution:

Companies started selling products to meet the market need before it had a name. Master Data Services, Master Data Maestro, Data Quality Services, and other products entered the market to meet the emerging demand and they needed something on which to base their marketing material. I don’t think there was a specific trigger I can say prompted it, though Sarbanes Oxley was contemporaneous with its emergence—which I think was more coincidence than coercion.

That makes two claims and only one of them is the interesting one. The first is about naming: the demand was already there, and vendors supplied the label and the frame because they needed something to point marketing at. The second is a refusal. Handed a menu of tidy causes, he declines all of them, including the one already on the table — and he hedges it twice, which I am keeping. Those hedges are doing real work; a causal negative is the hardest kind of claim to carry, and he carries it carefully.

It is also a dissent, so the received view deserves stating at full strength first.

The Sarbanes-Oxley Act of 2002 — the United States corporate accounting and auditing statute, enacted 30 July 2002 as Public Law 107-204, to improve the accuracy and reliability of corporate disclosures [1] — is, on the standard account, what made data accuracy an executive liability rather than an operational annoyance. And the standard account has contemporaneous evidence behind it, which is more than most received wisdom manages. Writing in Enterprise Systems in September 2003, thirteen months after enactment, Vinod Badami set out what compliance was going to demand of the underlying systems: consolidate and standardize financial data across the operational applications that feed it; provide the means to track and audit that data; give users the ability to standardize and integrate multiple versions of financial data to present a consistent picture; supply end-to-end metadata so the data's origins can be traced [2]. That is a specification for master data management, published in 2003, by someone who does not use the term. I originally wrote because it was not the available term, and my own timeline further down this page refutes that: the phrase was in an SEC filing six months before Badami published. What is defensible is the absence itself — neither master data management nor even master data appears anywhere in his piece — not an explanation for it. He calls the answer business intelligence and data management, and locates it inside what he calls a framework of financial governance [2].

The academic literature carries the same association. Khatri and Brown open by noting that data governance had recently gained prominence at practitioner conferences in light of both the opportunity to leverage data assets and the need to ensure compliance with mandates such as Sarbanes-Oxley and Basel II [3]. Two drivers, given equal billing, in a peer-reviewed venue in 2010.

So here are the dates, which is what you actually need to form your own view. The statute is July 2002 [1]. The demand-side argument is in print by September 2003 [2]. Microsoft acquired the vendor whose product became Master Data Services in June 2007, and announced in May 2009 that it would ship in SQL Server 2008 R2 [6]. Profisee's Master Data Maestro, built on top of Master Data Services by the team that originally wrote it, was in the market and being evaluated by working practitioners by early 2013 [7]. Data Quality Services arrived with SQL Server 2012 [9]. Seven years, minimum, between the statute and the first of the three products he names. The gap is real, and so is the contemporaneity across the middle of the decade, which is exactly why coincidence and causation are hard to separate here.

What happened when I tested him

A claim like his has an obvious test, so I ran it. If compliance created the category, the earliest launch and marketing material for those products would lead with compliance and name the statute. Go read it.

It partially fired, and against him. Microsoft's launch-era post announcing Master Data Services asks why master data management had suddenly become a top-three item for chief information officers, and answers with what it calls a perfect storm: economic downturn, corporate information ecosystem complexity, strict governmental regulations such as HIPAA and Sarbanes-Oxley, service-oriented architectures, and the recognition that solving the problem with custom code is extremely difficult [6]. The statute is there, by name, in the vendor's own first pitch. His claim as stated — contemporaneous but not causal — does not survive that intact, and I would rather correct him in public than build a section on a sentence I had not checked.

He came back on that, and his answer is better than my correction. His point is not that the statute went unmentioned. It is that the work was already underway before the regulation existed: business had identified the gap and was building to close it, and when the rules arrived they calcified a direction the market was already moving in. Then vendors shipped products into the newly mandated demand — products that must have been in development for years to have arrived in the shape they did. In his phrase, the regulation may have been the capstone, but the pebbles were already in motion.

That is a claim with a date on it, so I went and checked it rather than taking either of us on trust. It holds, and the evidence is not thin. In November 2001, nine months before the statute, Madnick, Wang, Dravis and Chen published a worked account of an engagement delivering a single view of customers — match rules, consolidation rules, confidence thresholds, and cross-functional sign-off on what a customer even was [11]. That is master data management and its governance, executed commercially, before the vocabulary. Two years earlier Ralph Kimball was describing a standing conforming meeting that a CIO should personally attend, because agreeing a definition of customer across an enterprise is more political than technical [12] — a data stewardship forum under another name, in 1999. Seven weeks before Sarbanes-Oxley was signed, the grocery industry had a live product-data registry with standardized attributes, subscription fees reaching six figures, and a Wal-Mart supplier mandate behind it [13]. And in February 2002 an industry survey of 647 practitioners was already costing bad data at $600 billion a year and using the phrase single version of the truth, with six commercial data-quality vendors paying for the research [14].

So the pebbles were in motion, demonstrably. Now the part that looks like it cuts the other way, and the reason it does not. The term itself has a traceable arrival date, and the shape of that arrival is exactly what a regulation-caused-it story predicts. Searching the full text of every electronically filed SEC document, the phrase master data management appears zero times in 2001 and 2002. It first appears in March 2003, in SAP's annual filing, and by that October SAP is describing it to investors as a new offering. Then: three filings in 2003, fifteen in 2004 across five companies, forty-one in 2006 across fifteen [16]. The takeoff begins immediately after the statute and steepens for four years.

Except that curve measures when a name entered corporate English, and the claim under test is about a practice. Those are different objects, and I had them confused: a category being named and sold after a statute is entirely compatible with the work predating it, which is precisely his thesis rather than a rival to it. The regulation may well have forced the industry to adopt a name. That is not the same as the regulation having caused the thing the name refers to.

Which suggests the query that would actually discriminate. If the statute raised the commercial salience of the practice rather than just supplying its label, then the language for the underlying capability — the words available before anyone coined the category — should surge alongside the name. So I ran the same method on capability language [16]. It does not surge. Single view of the customer runs 17 documents in 2001, 5 in 2002, 9 in 2003, 17 in 2004 — and normalized against filing volume it falls after 2001 and never recovers its own starting rate. Customer data integration is flat across 2001–2004 (20, 20, 26, 25) and only moves in 2006, three years after the name arrived. The enterprise-data-management share of data quality is flat from 2001 to 2006 — its raw growth turns out to be mining assay reports and index boilerplate. Data warehouse declines on a normalized basis.

And the capability language is real where it appears, which I checked at sentence level rather than by counting. Acxiom's 2001 annual report sells Customer Data Integration as a named product category that lies at the core of effective CRM, providing a single view of the customer, in real time, across multiple data sources — two years before master data management appears anywhere in the corpus. In July 2001 a property-casualty insurer wrote into an inter-company cost-sharing agreement that its platform provides a single view of the customer across the enterprise in support of all customer processes and touchpoints [16]. That is a buyer, not a vendor, putting the capability in a contract before the statute.

Three limits, because a result this tidy deserves them and one of them is severe. First, the volume is small and vendor-concentrated: every 2001 hit for customer data integration traces back to four filers, fifteen of the twenty being Acxiom filing about itself. Early is not the same as widespread. Second, two phrases I would have guessed were older turn out to be younger: single version of the truth is absent in 2001 and 2002 and first appears in 2003, the same year as the name, and data governance is absent through 2004 with a single document in all of 2006. Third and most important: SEC full-text search only covers filings from 2001 onward [16], and a corpus of filings is a record of what became worth saying to investors, not a record of what practitioners were doing. The Kalido case makes the seam visible: the release describes the suite as a data warehouse and master data management software portfolio and the underlying platform as supporting master-data-aligned warehouses — and that platform was already in production at Shell, Unilever and Philips. A module can be built in fifteen months; the master-data capability it was packaging cannot have started after July 2002 — and the same press release puts improved corporate transparency in its subhead and reaches for proactive management of regulatory compliance in its body [15]. The engineering predates the regulation. The marketing does not. That distinction is the whole of it.

What survives is narrower and, I think, more useful. Compliance was named as one pressure among five, in a hedged sentence beginning with the admission that the answer lies somewhere in the list. It was not the frame. The frame of that entire post is operational: master lists that get corrupted in myriad ways across divisions and locations, and four forces acting on them — decay, conflict, corruption, inconsistency — against which the goal is not truth but an authoritative source of master data [6]. The investment case offered is cost savings and revenue recovery, not audit survival.

That narrower claim forbids something specific, so I tested it too: no launch or positioning material for the other two products should lead with compliance either. Microsoft's introduction to Data Quality Services frames the business need as incorrect data arising from entry errors, transmission corruption, mismatched dictionary definitions and the aggregation of sources using different standards; compliance appears once, as one consequence among several alongside lost credibility, lost revenue and unhappy customers, with no statute named [9]. The most detailed contemporaneous account of Master Data Maestro I could find — a practitioner's, written in March 2013 by someone using the products daily — runs to a capability-by-capability comparison covering stewardship interfaces, matching strategies, survivorship, golden records, address verification and a software development kit, and does not mention compliance once [7]. Both hold.

The honest summary: compliance was in the room and named out loud in the vendor's own first pitch, so coincidence is too strong a word for what happened. But it was never the frame these products were sold on. Every one of them was pitched on operational consolidation — one authoritative version of the customer, the product, the part — and that is the frame the category inherited. I would rather hand you the dates and the actual sentences than a verdict.

One standard for the whole enterprise, and what it cost to learn better

Master data does not become the foundation of a governance framework because a framework says so. It becomes the foundation when somebody tries to run the framework without it and finds out. My co-author's version of finding out is the most useful paragraph in this piece, partly because it is at his own expense:

Much to my chagrin, early on I tried to shoehorn the entirety of the Enterprise onto a single standard.

What broke it was not architecture. It was, in his words, the amount of pushback we got—and continue to get to make downstream reports match what individual departments expected to see. Note the tense. That pressure did not stop when he changed his approach; it is a condition of the work, managed rather than solved, and any account that reports it as resolved has flattened the only honest thing about it.

The correction he made is narrower than the failure suggests, and this is where I have watched people take the wrong lesson. He did not conclude that enterprise standards don't work. He remains, as he puts it, still a champion of a single, Enterprise-wide ‘standard’ for attributes like client type, product line, inventory category. What he conceded was the presentation layer: the flexibility to show alternatives to particular segments of the business, for backward compatibility or for a department's own reporting.

That concession has two possible readings and they are not equally defensible, so I will say which one I am taking. The defensible reading is a governed alias: one authoritative underlying value, with multiple sanctioned labels or roll-ups presented over it. The other reading — genuinely different values in different places, unreconciled — is just local standards with better manners, and it is the thing he says he was trying to eliminate. I read him as meaning the first, because it is the only reading under which a single enterprise-wide standard is still true as written, and because a concession that dissolves the standard it claims to preserve is not a concession, it is a retreat with good PR. If I have him wrong, the fault is my inference and not his sentence.

He confirmed the reading and then narrowed where it holds, which is the more useful answer. It is right wherever categorization can happen through direct matches at the same grain and the same category. Where it stops being right is cross-hierarchy references and rollups: when all zits are zats but not all zats are zits, the rollups stop being trustworthy, and there is no clean way around it. Concretely: a member that belongs to two parents double-counts unless somebody picks a primary, a member with no parent at the reporting level falls out of the total silently, and a ragged hierarchy makes the sum at one level disagree with the sum at the next. None of those announce themselves — the report renders, the numbers are wrong, and the variance is small enough to be argued about for a quarter. And the cost is not technical. A business that has categorized by zits for decades gets very cross when it is told to move to the zat paradigm — which is the same political problem as the definitional one, arriving one level up and with a longer institutional memory behind it.

And that is the mechanism by which master data ends up at the bottom of the framework. Not because a body of knowledge listed it. Because the standard is the only thing that survives the pushback, and the standard has to sit on something. Governance without an authoritative underlying value is a policy about a value nobody agrees on.

What twenty years of frameworks did not touch

Ask what all that framework-building has failed to solve and the answer comes back in three words: politics and personality conflicts. Which sounds like a grumble until the mechanism arrives with it. Territoriality persists because budget, area of responsibility and domain of control remain welded to position and power. Nothing has changed that, so nothing has changed the behavior it produces. Then the sentence I would put on a wall:

While on paper each SVP is a peer to each other SVP, their various domains and areas of responsibility, not to mention the size of their departmental budgets, established a very real power hierarchy.

SVP is senior vice president, and the observation is that the org chart shows a flat row of them while the actual hierarchy is set by budget and scope. That gap is checkable on your own org chart this afternoon, which is more than most observations about organizational politics can say for themselves.

The generic version of this — governance is a people problem, not a technology problem — is thirty years old and it wastes the answer. His claim is narrower and structural. Governance asks executives to give up control over the exact things their standing is measured in. That is a statement about incentives, and incentives imply a remedy that training and communication cannot reach. You either change what confers power, or you route around it. The diagnosis has been published a thousand times; that implication is the part worth taking away.

Now, nothing has changed is an unqualified universal across two decades, and there is a real candidate against it. Data mesh, as Zhamak Dehghani set it out in 2019, does not negotiate over domain ownership — it assigns it, arguing that domains should host and serve their own datasets rather than flowing them into a centrally owned platform, and naming the loss of domain ownership after ingestion as the specific failure of the monolithic model [10]. That is a formal answer to exactly the territorial problem described above — and my co-author's response to it was not that it is wrong but that it is his: the federated ownership model he had already been running. Which is this article's own thesis arriving one more time, in a place I had not planned for it. The formal answer published in 2019 is a practice somebody was operating before it had that name either.

So the claim narrows, and here is what the narrowed version forbids. It concedes that the formal structure has changed — ownership can now be assigned on paper, and often is. It holds that the informal one has not. What that forbids is straightforward: it forbids evidence that assigning domain ownership on paper measurably reduces territorial conflict. I went looking for that evidence and did not find it. Dehghani's article is an architectural argument, not a study of organizational behavior, and it makes no claim about conflict rates [10]. I could not locate a measurement of territorial conflict before and after a domain-ownership assignment in any source I opened. That is an absence of evidence rather than evidence of absence, and I would rather leave it standing as an open question than close it with a framework. If you have run that comparison, I would genuinely like to see it.

The naming is happening again

One observation before I hand this over, and I will keep it short because it is an aside rather than the argument.

The pattern traced above is not finished. It is running right now around AI governance: the demand is genuine, the category is being assembled in front of us largely by people with something to sell, and it is acquiring vocabulary faster than it is acquiring practice.

The useful thing to take from that is not cynicism. Someone has to name a category, and the products that carried master data management into the enterprise were good products. It is that the name arrives after the work and is shaped by whoever needs to sell something. So when the AI governance framework lands on your desk with its domains neatly laid out, ask the question this history answers: which parts of this were people already doing, in three places at once, because the business needed them? Those are the parts that will hold.

What I Would Watch For

Practitioner layer — the curator's read on the consensus above. Governance is my lane; master data is my co-author's, and I have kept my judgement to where the two meet.

The failure mode I'd watch hardest

A framework that lists master data as a foundational domain and changes nothing about who decides what a customer is. Listing is free. The whole force of Khatri and Brown's framing is that governance is about decision rights and the locus of accountability [3] — so the test of whether master data is genuinely foundational in your framework is not whether it appears in the diagram. It is whether anyone can name the person who decides the enterprise definition of client type, and whether that person's decision survives contact with a department that wanted a different one. If the answer is that the framework document was updated, you have a domain on a slide. Eight strong rowers all pulling hard on their own count still lose.

The trade-off that usually bites

The one my co-author walked into, and it bites in both directions. Hold the single standard rigidly and you will spend your political capital on report-formatting arguments you cannot win, because the departments are not being difficult — their numbers genuinely have to reconcile to things they have published before. Concede too broadly and you have local standards wearing an enterprise badge. The workable middle is narrow and it is a governance decision rather than a technical one: one authoritative value, with sanctioned alternate presentations that are themselves governed and documented as derivations. If nobody can tell you which value is the authoritative one and which is the alias, the middle has already collapsed.

The claim I'd be sceptical of

Any clean causal history of this discipline, including the one I have just partly defended. The received account — regulation forced it, MDM followed — is not fabricated; it has real contemporaneous evidence, in a 2003 trade article that specified the requirement before the vocabulary existed [2], and in the vendor's own naming of the statute six years later [6]. But a category emerging across a decade, in a market with a downturn, a complexity problem, an architectural fashion and a compliance regime all pressing at once, is overdetermined by construction. Anyone selling you the single trigger is selling you a story. My co-author's real contribution here is not that he identified the right cause; it is that he declined to name one at all, which is the harder and rarer move.

Where to Go Deeper

In the order I would read them. Khatri and Brown's Designing Data Governance [3] first — it is five pages, it is open access, and its governance-versus-management distinction will do more for your programme than a maturity model will. Then DAMA International's framework [4], for the eleven knowledge areas and specifically for how it positions reference and master data management relative to governance. Kimball's 2003 column on drilling across [5] is twenty-three years old, free, and the clearest short statement of conformed dimensions anyone has written; read it to see how much of what we now call governance was already solved inside warehouse design. For the vendor side of the history, Microsoft's 2009 announcement post [6] is worth reading whole, because you can watch a category being framed in real time. And Dehghani on data mesh [10] for the strongest modern argument that ownership should be assigned rather than negotiated — I am unconvinced it changes the behavior, and it is the best case against my scepticism.

What I would take from it

The repeated version of this history has a statute at the front and a straight line running out of it. The version the sources support has several pressures arriving together across a decade, practices already named individually inside warehouse design, and vendors supplying a name because a name is what you build marketing on. Master data reached the bottom of the framework by holding weight, not by being listed.

If that sounds like a demotion of governance, it isn't. It is the opposite. Anybody can put a domain on a diagram. What is hard, and what nothing in twenty years of frameworks has made easier, is getting a room of nominal peers to accept a definition that costs one of them something. That was the job before it had a name, and it is the job now.

References

[1] Sarbanes-Oxley Act of 2002, Public Law 107-204, 116 Stat. 745, enacted 30 July 2002 (H.R. 3763), U.S. Government Publishing Office — cited for the enactment date, the public law number, and the Act's stated purpose, “To protect investors by improving the accuracy and reliability of corporate disclosures made pursuant to the securities laws.” I read the enacting clause, the short title and the table of contents, not the full statute; nothing in this article rests on the text of a particular section. https://www.govinfo.gov/content/pkg/PLAW-107publ204/html/PLAW-107publ204.htm

[2] Vinod Badami, “Sarbanes-Oxley: The Role of Business Intelligence and Data Management,” Enterprise Systems, 17 September 2003 — a trade article written thirteen months after enactment, by the then national director for business intelligence at an IT professional services firm. Source of the consolidation, tracking-and-auditing, multiple-versions-standardization and end-to-end-metadata requirements, and of the phrase “framework of financial governance.” Cited as evidence of what was being asked for in 2003 and in what vocabulary, not as an independent finding. Read in full 2026-08-22. https://esj.com/articles/2003/09/17/sarbanesoxley-the-role-of-business-intelligence-and-data-management.aspx

[3] Vijay Khatri and Carol V. Brown, “Designing Data Governance,” Communications of the ACM 53, no. 1 (January 2010): 148–152 — source of the governance-versus-management distinction adapted from Weill and Ross, the data-quality worked example, the five decision domains (data principles, data quality, metadata, data access, data lifecycle), and the opening sentence associating data governance's rise with both data-asset opportunity and compliance mandates including Sarbanes-Oxley and Basel II. Read in full 2026-08-22 via the open-access edition. https://cacm.acm.org/research/designing-data-governance/

[4] DAMA International, “DAMA-DMBOK Framework — Core Knowledge Areas” — the association's own current statement of its framework, cited for the eleven core knowledge areas and for the position of Reference & Master Data Management as area eight, listed alongside Data Governance as area one. This is the framework as published today; the page does not carry edition dates and no dated claim is made from it here. Read 2026-08-22. https://www.damadmbok.org/copy-of-about-dama-dmbok

[5] Ralph Kimball, “The Soul of the Data Warehouse, Part Two: Drilling Across,” Kimball Group, April 2003 — source of the precise definition of conformed dimensions (“Two dimensions are conformed if the fields that you use as common row headers have the same domain”), the consequence of attempting a drill-across on unconformed dimensions, and the framing of the whole discussion as data warehouse architecture. Read in full 2026-08-22. https://www.kimballgroup.com/2003/04/the-soul-of-the-data-warehouse-part-two-drilling-across/

[6] Kirk Haselden (SQL Server Team), “Master Data Services – What's the big deal?”, Microsoft SQL Server Blog, 13 May 2009 — written by the then product unit manager for Master Data Services. Source of the June 2007 Stratature acquisition, the announcement that the product would ship as Master Data Services in SQL Server 2008 R2, the four forces (decay, conflict, corruption, inconsistency), the preference for “authoritative source of master data” over “one source of the truth,” and the “perfect storm” sentence naming economic downturn, ecosystem complexity, HIPAA and Sarbanes-Oxley, service-oriented architectures and the difficulty of custom solutions. Cited as vendor positioning — evidence of how the category was framed, not evidence of the product's value. Read in full 2026-08-22. https://www.microsoft.com/en-us/sql-server/blog/2009/05/13/master-data-services-whats-the-big-deal/

[7] James Serra, “Profisee Master Data Maestro,” jamesserra.com, 7 March 2013 — a contemporaneous practitioner account by a then-independent data warehouse consultant and Microsoft SQL Server MVP, writing while using Master Data Services and Data Quality Services daily. Cited for three things: that Master Data Maestro is Profisee's and is built on top of Master Data Services; that Profisee's principals originally built the product Microsoft acquired from Stratature in 2007; and for the content of its capability comparison, in which compliance does not appear. This is a personal blog, not a peer-reviewed or editorially controlled source, and it is used here as period evidence of how the product was described rather than as an independent verification of any vendor claim. Read in full 2026-08-22. https://www.jamesserra.com/archive/2013/03/profisee-master-data-maestro/

[8] Microsoft, “What is the medallion lakehouse architecture?”, Azure Databricks documentation, Microsoft Learn — cited for what the bronze, silver and gold layers denote (raw ingestion; cleaning and validation; dimensional modeling and aggregation) and for the description of the pattern as a data design pattern for progressive refinement. Used only to establish what the term means and where it comes from, which is what makes its use in this article an identified anachronism. Read in full 2026-08-22. https://learn.microsoft.com/en-us/azure/databricks/lakehouse/medallion

[9] Microsoft, “Introduction to Data Quality Services,” Microsoft Learn (SQL Server documentation) — the vendor's own documentation, cited for what Data Quality Services is and does (knowledge-driven cleansing, matching, reference data services, profiling, monitoring), for its arrival as a SQL Server feature installed with the product, and for the framing of “The Business Need for DQS,” in which compliance appears once as one consequence of incorrect data among several. The page's article metadata carries an original date of 5 March 2012, consistent with the SQL Server 2012 release. Read in full 2026-08-22. https://learn.microsoft.com/en-us/sql/data-quality-services/introduction-to-data-quality-services

[10] Zhamak Dehghani, “How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh,” martinfowler.com, 20 May 2019 — cited for the argument that domains should host and serve their own datasets rather than flowing them into a centrally owned platform, and for the identification of lost domain ownership after ingestion as the failure of the monolithic model. My reading was targeted rather than complete: I read the sections on centralized data platforms and on domain-oriented decentralized data ownership, which are the passages the claim rests on, and did not read the article end to end. Read 2026-08-22. https://martinfowler.com/articles/data-monolith-to-mesh.html

[11] Stuart Madnick, Richard Wang, Frank Dravis and Xinping Chen, “Improving the Quality of Corporate Household Data: Current Practices and Research Directions,” Proceedings of the Sixth International Conference on Information Quality (ICIQ), November 2001, 92–104 — the load-bearing pre-statute source in this article, and the reason the “pebbles were already in motion” claim is checkable rather than a memory. Source of the eleven-step commercial engagement delivering a single view of customers, the match and consolidation rules stored in a rules matrix, the confidence thresholds, and the cross-functional definitional sign-off (“The terms had different meanings to different people”). What this paper is NOT cited for, because an earlier draft got it wrong: its ABI/INFORM literature-search sentence is about the authors' own proposed coinage corporate householding, and it reports that no corresponding concept existed — which is close to the opposite of the concept existed and had no name. The engagement itself is the evidence here; that sentence is not, and a blind reader caught the draft that said otherwise. Read in full 2026-08-25 from MIT's own Total Data Quality Management publications archive; dated November 2001 there and corroborated by internal citations to 2001 work. https://web.mit.edu/tdqm/www/tdqmpub/CorporateHouseholdNov01.pdf

[12] Ralph Kimball, “The Matrix,” Intelligent Enterprise column, December 1999, in the Kimball Group archive — source of the “conforming meeting” as a standing forum that is “probably more political than technical” and that a CIO should personally attend, and of the framing of a common definition of customer as “a major litmus test for an organization.” Read in full 2026-08-25. A date caveat, because this is a history piece and the date is the point: the page displays December 7, 1999 beneath the byline and encodes it as 1999-12-07T22:40:26-06:00, so the date is attested to the day. What it is not is a 1999 artifact: the same head carries a modified time of 2016-01-26, and the markup is a modern content-management system. So this is a migrated record rather than a primary one — good enough to date the column, not the same thing as holding the December 1999 issue. Read in full 2026-08-25. https://www.kimballgroup.com/1999/12/the-matrix/

[13] Brian Sullivan and Michael Meehan, “UCCnet's Promise: Synchronized Product Data,” Computerworld, 10 June 2002 — cited for a live product-data registry with 62 standardized attributes, subscription fees reaching $400,000 a year, named participants including Procter & Gamble, and a Wal-Mart supplier mandate, all seven weeks before Sarbanes-Oxley was signed. Note the structural point it carries: the forcing function here was a retailer's mandate, not a statute. Read in full 2026-08-25. The date is publisher-asserted in the page metadata of a modern re-publication of a 2002 article rather than an original artifact. https://www.computerworld.com/article/1328390/uccnet-s-promise-synchronized-product-data.html

[14] The Data Warehousing Institute, press release for Wayne W. Eckerson, Data Quality and the Bottom Line, 1 February 2002 — cited for the $600 billion annual cost figure, the 647-respondent survey base, the circulation of the phrase “single version of the truth,” and the existence of six paying commercial data-quality sponsors, six months before the statute. The REPORT itself was not read — the PDF would not retrieve — so every figure here is cited to the dated press release, which was read in full 2026-08-25. Also worth recording because it cuts against this article: the same release notes that almost half of respondents had no current plans to improve data quality, so practice existed but adoption was shallow, which leaves real room for a post-2002 forcing function. https://insurance-canada.ca/2002/02/01/the-data-warehousing-institutes-recent-study-finds-high-quality-data-is-critical-to-the-success-of-businesses-worldwide/

[15] Kalido, “Kalido Launches Data Warehouse Lifecycle Management Software Suite,” press release, London, 27 October 2003, carried by Enterprise Systems Journal — cited for a master data management module at beta stage in October 2003, sitting on a warehouse platform already in production at Cadbury Schweppes, HBOS, Intelsat, Royal Dutch/Shell, Philips and Unilever. That timeline is the “years in development” evidence, and the release's subhead reaching for “improved corporate transparency” — with “regulatory compliance” appearing once in the body, at about the same depth as the vendor documentation this article scores as NOT leading with compliance — is the evidence that marketing followed the statute where engineering could not have. Page 1 read in full 2026-08-25; page 2 not reached. Dateline, URL path and on-page stamp agree on the date. https://esj.com/articles/2003/10/27/kalido-launches-data-warehouse-lifecycle-management-software-suite.aspx

[16] U.S. Securities and Exchange Commission, EDGAR full-text search — queries run 2026-08-25 for the exact phrase “master data management” across all electronically filed documents, returning zero filings for 2001–2002, three in 2003 (all SAP AG, the earliest being SAP's Form 20-F filed 21 March 2003), fifteen in 2004 across five filers, and forty-one in 2006 across fifteen. The SAP 6-K exhibit filed 16 October 2003 describing “Master Data Management (SAP MDM), a new offering” was read in full. Also cited for the capability-language queries run the same day, which are the discriminating test: exact-phrase counts for 2001, 2002, 2003, 2004 and 2006 on single view of the customer (17 / 5 / 9 / 17 / 10), customer data integration (20 / 20 / 26 / 25 / 66), single version of the truth (0 / 0 / 8 / 2 / 3), data governance (0 / 0 / 0 / 0 / 1), plus data quality, reference data and data warehouse. Normalization uses non-ownership filing counts derived from EDGAR's quarterly form.idx full-index files, because raw totals are inflated 3.12× by Forms 3/4/5 after Sarbanes-Oxley section 403 mandated electronic Section 16 filing — tiny XML documents that can never contain these phrases. Stripping them, real growth is 1.71×. Source of the Acxiom 10-K passage (accession 0000733269-01-500016, filed 2001-06-27) and the Erie Indemnity cost-sharing exhibit (accession 0000922621-01-500002, filed 2001-07-19), both opened and read rather than counted. Three limits recorded because they bound everything above: coverage begins in 2001 per the SEC's own FAQ, so the zero says nothing about the 1990s; the endpoint counts documents while the denominator counts filings, so normalized figures overstate growth by an unmeasured factor; and reference data proved unusable as a proxy at all, its early counts dominated by a Boston commercial lease form whose opening article is headed “REFERENCE DATA.” https://www.sec.gov/edgar/search/