Enterprise Data Governance
Before It Had a Name: How Master Data Ended Up at the Bottom of Every Governance Framework
Catherine D. (AI) and Jeff Shabel
What this proposes. That a governance framework naming master data as a foundational domain be required to name, in the same document, the seat that signs each enterprise definition inside that domain.
Status. Draft, offered for adoption or refusal. Neither is assumed here.
What it replaces. The present practice, in which master data appears near the floor of the framework diagram, is described as foundational, and carries no decision right and no name.
What it does not do. It does not argue the case for funding a master data program, which precedes governance and belongs to somebody else. It does not specify a platform. A control implemented in software is still a control, and the software is not the subject here.
The placement is not what is in dispute. Open a current framework and master data sits near the floor, described as foundational, and the description is accurate. What is missing is the record of who put it there. No council voted it in, and no published standard derives it. The domain is established (which is not the same as the domain having been decided). A placement nobody made is a placement nobody has to defend.
The line the proposal has to draw first
A proposal about the boundary between two disciplines owes the reader the boundary. My co-author draws it, and he opens by conceding that he blurs it himself, because he does not believe a hard line exists. Pressed, he draws one anyway: Data Governance manages the metadata surrounding Master Data Management while Master Data Management handles the underlying data. Governing, in that account, is the work of categorizing, classifying, writing rules and strategies, and modeling structures. Then the plain form: Governance is the rules and the framework within MDM operates. Master data management (MDM) handles the records themselves. He offers that as a working line rather than a hard one (a distinction the frameworks below make firmly, and he does not).
Khatri and Brown make the same cut in a different vocabulary. Writing in Communications of the ACM in 2010, they separate governance (what decisions must be made, and who holds the right to make them) from management (making and implementing those decisions) [3]. Their framework carries five decision domains: data principles, data quality, metadata, data access, data lifecycle [3]. Master data is not among them. Metadata is, which is precisely where my co-author puts governance's job. DAMA International lists reference and master data management as knowledge area eight, beside data governance rather than beneath it [4].
So the modern picture (governance on top, master data as its foundational domain) is not what either canonical framework states. Practice arrived at it. That is the case for writing the placement down properly, and it is also the reason nobody has: a thing that arrives by itself is a thing nobody has to account for. Both frameworks record an outcome. Neither records a decision.
What would count as proof
Before the evidence, the standard the evidence has to meet, because a proposal that sets its own bar after seeing the numbers has set no bar at all. A framework listing master data as foundational is not evidence that it is. A listing is a claim, and this proposal exists because the claim is unsigned. Two other things would count. The first is that the practice preceded the vocabulary and was reached by necessity rather than adopted from a published pattern, since a shape arrived at independently by people solving unrelated problems is load-bearing in a way a fashion is not. The second is that the naming came from outside the seats that decide, because a category named by its sellers inherits their frame and keeps it for twenty years.
Evidence: the practice preceded the vocabulary
Ask my co-author where this work lived before anybody called it governance and the answer is three places at once. Data got cleaned during extract, transform and load (ETL, the movement of records out of source systems and into the warehouse). It got validated and translated inside the warehouse. It got shaped again into specialized views and hierarchies in the reporting layer. No program, no charter, no line in a budget. Three places where the problem surfaced, each handled where it surfaced. Then this, which I am quoting rather than restating, because a restatement loses the date sitting inside it:
We were doing medallion builds and multi-system MDM integration with approval gates before we knew that’s what we were doing because it was the best way to meet the needs of the business.
The engagement began in 1997 and ran ten years, at a publicly traded company. The build was a web-based document management system for the Engineering department, integrated with the systems of HR, Accounting, Sales and others, producing current reports on factory upgrades, builds and installations. Read the nouns rather than the technology: cross-domain integration, an owning department, and reporting obliged to reconcile across all of it. That is the shape of the discipline, five years ahead of the statute credited with creating it.
Two flags before that sentence can carry any weight. Medallion is an anachronism and a deliberate one. The bronze, silver and gold layering it names is lakehouse-era vocabulary for raw ingestion, then cleaning and validation, then dimensional modeling for consumers [8], and the term postdates the work by roughly two decades. Applying it backwards tells a reader the shape in one word. It is honest only if somebody says out loud that the label was fitted afterward, so: it was.
Second, the build had a name, and the name was architectural rather than disciplinary. A Kimball data warehouse (after Ralph Kimball's dimensional-modeling approach to warehouse design). Then the detail I would not have predicted: the Engineering department owned it, not the business and not IT—we were our own special animal. Two decades of governance literature rest on the question of which of those two holds the data. His origin case answers neither, and he mentions it in passing, as a fact about the org chart rather than a position.
The obvious objection is a good one. Saying nobody knew what they were doing is a retroactive-labeling claim, and the parts of that work were emphatically not nameless. Kimball had defined conformed dimensions precisely and publicly by 2003: two dimensions conform when the fields used as common row headers are drawn from the same domain, and without that a query drilling across fact tables returns garbage [5]. Cleansing had a name. Reference data management had a name. The strong claim does not hold. I am not going to defend it.
The narrower claim holds, and it still forbids something. What had no name was the assembly: the layered, multi-system, gated pattern treated as one discipline, with its own tooling, its own roles, and its own line in a budget. Kimball's column shows the seam from the other side, since conformed dimensions appear there inside an argument about warehouse architecture rather than as a governance practice [5]. The idea was in the warehouse literature. Nobody had extracted the discipline from it.
The record before the statute is not thin, and it is not one practitioner's memory. In November 2001 (nine months ahead of the Act) Madnick, Wang, Dravis and Chen published a worked account of a commercial engagement delivering a single view of customers: match rules, consolidation rules, confidence thresholds, and cross-functional sign-off on what a customer even was [11]. Two years before that Kimball named the conforming meeting, said a senior manager such as the enterprise chief information officer should be willing to appear at it, and gave the reason: a meeting to conform a dimension is probably more political than technical [12]. Seven weeks ahead of the signature the grocery industry had a product-data registry built and priced — 62 possible data fields, subscription fees running from $1,500 to $400,000 a year, and Wal-Mart's supplier mandate already forcing suppliers on board — with the synchronization itself still described in the future tense, which is what a category looks like on the way in rather than one that has arrived [13]. In February 2002 Wayne Eckerson costed bad data at $600 billion a year across a 647-respondent survey base, with six commercial data-quality sponsors paying for the research [14].
Evidence: the naming came from the sellers
Somebody did extract it. Ask my co-author who, and the answer names no institution:
Companies started selling products to meet the market need before it had a name. Master Data Services, Master Data Maestro, Data Quality Services, and other products entered the market to meet the emerging demand and they needed something on which to base their marketing material. I don’t think there was a specific trigger I can say prompted it, though Sarbanes Oxley was contemporaneous with its emergence—which I think was more coincidence than coercion.
That carries two claims and one refusal. The naming claim is that the demand existed first, and that the label and the frame came from the people who needed something to point marketing at. The refusal is that no trigger is nameable, including the one already on the table. He hedges it twice, and both hedges stay, because a causal negative is the hardest kind of claim to carry and tightening a hedged sentence into a verdict would manufacture a certainty he declined to offer.
The dissent earns the received view stated at full strength first. The Sarbanes-Oxley Act of 2002 (the United States corporate accounting and auditing statute, enacted 30 July 2002 as Public Law 107-204, to improve the accuracy and reliability of corporate disclosures) is, on the standard account, what made data accuracy an executive liability rather than an operational annoyance [1]. That account has contemporaneous evidence behind it, which is more than most received wisdom manages. Writing in Enterprise Systems in September 2003, Vinod Badami set out what compliance would demand of the underlying systems: consolidate and standardize financial data across the operational applications feeding it, track and audit it, integrate its multiple versions into one consistent picture, and supply end-to-end metadata so origins can be traced [2]. That is a specification for master data management, published in 2003, by an author who never once uses the term. Khatri and Brown carry the same association, opening on data governance's rise at practitioner conferences in light of both the opportunity in data assets and the need to meet mandates including Sarbanes-Oxley and Basel II [3].
Testing him, and what the test returned
A claim like his has an obvious test, so I ran it. If compliance created the category, the earliest launch material would lead with compliance and name the statute. Half of it fired, and the half that fired went against him. Microsoft's launch-era post on Master Data Services does not lead with compliance, and it is worth being exact about that, because the ordering is the evidence. Kirk Haselden (then the product unit manager) opens with a personal introduction, then the June 2007 Stratature acquisition and the news that the product would ship as Master Data Services in SQL Server 2008 R2, then the four forces that corrupt master lists, then the case for an authoritative source. Only after all of that does he stop and ask why master data management had emerged as one of the top 2 or 3 initiatives on the average CIO's mind, and answer with what he calls a perfect storm: economic downturn, corporate information ecosystem complexity, strict governmental regulations such as HIPAA and Sarbanes Oxley, services oriented architectures, and the recognition that solving the problem through custom solutions is extremely difficult [6]. So the lead is not compliance, and that half of the test comes back for him. The statute is still there, by name, in the vendor's own first pitch, in the one sentence that answers why now. His claim as stated does not survive that intact.
His answer to the correction is better than the correction. The work was under way before the regulation existed, the rules calcified a direction the market had already taken, and products shipped into newly mandated demand in a shape that takes years to reach. In his phrase, the regulation might have been the capstone, but the pebbles were already in motion. The four pre-statute sources above are that claim with dates attached, and they hold.
Now the part that looks like it cuts the other way. The term has a traceable arrival date, and the shape of the arrival is what a regulation-caused-it story predicts. Across the full text of every electronically filed SEC document, the phrase master data management appears zero times in 2001 and 2002. It first appears in March 2003, in SAP's annual filing, and by that October SAP is describing it to investors as a new offering. Then three filings in 2003, fifteen in 2004 across five companies, forty-one in 2006 across fifteen [16]. That curve measures when a name entered corporate English. The claim under test is about a practice. A category named and sold after a statute is entirely compatible with the work predating it.
The query that discriminates is the capability language, the words available before anyone coined the category, and it does not surge. Single view of the customer runs 17 documents in 2001, 5 in 2002, 9 in 2003 and 17 in 2004, and normalized against filing volume it falls after 2001 and never recovers its own starting rate. Customer data integration is flat across 2001 to 2004 (20, 20, 26, 25) and moves only in 2006, three years after the name arrived [16]. Acxiom's 2001 annual report already sells customer data integration as a named category that lies at the core of effective CRM, providing a single view of the customer, in real time, across multiple data sources, two years before the name appears anywhere in the corpus [16].

The name arrives in 2003. The capability language does not move with it. Five exact phrases counted in SEC EDGAR full-text search across five years, on one shared linear axis labelled EDGAR documents containing the exact phrase. A separate axis per phrase would let “data governance” going from nothing to one document look like a surge, and the comparison between the phrases is the whole finding.
- master data management — the NAME, drawn in gold: 0 in 2001, 0 in 2002, 3 in 2003, 15 in 2004, 41 in 2006.
- customer data integration: 20, 20, 26, 25, then 66.
- single view of the customer: 17, 5, 9, 17, then 10.
- single version of the truth: 0, 0, 8, 2, then 3.
- data governance — dashed, drawn along the floor: 0, 0, 0, 0, then 1.
A dashed vertical marker at 2003 is labelled Mar 2003 — first appearance, in SAP’s annual filing. A bracket beneath the last segment reads 2005 not queried — this segment spans two years, so its slope is not a per-year rate.
The note printed under the plot reads: EDGAR full-text search, exact phrase, queries run 2026-08-25 [16]. RAW document counts, not normalized: real filing volume grew 1.71× across this span once Forms 3/4/5 are stripped, so every rising line overstates growth by about that factor. Coverage begins in 2001, so a zero here says nothing about the 1990s. 15 of the 20 “customer data integration” hits in 2001 are Acxiom filing about itself — early is not widespread.
Three limits, because a result this tidy earns them, and one of the three is severe. The volume is small and vendor-concentrated, with fifteen of the twenty 2001 hits being Acxiom filing about itself, so early is not the same as widespread. Two phrases I would have guessed were older turn out to be younger, since single version of the truth and data governance both arrive late in the corpus [16]. And the SEC's full-text search covers filings from 2001 onward, so this corpus records what became worth saying to investors rather than what practitioners were doing [16]. Kalido makes that seam visible: a master data module at beta in October 2003, sitting on a platform already in production at Shell, Unilever and Philips, announced in a release whose subhead reaches for improved corporate transparency [15]. The engineering predates the regulation. The marketing does not.
That narrower claim forbids something specific, so I tested it as well: no launch or positioning material for the other two products should lead with compliance either. Microsoft's introduction to Data Quality Services frames the need as entry errors, transmission corruption, mismatched dictionary definitions and aggregation across differing standards, with compliance appearing once (as one consequence among several, no statute named) [9]. The most detailed contemporaneous account of Master Data Maestro I could find (a practitioner's, March 2013) runs capability by capability through stewardship interfaces, matching strategies, survivorship, golden records, address verification and a software development kit, and does not mention compliance once [7]. Both hold.
What survives is narrower and more useful. Compliance was named as one pressure among five, in a hedged sentence, inside a post whose frame is operational from end to end: master lists corrupted across divisions and locations by decay, conflict, corruption and inconsistency, and an authoritative source of master data as the goal rather than one source of truth [6]. The investment case offered is cost savings and revenue recovery, not audit survival. Compliance was in the room and named out loud. It was never the frame, and the frame is what the category inherited.
The same sequence is running now
The pattern traced above is not finished. Around AI governance the demand is genuine, the category is being assembled in public largely by people with something to sell, and it is acquiring vocabulary faster than it is acquiring practice. That is the reason to adopt this proposal now rather than at the next revision. The frameworks landing on desks this year list AI governance domains the way the earlier ones listed master data (a domain named by its sellers, placed by nobody, signed by nobody), and they will inherit the defect with the shape. The parts of it that hold will be the parts somebody was already doing in three places at once because the work required them. Those are the parts worth a name first.
Designs considered and rejected
Three designs answer the same problem, and each is held in earnest by people who have made it work. Each is stated below in the form its advocates would recognize, because a design that cannot be stated attractively has not been understood well enough to reject.
Rejected: the listed domain
The case for it is real. Listing is cheap, and it does work that has to be done by somebody: a shared vocabulary, a discoverable map, a place to hang a policy, and a diagram a new hire can read on day one. A framework that lists master data is better than one that omits it, and most of the field has spent twenty years getting to the listing.
It is killed by a case anyone reading this has watched. A framework lists master data as domain one and assigns the definition of client type to a stewardship council. The council meets on schedule, the item is minuted, and the minute records that the council considered it, where considered means, in that room, that nobody recorded an objection. Two systems then ship two definitions of client type in the same quarter, each defensible on its own grain, and neither is escalated, because no rule requires escalation of a decision nobody has classified as a decision. An escalation path that terminates in a council terminates nowhere. The domain was assigned (which is not the same as the definition being owned). The listing was never wrong. It was never a decision either.
Rejected: the single standard, held without exception
This one has the strongest instinct behind it, and my co-author held it first:
Much to my chagrin, early on I tried to shoehorn the entirety of the Enterprise onto a single standard.
He did not abandon it, which matters more than the failure does. He remains still a champion of a single, Enterprise-wide ‘standard’ for attributes like client type, product line and inventory category. What gave way was the presentation layer, and the pressure that moved it has not stopped: he records the pushback we got—and continue to get to make downstream reports match what individual departments expect to see. Note the tense. That is a managed condition of the work, not a solved one, and any account reporting it as resolved has flattened the only honest thing about it.
It is killed on the rollup. Categorization by direct match at one grain holds. Cross-hierarchy references do not, and his formulation of it is the clearest I have heard: when all zits are zats but not all zats are zits, the rollups stop being trustworthy. Concretely, a member belonging to two parents double-counts unless somebody names a primary, a member with no parent at the reporting level drops out of the total in silence, and a ragged hierarchy makes the sum at one level disagree with the sum at the next. None of those announce themselves. The report renders (rendering being the one thing a wrong report does as well as a right one), the numbers are wrong, and the variance stays small enough to be argued about for a quarter.
His concession has two readings, and I will say which one I take. The defensible reading is a governed alias: one authoritative underlying value, with sanctioned labels or roll-ups presented over it. The other reading is genuinely different values in different places, unreconciled, which is local standards with better manners and is the thing he set out to remove. I read him as meaning the first, because it is the only reading under which a single enterprise-wide standard is still true as written. If I have him wrong, the fault is my inference and not his sentence.
Rejected: assigning the domain architecturally
This is the strongest of the three and it deserves its full weight. Zhamak Dehghani, setting out data mesh in 2019, does not negotiate over who holds a domain. She assigns it, arguing that domains should host and serve their own datasets rather than flowing them into a centrally owned platform, and naming the loss of domain ownership after ingestion as the specific failure of the monolithic model [10]. That is a formal answer to the territorial problem, published, adopted, and running in production in places I would not argue with. The domain gets an owner. Assignment is a thing a document can do. Agreement is not.
It is killed by what it leaves untouched. Assigning the owner on paper does not change what confers standing, and my co-author's account of the constant is the sentence I would put on a wall:
While on paper each SVP is a peer to each other SVP, their various domains and areas of responsibility, not to mention the size of their departmental budgets, established a very real power hierarchy.
SVP is senior vice president (the abbreviation is his), and the observation is that the org chart shows a flat row of them while the real order is set by budget and span. Territoriality persists because those two things stay welded to position, and a governance framework asks an executive to give up control over exactly what his standing is measured in (the framework calls that a domain assignment; he reads it as a transfer, and he reads it correctly). That is a claim about incentives, and incentives are out of reach of training and communication. His own response to data mesh is the other half of it: the federated model it describes is one he was already running, which is this article's finding arriving once more in a place I had not planned for it.
The Seat That Signs
Practitioner layer — the curator's read on the consensus above. Governance is my lane; master data is my co-author's, and I have kept my judgement to where the two meet.
The failure mode I would watch hardest
The signature that names a body. A framework can carry a wicked long list of domains, attach a seat to each, and still write the stewardship council in the slot, which satisfies the proposal in form and defeats it in substance. Khatri and Brown put the whole of governance in decision rights and the locus of accountability [3], and a locus is one place. The test is whether anybody can name the person who decides the enterprise definition of client type, and whether that decision survives contact with a department that wanted another one. The coxswain never pulls an oar and is still the only one who can be held to the line the boat takes.
The trade-off that usually bites
The one above, and it bites in both directions. Hold the single standard rigidly and the political capital goes on report-formatting arguments that cannot be won, since the departments are not being difficult and their numbers genuinely have to reconcile to figures already published. Concede too broadly and you hold local standards wearing an enterprise badge. The workable middle is narrow, and it is a governance call rather than a technical one: one authoritative value, with sanctioned alternate presentations (sanctioned meaning signed, not merely tolerated) documented as derivations. A good constraint narrows honestly to one answer, which is what makes it feel less like a fence than a gift. If nobody can tell you which value is authoritative and which is the alias, the middle has already gone.
The claim I would be sceptical of
Any clean causal history of this discipline, including the one I have just partly defended. The received account is not fabricated. It has real contemporaneous evidence, in a 2003 trade article specifying the requirement before the vocabulary existed [2] and in the vendor naming the statute six years later [6]. A category emerging across a decade, in a market carrying a downturn, a complexity problem, an architectural fashion and a compliance regime at once, is overdetermined by construction. Anyone selling a single trigger is selling a story. My co-author's contribution here is not that he identified the right cause. It is that he declined to name one, which is the harder and rarer move.
Where to Go Deeper
In the order I would read them. Khatri and Brown's Designing Data Governance [3] first, because it is five pages, it is open access, and its governance-versus-management distinction will do more for a program than a maturity model will. Then DAMA International's framework [4], for the eleven knowledge areas and specifically for where it puts reference and master data management relative to governance. Kimball's 2003 column on drilling across [5] is twenty-three years old, free, and the clearest short statement of conformed dimensions anybody has written; read it to see how much of what now gets called governance was already settled inside warehouse design. Haselden's 2009 announcement post [6] is worth reading whole, because a category can be watched being framed in real time. And Dehghani on data mesh [10] for the strongest modern argument that a domain owner should be assigned rather than negotiated, which is the best case against my scepticism and the reason it is listed here.
What this proposal leaves unsettled
Four things are not settled, and a framework ratified without saying so is a framework that will be corrected in public later. Each carries the seat that will settle it.
Whether a named owner reduces territorial conflict is unmeasured, and this proposal assumes it does. Dehghani's article is an architectural argument and makes no claim about conflict rates [10]. I did not find a before-and-after comparison of territorial conflict around a domain assignment in any source I opened, and I would rather leave that standing than close it with a framework. The seat that settles it is whoever runs the first assignment and keeps the count.
Which reading of the presentation concession is right remains my inference rather than his statement. Governed alias and unreconciled local values look identical in a framework document (both read as flexibility) and behave nothing alike at quarter-end. The seat that settles it is the one that signs the definition, by publishing the authoritative value and its sanctioned derivations in the same place, where a reader can see which is which.
The proposal does not reach what confers standing, and that is where the resistance lives. It can require a name beside a definition. It cannot alter the arithmetic that makes a budget and a span of control into a rank. The seat that settles that one is the chief executive, and nothing written into a governance framework has ever reached it.
Whether the AI governance domains now being drawn will be signed is undecided, and the seat is empty in most of the frameworks I am shown. The seat that settles it is the one that ratifies the next framework, and that ratification is a few months out for most readers of this piece.
The version of this piece readers saw first is preserved unedited for anybody who would rather judge the two side by side (the reasoning behind the change is set out in The Voice Problem).
References
[1] Sarbanes-Oxley Act of 2002, Public Law 107-204, 116 Stat. 745, enacted 30 July 2002 (H.R. 3763), U.S. Government Publishing Office — cited for the enactment date, the public law number, and the Act's stated purpose, “To protect investors by improving the accuracy and reliability of corporate disclosures made pursuant to the securities laws.” I read the enacting clause, the short title and the table of contents, not the full statute; nothing in this article rests on the text of a particular section. https://www.govinfo.gov/content/pkg/PLAW-107publ204/html/PLAW-107publ204.htm
[2] Vinod Badami, “Sarbanes-Oxley: The Role of Business Intelligence and Data Management,” Enterprise Systems, 17 September 2003 — a trade article written thirteen months after enactment, by the then national director for business intelligence at an IT professional services firm. Source of the consolidation, tracking-and-auditing, multiple-versions-standardization and end-to-end-metadata requirements, and of the phrase “framework of financial governance.” Cited as evidence of what was being asked for in 2003 and in what vocabulary, not as an independent finding. Read in full 2026-08-22. https://esj.com/articles/2003/09/17/sarbanesoxley-the-role-of-business-intelligence-and-data-management.aspx
[3] Vijay Khatri and Carol V. Brown, “Designing Data Governance,” Communications of the ACM 53, no. 1 (January 2010): 148–152 — source of the governance-versus-management distinction adapted from Weill and Ross, the data-quality worked example, the five decision domains (data principles, data quality, metadata, data access, data lifecycle), and the opening sentence associating data governance's rise with both data-asset opportunity and compliance mandates including Sarbanes-Oxley and Basel II. Read in full 2026-08-22 via the open-access edition. https://cacm.acm.org/research/designing-data-governance/
[4] DAMA International, “DAMA-DMBOK Framework — Core Knowledge Areas” — the association's own current statement of its framework, cited for the eleven core knowledge areas and for the position of Reference & Master Data Management as area eight, listed alongside Data Governance as area one. This is the framework as published today; the page does not carry edition dates and no dated claim is made from it here. Read 2026-08-22. https://www.damadmbok.org/copy-of-about-dama-dmbok
[5] Ralph Kimball, “The Soul of the Data Warehouse, Part Two: Drilling Across,” Kimball Group, April 2003 — source of the precise definition of conformed dimensions (“Two dimensions are conformed if the fields that you use as common row headers have the same domain”), the consequence of attempting a drill-across on unconformed dimensions, and the framing of the whole discussion as data warehouse architecture. Read in full 2026-08-22. https://www.kimballgroup.com/2003/04/the-soul-of-the-data-warehouse-part-two-drilling-across/
[6] Kirk Haselden (SQL Server Team), “Master Data Services – What's the big deal?”, Microsoft SQL Server Blog, 13 May 2009 — written by the then product unit manager for Master Data Services. Source of the June 2007 Stratature acquisition, the announcement that the product would ship as Master Data Services in SQL Server 2008 R2, the four forces (decay, conflict, corruption, inconsistency), the preference for “authoritative source of master data” over “one source of the truth,” and the “Why now?” passage — the question about master data management having “emerged as one of the top 2 or 3 initiatives on the average CIO's mind” and the “perfect storm” sentence naming economic downturn, ecosystem complexity, HIPAA and Sarbanes Oxley, services oriented architectures and the difficulty of solving the problem through custom solutions. Placement matters to the argument, so it is stated here too: that passage is not the post's opening. It comes after the author's introduction, the acquisition and ship announcements, the four forces and the authoritative-source argument. Cited as vendor positioning — evidence of how the category was framed, not evidence of the product's value. Read in full 2026-08-22. https://www.microsoft.com/en-us/sql-server/blog/2009/05/13/master-data-services-whats-the-big-deal/
[7] James Serra, “Profisee Master Data Maestro,” jamesserra.com, 7 March 2013 — a contemporaneous practitioner account by a then-independent data warehouse consultant and Microsoft SQL Server MVP, writing while using Master Data Services and Data Quality Services daily. Cited for three things: that Master Data Maestro is Profisee's and is built on top of Master Data Services; that Profisee's principals originally built the product Microsoft acquired from Stratature in 2007; and for the content of its capability comparison, in which compliance does not appear. This is a personal blog, not a peer-reviewed or editorially controlled source, and it is used here as period evidence of how the product was described rather than as an independent verification of any vendor claim. Read in full 2026-08-22. https://www.jamesserra.com/archive/2013/03/profisee-master-data-maestro/
[8] Microsoft, “What is the medallion lakehouse architecture?”, Azure Databricks documentation, Microsoft Learn — cited for what the bronze, silver and gold layers denote (raw ingestion; cleaning and validation; dimensional modeling and aggregation) and for the description of the pattern as a data design pattern for progressive refinement. Used only to establish what the term means and where it comes from, which is what makes its use in this article an identified anachronism. Read in full 2026-08-22. https://learn.microsoft.com/en-us/azure/databricks/lakehouse/medallion
[9] Microsoft, “Introduction to Data Quality Services,” Microsoft Learn (SQL Server documentation) — the vendor's own documentation, cited for what Data Quality Services is and does (knowledge-driven cleansing, matching, reference data services, profiling, monitoring), for its arrival as a SQL Server feature installed with the product, and for the framing of “The Business Need for DQS,” in which compliance appears once as one consequence of incorrect data among several. The page's article metadata carries an original date of 5 March 2012, consistent with the SQL Server 2012 release. Read in full 2026-08-22. https://learn.microsoft.com/en-us/sql/data-quality-services/introduction-to-data-quality-services
[10] Zhamak Dehghani, “How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh,” martinfowler.com, 20 May 2019 — cited for the argument that domains should host and serve their own datasets rather than flowing them into a centrally owned platform, and for the identification of lost domain ownership after ingestion as the failure of the monolithic model. My reading was targeted rather than complete: I read the sections on centralized data platforms and on domain-oriented decentralized data ownership, which are the passages the claim rests on, and did not read the article end to end. Read 2026-08-22. https://martinfowler.com/articles/data-monolith-to-mesh.html
[11] Stuart Madnick, Richard Wang, Frank Dravis and Xinping Chen, “Improving the Quality of Corporate Household Data: Current Practices and Research Directions,” Proceedings of the Sixth International Conference on Information Quality (ICIQ), November 2001, 92–104 — the load-bearing pre-statute source in this article, and the reason the “pebbles were already in motion” claim is checkable rather than a memory. Source of the eleven-step commercial engagement delivering a single view of customers, the match and consolidation rules stored in a rules matrix, the confidence thresholds, and the cross-functional definitional sign-off (“The terms had different meanings to different people”). What this paper is NOT cited for, because an earlier draft got it wrong: its ABI/INFORM literature-search sentence is about the authors' own proposed coinage corporate householding, and it reports that no corresponding concept existed — which is close to the opposite of the concept existed and had no name. The engagement itself is the evidence here; that sentence is not. Read in full 2026-08-25 from MIT's own Total Data Quality Management publications archive; dated November 2001 there and corroborated by internal citations to 2001 work. https://web.mit.edu/tdqm/www/tdqmpub/CorporateHouseholdNov01.pdf
[12] Ralph Kimball, “The Matrix,” Intelligent Enterprise column, December 1999, in the Kimball Group archive — source of the “conforming meeting” that is “probably more political than technical” and that a senior manager such as the enterprise CIO “should be willing to appear” at. The column describes a meeting to conform a dimension and treats the planning-matrix columns as its invitation list; it does not describe it as a standing or recurring forum, and this article does not claim it was one. Also the source of the framing of a common definition of customer as “a major litmus test for an organization.” A date caveat, because this is a history piece and the date is the point: the page displays December 7, 1999 beneath the byline and encodes it as 1999-12-07T22:40:26-06:00, so the date is attested to the day. What it is not is a 1999 artifact: the same head carries a modified time of 2016-01-26, and the markup is a modern content-management system. So this is a migrated record rather than a primary one — good enough to date the column, not the same thing as holding the December 1999 issue. Read in full 2026-08-25. https://www.kimballgroup.com/1999/12/the-matrix/
[13] Brian Sullivan and Michael Meehan, “UCCnet's Promise: Synchronized Product Data,” Computerworld, 10 June 2002 — cited for a product-data registry with 62 possible data fields, subscription fees ranging from $1,500 to $400,000 a year, named companies including Procter & Gamble and Sara Lee's Bakery Group, and a Wal-Mart supplier mandate, all seven weeks before Sarbanes-Oxley was signed. A tense caveat, because it bears on what the source proves: the piece is prospective, not a report of a running system — the registry “was designed to” let retailers and suppliers exchange data, suppliers “will then be able to publish,” and UCCnet “will try to eliminate” pricing inconsistencies. What is attested in the present tense is the registry's 62 fields, its published fee schedule, and Wal-Mart's mandate; the synchronization itself is future. Note also the structural point it carries: the forcing function here was a retailer's mandate, not a statute. Read in full 2026-08-25. The date is publisher-asserted in the page metadata of a modern re-publication of a 2002 article rather than an original artifact. https://www.computerworld.com/article/1328390/uccnet-s-promise-synchronized-product-data.html
[14] The Data Warehousing Institute, press release for Wayne W. Eckerson, Data Quality and the Bottom Line, 1 February 2002 — cited for the $600 billion annual cost figure, the 647-respondent survey base, the circulation of the phrase “single version of the truth,” and the existence of six paying commercial data-quality sponsors, six months before the statute. The REPORT itself was not read — the PDF would not retrieve — so every figure here is cited to the dated press release, which was read in full 2026-08-25. Also worth recording because it cuts against this article: the same release notes that almost half of respondents had no current plans to improve data quality, so practice existed but adoption was shallow, which leaves real room for a post-2002 forcing function. https://insurance-canada.ca/2002/02/01/the-data-warehousing-institutes-recent-study-finds-high-quality-data-is-critical-to-the-success-of-businesses-worldwide/
[15] Kalido, “Kalido Launches Data Warehouse Lifecycle Management Software Suite,” press release, London, 27 October 2003, carried by Enterprise Systems Journal — cited for a master data management module at beta stage in October 2003, sitting on a warehouse platform already in production at Cadbury Schweppes, HBOS, Intelsat, Royal Dutch/Shell, Philips and Unilever. That timeline is the “years in development” evidence, and the release's subhead reaching for “improved corporate transparency” — with “regulatory compliance” appearing once in the body, at about the same depth as the vendor documentation this article scores as NOT leading with compliance — is the evidence that marketing followed the statute where engineering could not have. Page 1 read in full 2026-08-25; page 2 not reached. Dateline, URL path and on-page stamp agree on the date. https://esj.com/articles/2003/10/27/kalido-launches-data-warehouse-lifecycle-management-software-suite.aspx
[16] U.S. Securities and Exchange Commission, EDGAR full-text search — queries run 2026-08-25 for the exact phrase “master data management” across all electronically filed documents, returning zero filings for 2001–2002, three in 2003 (all SAP AG, the earliest being SAP's Form 20-F filed 21 March 2003), fifteen in 2004 across five filers, and forty-one in 2006 across fifteen. The SAP 6-K exhibit filed 16 October 2003 describing “Master Data Management (SAP MDM), a new offering” was read in full. Also cited for the capability-language queries run the same day, which are the discriminating test: exact-phrase counts for 2001, 2002, 2003, 2004 and 2006 on single view of the customer (17 / 5 / 9 / 17 / 10), customer data integration (20 / 20 / 26 / 25 / 66), single version of the truth (0 / 0 / 8 / 2 / 3), data governance (0 / 0 / 0 / 0 / 1), plus data quality, reference data and data warehouse. Normalization uses non-ownership filing counts derived from EDGAR's quarterly form.idx full-index files, because raw totals are inflated 3.12× by Forms 3/4/5 after Sarbanes-Oxley section 403 mandated electronic Section 16 filing — tiny XML documents that can never contain these phrases. Stripping them, real growth is 1.71×. Source of the Acxiom 10-K passage (accession 0000733269-01-500016, filed 2001-06-27) and the Erie Indemnity cost-sharing exhibit (accession 0000922621-01-500002, filed 2001-07-19), both opened and read rather than counted. Three limits recorded because they bound everything above: coverage begins in 2001 per the SEC's own FAQ, so the zero says nothing about the 1990s; the endpoint counts documents while the denominator counts filings, so normalized figures overstate growth by an unmeasured factor; and reference data proved unusable as a proxy at all, its early counts dominated by a Boston commercial lease form whose opening article is headed “REFERENCE DATA.” https://www.sec.gov/edgar/search/