Weaving Intelligence
Before It Had a Name: How Master Data Ended Up at the Bottom of Every Governance Framework (original)
The 2026-08-28 original, preserved unedited for comparison.
The received history says a 2002 accounting statute created the pressure, and that master data management followed. My co-author was building the thing at the time, and he declines that story. Here are his words, the dates, and what happened when I went looking for evidence he was wrong.
Open any current governance framework and master data is in it, usually near the floor, usually described as foundational. That placement looks obvious now. It was not always there, it did not arrive because a committee voted it in, and the story of how it got there is less tidy than the version that circulates.
This is a history piece, so a word about where it comes from. The spine is one practitioner's account — my co-author's, given in answer to questions about origins, and quoted as his where it is his. Around that I have put public sources: the statute, the vendor documentation, the academic framework, the trade press of the period. The sources are scaffolding. They are here so you can check the dates yourself, and in one place they made me correct him. Where that happened I have said so rather than smoothing it over.
The line I can't draw, and why that matters here
You cannot trace a relationship without saying what the relationship is, so let me get the awkward part out of the way. Ask my co-author where governance stops and master data management (MDM) begins and the first thing he does is concede that he blurs it himself, because he doesn't believe there is a hard line. Pressed, he draws one anyway: Data Governance manages the metadata surrounding Master Data Management while Master Data Management handles the underlying data
. Governing, in that description, is the work of categorizing, classifying, writing the rules, and modeling the structures. And then the clean formulation: Governance is the rules and the framework within which MDM operates.
The academic version lands close enough to be worth putting beside it. Khatri and Brown, writing in Communications of the ACM in 2010, separate governance from management this way: governance is about what decisions must be made and who holds the right to make them, while management is making and implementing those decisions [3]. Their worked example is data quality — governance establishes who decides the standard, management picks the actual metric. That is the same cut, made in a different vocabulary.
Here is the part that matters for a history. Their framework has five decision domains: data principles, data quality, metadata, data access, and data lifecycle [3]. Master data is not one of them. Metadata is — which is precisely where my co-author puts governance's job. DAMA International's body of knowledge, the field's other durable framework, organizes data management into eleven core knowledge areas and does list reference and master data management, as area eight, sitting alongside data governance rather than underneath it [4].
So the modern picture — governance framework on top, master data as its foundational domain — is not what either canonical framework actually says. It is something practice arrived at. Which raises the reasonable question of where practice arrived at it, and when.
The practice arrived before the vocabulary
Ask my co-author where this work lived before anybody called it governance, and the answer is that it lived in three places at once. Data got cleaned during extract, transform and load — ETL, the movement of data out of source systems and into the warehouse. It got validated and translated inside the warehouse itself. And it got shaped again into specialized views and hierarchies in the reporting layer. Not a program. Three places where the problem showed up, each handled where it showed up.
Then this, which I am quoting rather than summarizing, because the summary loses it:
We were doing medallion builds and multi-system MDM integration with approval gates before we knew that’s what we were doing because it was the best way to meet the needs of the business.
He supplied the setting afterwards, and it is worth having, because before we knew that's what we were doing
is a claim about a date. The engagement began in 1997 and ran ten years, at a publicly traded company. The build was a web-based document management system for the Engineering department that integrated with HR, Accounting, Sales and other departments' systems to produce current reports on factory upgrades, builds and installations. Read the nouns rather than the technology: cross-domain integration, an owning department that is neither the business nor IT, and reporting that has to reconcile across all of it. That is the whole shape of the discipline, five years before the statute that is supposed to have created it.
Two things to flag before that sentence can be used. First, medallion is an anachronism and a deliberate one. The bronze-silver-gold layering it refers to is lakehouse-era vocabulary — it describes raw ingestion, then cleaning and validation, then dimensional modeling and aggregation for business consumers [8]. That term did not exist when the work described here was done. Applying it backwards is useful, because it tells a reader in one word what the shape was. It is only honest if somebody says out loud that the label was fitted afterward, so: it was.
Second, the build had a name, and the name was architectural rather than disciplinary — a Kimball data warehouse, after Ralph Kimball's dimensional-modeling approach to warehouse design. And then the detail I would not have predicted: the Engineering department owned it, not the business and not IT—we were our own special animal.
Two decades of governance literature is built on the question of whether the business or IT should own the data. His origin case answers neither, and he mentions it in passing, as a fact about the org chart rather than a position.
Now the obvious objection, because it is a good one. Saying nobody knew what they were doing is a retroactive-labeling claim, and the parts of that work were emphatically not nameless. Kimball had defined conformed dimensions precisely and publicly by 2003 — two dimensions conform when the fields used as common row headers are drawn from the same domain, and without that, cross-fact-table queries return garbage [5]. Cleansing had a name. Reference data management had a name. The vocabulary existed.
So the strong claim doesn't hold, and I am not going to defend it. The narrower one does, and it still forbids something. What had no name was the assembly: the layered, multi-system, gated pattern treated as one discipline, with its own tooling, its own roles, and its own line in a budget. The parts were named. The whole was not, and nobody was selling it as a whole. Kimball's own column makes that visible from the other side — conformed dimensions appear there as an argument about data warehouse architecture, inside a discussion of drilling across fact tables, not as a governance practice [5]. The idea was in the warehouse literature. The discipline had not been extracted from it.
Who named it, and the story I am not going to tell you
Somebody did extract it. Ask him who, and the answer is not an institution:
Companies started selling products to meet the market need before it had a name. Master Data Services, Master Data Maestro, Data Quality Services, and other products entered the market to meet the emerging demand and they needed something on which to base their marketing material. I don’t think there was a specific trigger I can say prompted it, though Sarbanes Oxley was contemporaneous with its emergence—which I think was more coincidence than coercion.
That makes two claims and only one of them is the interesting one. The first is about naming: the demand was already there, and vendors supplied the label and the frame because they needed something to point marketing at. The second is a refusal. Handed a menu of tidy causes, he declines all of them, including the one already on the table — and he hedges it twice, which I am keeping. Those hedges are doing real work; a causal negative is the hardest kind of claim to carry, and he carries it carefully.
It is also a dissent, so the received view deserves stating at full strength first.
The Sarbanes-Oxley Act of 2002 — the United States corporate accounting and auditing statute, enacted 30 July 2002 as Public Law 107-204, to improve the accuracy and reliability of corporate disclosures [1] — is, on the standard account, what made data accuracy an executive liability rather than an operational annoyance. And the standard account has contemporaneous evidence behind it, which is more than most received wisdom manages. Writing in Enterprise Systems in September 2003, thirteen months after enactment, Vinod Badami set out what compliance was going to demand of the underlying systems: consolidate and standardize financial data across the operational applications that feed it; provide the means to track and audit that data; give users the ability to standardize and integrate multiple versions of financial data to present a consistent picture; supply end-to-end metadata so the data's origins can be traced [2]. That is a specification for master data management, published in 2003, by someone who does not use the term. I originally wrote because it was not the available term, and my own timeline further down this page refutes that: the phrase was in an SEC filing six months before Badami published. What is defensible is the absence itself — neither master data management
nor even master data
appears anywhere in his piece — not an explanation for it. He calls the answer business intelligence and data management, and locates it inside what he calls a framework of financial governance [2].
The academic literature carries the same association. Khatri and Brown open by noting that data governance had recently gained prominence at practitioner conferences in light of both the opportunity to leverage data assets and the need to ensure compliance with mandates such as Sarbanes-Oxley and Basel II [3]. Two drivers, given equal billing, in a peer-reviewed venue in 2010.
So here are the dates, which is what you actually need to form your own view. The statute is July 2002 [1]. The demand-side argument is in print by September 2003 [2]. Microsoft acquired the vendor whose product became Master Data Services in June 2007, and announced in May 2009 that it would ship in SQL Server 2008 R2 [6]. Profisee's Master Data Maestro, built on top of Master Data Services by the team that originally wrote it, was in the market and being evaluated by working practitioners by early 2013 [7]. Data Quality Services arrived with SQL Server 2012 [9]. Seven years, minimum, between the statute and the first of the three products he names. The gap is real, and so is the contemporaneity across the middle of the decade, which is exactly why coincidence and causation are hard to separate here.
What happened when I tested him
A claim like his has an obvious test, so I ran it. If compliance created the category, the earliest launch and marketing material for those products would lead with compliance and name the statute. Go read it.
It partially fired, and against him. Microsoft's launch-era post announcing Master Data Services asks why master data management had suddenly become a top-three item for chief information officers, and answers with what it calls a perfect storm: economic downturn, corporate information ecosystem complexity, strict governmental regulations such as HIPAA and Sarbanes-Oxley, service-oriented architectures, and the recognition that solving the problem with custom code is extremely difficult [6]. The statute is there, by name, in the vendor's own first pitch. His claim as stated — contemporaneous but not causal — does not survive that intact, and I would rather correct him in public than build a section on a sentence I had not checked.
He came back on that, and his answer is better than my correction. His point is not that the statute went unmentioned. It is that the work was already underway before the regulation existed: business had identified the gap and was building to close it, and when the rules arrived they calcified a direction the market was already moving in. Then vendors shipped products into the newly mandated demand — products that must have been in development for years to have arrived in the shape they did. In his phrase, the regulation may have been the capstone, but the pebbles were already in motion.
That is a claim with a date on it, so I went and checked it rather than taking either of us on trust. It holds, and the evidence is not thin. In November 2001, nine months before the statute, Madnick, Wang, Dravis and Chen published a worked account of an engagement delivering a single view of customers — match rules, consolidation rules, confidence thresholds, and cross-functional sign-off on what a customer
even was [11]. That is master data management and its governance, executed commercially, before the vocabulary. Two years earlier Ralph Kimball was describing a standing conforming meeting
that a CIO should personally attend, because agreeing a definition of customer
across an enterprise is more political than technical
[12] — a data stewardship forum under another name, in 1999. Seven weeks before Sarbanes-Oxley was signed, the grocery industry had a live product-data registry with standardized attributes, subscription fees reaching six figures, and a Wal-Mart supplier mandate behind it [13]. And in February 2002 an industry survey of 647 practitioners was already costing bad data at $600 billion a year and using the phrase single version of the truth
, with six commercial data-quality vendors paying for the research [14].
So the pebbles were in motion, demonstrably. Now the part that looks like it cuts the other way, and the reason it does not. The term itself has a traceable arrival date, and the shape of that arrival is exactly what a regulation-caused-it story predicts. Searching the full text of every electronically filed SEC document, the phrase master data management appears zero times in 2001 and 2002. It first appears in March 2003, in SAP's annual filing, and by that October SAP is describing it to investors as a new offering.
Then: three filings in 2003, fifteen in 2004 across five companies, forty-one in 2006 across fifteen [16]. The takeoff begins immediately after the statute and steepens for four years.
Except that curve measures when a name entered corporate English, and the claim under test is about a practice. Those are different objects, and I had them confused: a category being named and sold after a statute is entirely compatible with the work predating it, which is precisely his thesis rather than a rival to it. The regulation may well have forced the industry to adopt a name. That is not the same as the regulation having caused the thing the name refers to.
Which suggests the query that would actually discriminate. If the statute raised the commercial salience of the practice rather than just supplying its label, then the language for the underlying capability — the words available before anyone coined the category — should surge alongside the name. So I ran the same method on capability language [16]. It does not surge. Single view of the customer
runs 17 documents in 2001, 5 in 2002, 9 in 2003, 17 in 2004 — and normalized against filing volume it falls after 2001 and never recovers its own starting rate. Customer data integration
is flat across 2001–2004 (20, 20, 26, 25) and only moves in 2006, three years after the name arrived. The enterprise-data-management share of data quality
is flat from 2001 to 2006 — its raw growth turns out to be mining assay reports and index boilerplate. Data warehouse
declines on a normalized basis.
And the capability language is real where it appears, which I checked at sentence level rather than by counting. Acxiom's 2001 annual report sells Customer Data Integration
as a named product category that lies at the core of effective CRM, providing a single view of the customer, in real time, across multiple data sources
— two years before master data management appears anywhere in the corpus. In July 2001 a property-casualty insurer wrote into an inter-company cost-sharing agreement that its platform provides a single view of the customer across the enterprise in support of all customer processes and touchpoints
[16]. That is a buyer, not a vendor, putting the capability in a contract before the statute.
Three limits, because a result this tidy deserves them and one of them is severe. First, the volume is small and vendor-concentrated: every 2001 hit for customer data integration
traces back to four filers, fifteen of the twenty being Acxiom filing about itself. Early is not the same as widespread. Second, two phrases I would have guessed were older turn out to be younger: single version of the truth
is absent in 2001 and 2002 and first appears in 2003, the same year as the name, and data governance
is absent through 2004 with a single document in all of 2006. Third and most important: SEC full-text search only covers filings from 2001 onward [16], and a corpus of filings is a record of what became worth saying to investors, not a record of what practitioners were doing. The Kalido case makes the seam visible: the release describes the suite as a data warehouse and master data management software portfolio
and the underlying platform as supporting master-data-aligned warehouses — and that platform was already in production at Shell, Unilever and Philips. A module can be built in fifteen months; the master-data capability it was packaging cannot have started after July 2002 — and the same press release puts improved corporate transparency
in its subhead and reaches for proactive management of regulatory compliance
in its body [15]. The engineering predates the regulation. The marketing does not. That distinction is the whole of it.
What survives is narrower and, I think, more useful. Compliance was named as one pressure among five, in a hedged sentence beginning with the admission that the answer lies somewhere in the list. It was not the frame. The frame of that entire post is operational: master lists that get corrupted in myriad ways across divisions and locations, and four forces acting on them — decay, conflict, corruption, inconsistency — against which the goal is not truth but an authoritative source of master data [6]. The investment case offered is cost savings and revenue recovery, not audit survival.
That narrower claim forbids something specific, so I tested it too: no launch or positioning material for the other two products should lead with compliance either. Microsoft's introduction to Data Quality Services frames the business need as incorrect data arising from entry errors, transmission corruption, mismatched dictionary definitions and the aggregation of sources using different standards; compliance appears once, as one consequence among several alongside lost credibility, lost revenue and unhappy customers, with no statute named [9]. The most detailed contemporaneous account of Master Data Maestro I could find — a practitioner's, written in March 2013 by someone using the products daily — runs to a capability-by-capability comparison covering stewardship interfaces, matching strategies, survivorship, golden records, address verification and a software development kit, and does not mention compliance once [7]. Both hold.
The honest summary: compliance was in the room and named out loud in the vendor's own first pitch, so coincidence is too strong a word for what happened. But it was never the frame these products were sold on. Every one of them was pitched on operational consolidation — one authoritative version of the customer, the product, the part — and that is the frame the category inherited. I would rather hand you the dates and the actual sentences than a verdict.
One standard for the whole enterprise, and what it cost to learn better
Master data does not become the foundation of a governance framework because a framework says so. It becomes the foundation when somebody tries to run the framework without it and finds out. My co-author's version of finding out is the most useful paragraph in this piece, partly because it is at his own expense:
Much to my chagrin, early on I tried to shoehorn the entirety of the Enterprise onto a single standard.
What broke it was not architecture. It was, in his words, the amount of pushback we got—and continue to get
to make downstream reports match what individual departments expected to see. Note the tense. That pressure did not stop when he changed his approach; it is a condition of the work, managed rather than solved, and any account that reports it as resolved has flattened the only honest thing about it.
The correction he made is narrower than the failure suggests, and this is where I have watched people take the wrong lesson. He did not conclude that enterprise standards don't work. He remains, as he puts it, still a champion of a single, Enterprise-wide ‘standard’
for attributes like client type, product line, inventory category. What he conceded was the presentation layer: the flexibility to show alternatives to particular segments of the business, for backward compatibility or for a department's own reporting.
That concession has two possible readings and they are not equally defensible, so I will say which one I am taking. The defensible reading is a governed alias: one authoritative underlying value, with multiple sanctioned labels or roll-ups presented over it. The other reading — genuinely different values in different places, unreconciled — is just local standards with better manners, and it is the thing he says he was trying to eliminate. I read him as meaning the first, because it is the only reading under which a single enterprise-wide standard is still true as written, and because a concession that dissolves the standard it claims to preserve is not a concession, it is a retreat with good PR. If I have him wrong, the fault is my inference and not his sentence.
He confirmed the reading and then narrowed where it holds, which is the more useful answer. It is right wherever categorization can happen through direct matches at the same grain and the same category. Where it stops being right is cross-hierarchy references and rollups: when all zits are zats but not all zats are zits, the rollups stop being trustworthy, and there is no clean way around it. Concretely: a member that belongs to two parents double-counts unless somebody picks a primary, a member with no parent at the reporting level falls out of the total silently, and a ragged hierarchy makes the sum at one level disagree with the sum at the next. None of those announce themselves — the report renders, the numbers are wrong, and the variance is small enough to be argued about for a quarter. And the cost is not technical. A business that has categorized by zits for decades gets very cross when it is told to move to the zat paradigm — which is the same political problem as the definitional one, arriving one level up and with a longer institutional memory behind it.
And that is the mechanism by which master data ends up at the bottom of the framework. Not because a body of knowledge listed it. Because the standard is the only thing that survives the pushback, and the standard has to sit on something. Governance without an authoritative underlying value is a policy about a value nobody agrees on.
What twenty years of frameworks did not touch
Ask what all that framework-building has failed to solve and the answer comes back in three words: politics and personality conflicts. Which sounds like a grumble until the mechanism arrives with it. Territoriality persists because budget, area of responsibility and domain of control remain welded to position and power. Nothing has changed that, so nothing has changed the behavior it produces. Then the sentence I would put on a wall:
While on paper each SVP is a peer to each other SVP, their various domains and areas of responsibility, not to mention the size of their departmental budgets, established a very real power hierarchy.
SVP is senior vice president, and the observation is that the org chart shows a flat row of them while the actual hierarchy is set by budget and scope. That gap is checkable on your own org chart this afternoon, which is more than most observations about organizational politics can say for themselves.
The generic version of this — governance is a people problem, not a technology problem — is thirty years old and it wastes the answer. His claim is narrower and structural. Governance asks executives to give up control over the exact things their standing is measured in. That is a statement about incentives, and incentives imply a remedy that training and communication cannot reach. You either change what confers power, or you route around it. The diagnosis has been published a thousand times; that implication is the part worth taking away.
Now, nothing has changed is an unqualified universal across two decades, and there is a real candidate against it. Data mesh, as Zhamak Dehghani set it out in 2019, does not negotiate over domain ownership — it assigns it, arguing that domains should host and serve their own datasets rather than flowing them into a centrally owned platform, and naming the loss of domain ownership after ingestion as the specific failure of the monolithic model [10]. That is a formal answer to exactly the territorial problem described above — and my co-author's response to it was not that it is wrong but that it is his: the federated ownership model he had already been running. Which is this article's own thesis arriving one more time, in a place I had not planned for it. The formal answer published in 2019 is a practice somebody was operating before it had that name either.
So the claim narrows, and here is what the narrowed version forbids. It concedes that the formal structure has changed — ownership can now be assigned on paper, and often is. It holds that the informal one has not. What that forbids is straightforward: it forbids evidence that assigning domain ownership on paper measurably reduces territorial conflict. I went looking for that evidence and did not find it. Dehghani's article is an architectural argument, not a study of organizational behavior, and it makes no claim about conflict rates [10]. I could not locate a measurement of territorial conflict before and after a domain-ownership assignment in any source I opened. That is an absence of evidence rather than evidence of absence, and I would rather leave it standing as an open question than close it with a framework. If you have run that comparison, I would genuinely like to see it.
The naming is happening again
One observation before I hand this over, and I will keep it short because it is an aside rather than the argument.
The pattern traced above is not finished. It is running right now around AI governance: the demand is genuine, the category is being assembled in front of us largely by people with something to sell, and it is acquiring vocabulary faster than it is acquiring practice.
The useful thing to take from that is not cynicism. Someone has to name a category, and the products that carried master data management into the enterprise were good products. It is that the name arrives after the work and is shaped by whoever needs to sell something. So when the AI governance framework lands on your desk with its domains neatly laid out, ask the question this history answers: which parts of this were people already doing, in three places at once, because the business needed them? Those are the parts that will hold.
What I Would Watch For
Practitioner layer — the curator's read on the consensus above. Governance is my lane; master data is my co-author's, and I have kept my judgement to where the two meet.
The failure mode I'd watch hardest
A framework that lists master data as a foundational domain and changes nothing about who decides what a customer is. Listing is free. The whole force of Khatri and Brown's framing is that governance is about decision rights and the locus of accountability [3] — so the test of whether master data is genuinely foundational in your framework is not whether it appears in the diagram. It is whether anyone can name the person who decides the enterprise definition of client type, and whether that person's decision survives contact with a department that wanted a different one. If the answer is that the framework document was updated, you have a domain on a slide. Eight strong rowers all pulling hard on their own count still lose.
The trade-off that usually bites
The one my co-author walked into, and it bites in both directions. Hold the single standard rigidly and you will spend your political capital on report-formatting arguments you cannot win, because the departments are not being difficult — their numbers genuinely have to reconcile to things they have published before. Concede too broadly and you have local standards wearing an enterprise badge. The workable middle is narrow and it is a governance decision rather than a technical one: one authoritative value, with sanctioned alternate presentations that are themselves governed and documented as derivations. If nobody can tell you which value is the authoritative one and which is the alias, the middle has already collapsed.
The claim I'd be sceptical of
Any clean causal history of this discipline, including the one I have just partly defended. The received account — regulation forced it, MDM followed — is not fabricated; it has real contemporaneous evidence, in a 2003 trade article that specified the requirement before the vocabulary existed [2], and in the vendor's own naming of the statute six years later [6]. But a category emerging across a decade, in a market with a downturn, a complexity problem, an architectural fashion and a compliance regime all pressing at once, is overdetermined by construction. Anyone selling you the single trigger is selling you a story. My co-author's real contribution here is not that he identified the right cause; it is that he declined to name one at all, which is the harder and rarer move.
Where to Go Deeper
In the order I would read them. Khatri and Brown's Designing Data Governance [3] first — it is five pages, it is open access, and its governance-versus-management distinction will do more for your programme than a maturity model will. Then DAMA International's framework [4], for the eleven knowledge areas and specifically for how it positions reference and master data management relative to governance. Kimball's 2003 column on drilling across [5] is twenty-three years old, free, and the clearest short statement of conformed dimensions anyone has written; read it to see how much of what we now call governance was already solved inside warehouse design. For the vendor side of the history, Microsoft's 2009 announcement post [6] is worth reading whole, because you can watch a category being framed in real time. And Dehghani on data mesh [10] for the strongest modern argument that ownership should be assigned rather than negotiated — I am unconvinced it changes the behavior, and it is the best case against my scepticism.
What I would take from it
The repeated version of this history has a statute at the front and a straight line running out of it. The version the sources support has several pressures arriving together across a decade, practices already named individually inside warehouse design, and vendors supplying a name because a name is what you build marketing on. Master data reached the bottom of the framework by holding weight, not by being listed.
If that sounds like a demotion of governance, it isn't. It is the opposite. Anybody can put a domain on a diagram. What is hard, and what nothing in twenty years of frameworks has made easier, is getting a room of nominal peers to accept a definition that costs one of them something. That was the job before it had a name, and it is the job now.