Weaving Intelligence
Costing the Promise: The Parts of an AI-MDM Business Case You Can Actually Check (original)
The 2026-08-19 original, preserved unedited for comparison.
Build the case from the mechanisms and prices you can check before you spend — not from outcomes someone else measured on a different estate.
The master-data business case has always been a hard room to work: real money to fix something nobody notices on a good day — the quiet plumbing under every report, every customer record, every model now queued up to drink from it. For years the pitch leaned on avoided pain. Then AI walked into the room and the deck got shinier: automated stewardship, self-healing data, matching backlogs cleared overnight.
So let me put the awkward part in the first hundred words rather than bury it: almost everyone costing an AI-augmented master data management (MDM) program right now is costing something they have not yet run. That includes this byline. When I drafted around the assumption that my co-author had one running, he stopped me:
I have not yet run governed LLM enrichment on a master-data foundation. Standing that up is a distinct piece of work from having run it, and the ordering matters: infrastructure has to exist before a proof of concept can prove anything, and portions can be proven on borrowed or temporary architecture long before the full platform is built.
A case built on somebody else's finished outcome is a forecast wearing a lab coat. Not an argument for waiting — an argument for building the case from the two things you can check before you spend: the mechanism, what the technique does and why, and the price, what it costs at your size. The outcome is the one thing you cannot check in advance.
The sequence, and where the money goes first
I asked how he sequences a build like this, half expecting a framework. What came back was a list, in the order he does it:
The process is the same for standing up any kind of new project: plan, scope, estimate, review, cost, review, infrastructure, prototype, implement, standardize. You only do what you are given time for. The key sequence is infrastructure before PoC before implementation before standardization. You cannot do anything without the infrastructure and, with cloud outsourcing, that is becoming a blocker.
Take those as one practitioner's order of operations. No published lifecycle matches them — the international standard for software life cycle processes says outright that it does not identify or require any specific software life cycle model
[3], and the closest analogue, the Launch phase of the AWS Cloud Adoption Framework, covers only the tail of it [2].
What governs the business case is the stretch in the middle — infrastructure, then proof of concept, then implementation, then standardization — because most decks skip the first item and start their arithmetic at the second. The clause he attached matters — portions get proven on borrowed architecture long before the full platform exists. Proving them early is often the smartest money in the program, but you cannot then book that result as proven on the platform you have not built. The proof transfers; the cost profile does not.
Infrastructure is the constraint now: with cloud provisioning in the path, the gap between "approved" and "an environment a team can build in" is a chargeable stretch of calendar, invisible in every ROI model because it produces no deliverable. Just a wait. So three line items belong in the case before a single model is called: the platform build, the temporary architecture, and the wait. Leave them out and the payback date is fictional. Gartner expects 80% of generative-AI business applications to run on existing data management platforms by 2028 [4] — if that holds, the governed layer is the launch pad, and launch pads are capital.

What you can cost before you have run it. Almost every input can be costed honestly up front. The one that cannot is the outcome — and it is the one you will be offered.
The large teal-edged panel takes most of the frame. It is headed Checkable before you have run anything and counted, at its right, as seven items. Each line names the item and how you get the number:
- The platform build — quoted
- The temporary architecture — quoted
- The wait for provisioning — scheduled
- Defining the output standard — chosen
- Continuous source access — quoted
- The labelled ground-truth set — counted
- The clerical review queue — measured
Set apart to the right, in grey and a fraction of the size, is the second panel: Not checkable, holding The outcome, one item. The note beneath it reads: this is the one a vendor-commissioned study offers to supply — and it arrives as a ceiling, not a measurement.
The data prerequisite is a standard, not a clean slate
Here is where I part company, gently, with the received wisdom, and only about one case. If you are pointing AI at your master data to fill in the gaps, "clean the data first" cannot be the prerequisite; the gaps are the work. What you need is narrower:
The two real prerequisites for gap-filling enrichment — and what each costs:
- The expected output structure. The target schema, required attributes, value sets, format — in most estates inherited and frankly legacy. This is the standard you are asking the model to help the data conform to; without it there is no way to say whether a filled gap was filled correctly. Cost: modelling and definition work, up front, and work you owed the MDM program anyway.
- Source access. Not once — continuously. The number of sources only ever goes up, and each carries its own negotiation, credentials and owner to be persuaded. Cost: a recurring line item, usually booked as a one-time integration and then re-bought every quarter.
That reframe leaves the benefit ledger untouched and changes the effort curve underneath it: define the standard, open the sources, then let the model close the distance between what you have and what the standard requires. The standard is what is under the model — skip it and what you have automated is agreement.
Gartner's figure — an average of at least $12.9 million a year in poor-data-quality cost per organization [1] — sits on more MDM slides than any other statistic, and it is weaker than the slides make it look: the population underneath is 154 reference customers across sixteen data-quality vendors, nominated by those vendors, asked to estimate what poor data quality was costing them. Self-reported, unaudited, from a panel already buying the software. Evidence the category is expensive, never your estimate. A borrowed benefit is the same defect as a borrowed outcome; it just arrives wearing a citation.
The cheapest dial in the deck is the one everybody turns to maximum
Now the price side, where most of these cases are wrong by an order of magnitude — in both directions, which is the interesting part. In most of them the work is repetitive and time-consuming rather than difficult: normalize a supplier name, classify a part, propose a value for a field empty since the last migration. Work of that shape needs basic domain competence, and competence of that kind you write into the prompt.
Ask my co-author to compress his cost approach and you get four moves:
If you wanted to encapsulate my optimization approach: start one generation behind the current main models and two generations behind the bleeding edge. Then use the minimum token budget available per request. Focus on prompt engineering, and organically build your gates as you go, starting with what you have already developed as what your company considers basic guardrails.
Three of those four I would still put in a cost model: the minimum viable token budget, measured rather than assumed; prompt engineering, the highest-leverage cost reduction available and a line item nowhere because it looks like writing; and gates grown from what your organization already treats as basic guardrails.
On how big a model needs to be, he declines the tidy answer and is right to: The data science answer is "lift" — and if that is unsatisfying, welcome to data science. The practical answer is that it depends, and that is the main part of model tuning.
Start small, scale up, stop where returns flatten.
The fourth move is a price claim, so it is checkable
Three of those four moves are judgement. "Start one generation behind" is not: it is a claim about published prices, checkable before you spend a dollar. So I checked it on three vendors' own price lists, all read August 18, 2026. They give three different answers, and the disagreement is the useful part.
| Vendor | Current | Going back | Shape of the ladder |
|---|---|---|---|
| OpenAI [16] | $5.00 / $30.00 | one back identical; two back $2.50 / $15.00; four back $1.25 / $10.00 | descends with age — the rule pays, up to 4× in and 3× out |
| Anthropic [17] | $5.00 / $25.00 | five consecutive versions at that same price; then the retired ones at $15.00 / $75.00 | flat, then a step up — the rule costs 3× on both |
| Google [10] | $0.75 / $3.75 | one back identical; two back $1.50 / $9.00; three back $0.30 / $2.50 | zigzag — nothing, then double, then 60% off |
So the heuristic is not wrong. It is local. On one of these sheets it is worth a multiple; on the second it would triple your bill; on the third the answer depends entirely on which rung you land on. And the direction is not even stable inside one vendor: on Anthropic's list the flagship tier steps up as you go back, the mid tier is 1.5× dearer one version back, and the small tier runs the other way — the retired one is 20% cheaper than the model that replaced it [17].
What the three lists have in common is structure rather than policy: each model carries its own price, set when it shipped. The slope you see across generations is therefore not a rule about age. It is the trace of what each successive launch happened to be priced at. Where that trace descends, "start one generation behind" is real money. Where a vendor holds one price across a whole family, or ships a successor cheaper than its predecessor, the trace flattens or inverts and the rule stops meaning anything. Vintage is not a lever. It is a lookup — per vendor, per tier, on the day.
Two dated step changes belong in your model rather than in your memory of it, and they point in opposite directions. Google's current rate is promotional and doubles on 2027-01-01, which collapses the two-generations-back premium to roughly parity [10]. Anthropic went the other way this month: an introductory rate on its mid tier was scheduled to rise by half on September 1 and has instead been made the standard price [17]. Either way the number is published, dated, and knowable before you commit — which makes it exactly the kind of fact this article is asking you to put in the spreadsheet.
The lever that holds everywhere
Which leaves the lever he added when I showed him the draft:
Intentionally "smaller" or "cheaper" current generation models are viable too. They're options to be considered in addition to older generation models. Your mileage (and token efficiency) may vary and really can only be guessed at without empirical tests.
He is right, and this is the lever that behaves the same way on all three sheets — including, unlike vintage, within a single generation, so there is no confound between age and size to argue about. OpenAI's current generation ships a flagship and a small variant released together: $5.00 in and $30.00 out against $0.20 and $1.20, which is 25× on both, same generation, same day [16]. Anthropic's current family runs $10.00 / $50.00 down to $1.00 / $5.00, a 10× spread [17]. Google prices Pro, Flash and Flash-Lite separately and runs 20× on input and 30× on output top to bottom of its sellable list, widening to 40× and 45× on long prompts [10]. Microsoft's catalogue shows the same shape as a product structure — a mini, a nano, a standard and a pro in one series [9]. The saving is in the tier. One caveat, and it is the ordinary one: three lists on one morning. Read your own rows, on your own vendor, and date them.
That leaves the older model as an option you price like any other rather than a lever you pull. The two choices fail differently:
- An older frontier model. Sometimes you buy last year's top-tier reasoning at a discount; sometimes, as above, at a premium. Check the rows rather than assuming the direction. Either way you give up what shipped after it: better instruction-following, current structured-output and tool behaviour, current safety defaults.
- A smaller model of the current generation. You keep the current tooling and pay a published fraction. What you give up is headroom: it holds less of your domain at once and degrades sooner as records get stranger — quietly, convincing on the easy ninety percent while missing the hard tail. That is the failure the next section is about.
Which one bites depends on the shape of your data, which is why I will not rank them, and neither will he. Token efficiency and quality-per-dollar cannot be known without a test on your own records, and a published benchmark will not settle it, because contamination inflates the score: when rephrased test data survives into training, a 13-billion-parameter model has been shown reaching drastically high performance, on par with GPT-4
[11].
So the line item in your cost model is a bake-off: a few thousand of your own records, labelled once, run through a frontier model as the ceiling, the generation behind it, and one or two small current models, scored against the same labels and divided by what each cost to run at the price published the week you ran it. What the bake-off adds on top of that labelled set is cheap; the labels never are.
The implication is blunt: a return computed at frontier-model prices is not a fact about the technique. It is a fact about a procurement decision.
But do not let that multiple make the decision for you. Whether you can run on the small model turns on things the price list has no opinion about: how long and how strange the tail of your records is, whether anyone exists to staff the review queue it will fill, whether the source data is there to give it context at all. A cheaper model that pushes twice as many records into a clerical queue you cannot staff has moved the cost to a line the price list does not print.

Three settings, not three givens. The cost side moves by a multiple. For repetitive, high-volume work the benefit side barely moves at all.
Three tracks run in parallel. On each, a hollow marker at the far left is what the deck assumes, a filled teal marker about two thirds along is where to start, and the teal bar between them is the move being recommended, labelled above with what that move does. Past the filled marker the track turns grey and a dashed gold tick marks diminishing returns: there is further to travel, and it stops paying.
- Model tier, which model you send it to. Over the bar, cost falls sharply. Assumed: bleeding edge. Start here: one generation back.
- Token budget, how much you send and accept. Over the bar, cost falls proportionally. Assumed: generous by default. Start here: the minimum the request tolerates.
- Prompt effort, how much domain competence is engineered in. Over the bar, the one dial you turn UP — the only one of the three the piece asks you to increase. Assumed: a thin prompt. Start here: domain competence engineered in.
A gold band beneath the three carries the Shared readout — benefit indicator, which holds across the sweep: a line of evenly spaced points crossing the same span the dials do, level from end to end. Under the band: holds for repetitive, high-volume work — which is exactly what measurement is for. Boundary: this is advice for cost-restricted environments.
Where the modelled savings leak back out
Every gain above ships with a failure mode. The cases that disappoint forgot to cost the catch.
The failure is the one the papers warn about
The failure he keeps meeting is the one the papers warn about:
You are going to be disappointed, because it is the one the papers warn about: confident matches based on limited, incorrect or insufficient domain knowledge. AIs do not know what they do not know, nor are they very good at questioning what they have been provided. If you ask an AI to match on a particular domain and the only context you give it is the raw data set, it will not know anything outside of it. So it will not see anything wrong with matching Buicks, GMCs and Fords in the same data set of deciduous trees as the Larch, or the Larch, or the Larch. It will not flag car brands as errors, because it does not know what a car brand is any more than it knows what a tree is. It just knows it was asked to classify data, so it will do that.
The larch is his Monty Python joke. And here is the part he would not soften, nor will I: That is not an AI problem. That is a user problem — someone set the AI up to fail by not giving it enough context to understand what it did not understand.
You cannot fix a model. You can fix a context window.
The confidence signal was there — the pipeline threw it away
This next part is mine rather than his, though he ratified it. "The model is overconfident" describes the model. The repair, if there is one, has to be somewhere in the pipeline. Scoring systems generally do produce an uncertainty signal — a probability, a distance, a link score — and in most pipelines it is dropped. A classify-this-list call asks for a label and gets one; nothing requests a confidence or routes a weak answer differently from a strong one. The system had the signal all along. Nobody asked it for it.
Nor is that new. The foundational decision rule for record linkage — Fellegi and Sunter's, from 1969 [8] — was three-way from the start: above an upper threshold a match, below a lower one a non-match, and in between, in the Census Bureau's summary, designate pair as a possible match and hold for clerical review
, with cutoffs determined by a priori error bounds on false matches and false nonmatches
[7]. That middle band is the design, not an admission of defeat. A pipeline emitting a label and no band has dropped the part that made it accountable.
So: surface the confidence, set thresholds from the error rates you can live with, route the middle to a human. Three things go on the ledger for that: two numbers, a queue, and a steward's time on the ambiguous minority. I am not going to tell you the queue is cheap, because I have not measured one and its volume is set by the threshold you have not chosen yet. What it prevents is easier to argue for: a false merge does not announce itself, and unwinding it costs more than every easy match it hid behind ever saved.
Be honest about what sets that threshold: it is never only the error economics. A threshold is where the arithmetic meets the review team you can staff, the risk your board and your regulator will wear, the mandate, and — plainly — how much money there is. Record which of those set yours, beside the value.
And measure properly: two numbers against a labelled set, not one comfortable percentage. Precision and recall pull against each other by construction — as Google's machine-learning course puts it, they often show an inverse relationship, where improving one of them worsens the other
, and which you favour depend[s] on the costs, benefits, and risks of the specific problem
[5]. AWS gives the asymmetry in business terms (healthcare favours precision, advertising favours recall) and is blunt about the prerequisite everybody skips: a manually annotated ground-truth set built from your edge cases, which most organizations do not have [6]. Budget for building it; every number above is meaningless without it.
The gate that is not enforced is a suggestion
Left alone, a model will cut every corner available to it: shortcut-taking against a proxy objective. That is why an enforced gate works: it changes what the proxy rewards. The costing point is one word long — enforced. If your case books reduced rework because "we will review the output," you have booked a saving against a hope. Book it against a gate that fails loudly, or do not book it.
The merge you cannot explain
“The model was confident” is no answer when a regulator asks why two identities were fused. I went looking for the rare-event story behind that sentence and was told the premise was wrong:
That question is constantly asked and if you don't have a ready answer, you're in a lot of trouble and headed down a rabbit hole without a flashlight. You either pay up front as part of the build or you really pay—and you don't know how much you pay until the bill comes due—down the road.
So I asked what the up-front half costs, expecting a fortnight of work to point at, and got a zero with a condition on it that carries the whole load: In a properly scoped and implemented project, there's no additional cost because implementation is the cost you have to pay up front.
Describing the mechanism behind a merge is part of building the merge. Skip it and you lose two things, and only the first is the compliance one: the logic is indefensible when challenged, and it also cannot be modified when the requirements change. A merge rule nobody can explain is a merge rule nobody can safely amend.
I asked for a per-merge retention figure and did not get one, which fits the account: if the description is part of the merge, there is no separate meter to read. That does not take the line off your model. It moves it. And the deferred version has no number either, for a harder reason — whoever is asking decides how far back the answer has to reach.
Where the line goes is the part a costing exercise has to settle, and he argues against the obvious home for it. He has never seen explainability budgeted anywhere and thinks it should be. The usual suggestion is documentation. His is architectural design, because that is where the implementation lives, and with it the flexibility to answer a question the business has not thought to ask yet. There is real work on explanation methods built for entity resolution too, producing saliency scores and counterfactuals showing which values would have flipped a decision [14] — useful, and a different artifact from the written rule an auditor asks to see.
The mess relabelled "AI-ready"
I have sat through the demo where a team pointed a model at a mess, called it AI-ready because the slide looked good, and shipped the errors faster. It is the most expensive slide in the deck, and the abandonment numbers are a running tally of it: 63% of organizations lack or are unsure of the data management practices their AI needs, and Gartner expects 60% of AI projects unsupported by AI-ready data to be abandoned through 2026 [12].
Read those as cost forecasts rather than warnings: the modal outcome of pointing AI at an ungoverned estate is that you pay for the pilot and keep the mess — which makes the governed layer not a precondition to the business case but a large, uncredited share of the return inside it. A companion prediction is checkable: in 2024 Gartner held that at least 30% of generative AI (GenAI) projects will be abandoned after proof of concept by the end of 2025, due to poor data quality, inadequate risk controls, escalating costs or unclear business value
[13]. That horizon has passed and I have no independent measurement of how it landed. But it is the kind of claim that can be marked, which is the whole difference between it and a vendor's slide: prefer a forecast with an expiry date.
What I Would Watch For
Most of what is above is the field's argument as it is usually made, sourced where a source exists; where I step outside it, I say so. Everything in this box is judgement.
Here is where I would expect a case like this to go wrong.
- The vendor-commissioned ROI study. The Total Economic Impact-style report is the genre to be most careful with, and its problem is not dishonesty — it is selection. The vendor commissions it and the analyst interviews customers the vendor supplied — by construction, the ones for whom it worked. The mechanism described is usually sound; the number attached is a ceiling somebody was paid to find.
- A case carried by a single figure. When the only thing holding a request up is a payback month, ask what becomes of it when the sponsor changes, when the integration team is booked, or when the cash is not there. And when somebody quotes a threshold, ask what it is made of.
- The one-time install. Thresholds need re-tuning, ground-truth sets maintaining, and somebody watching precision and recall so a quiet regression is not next year's incident. Budget for the feeding as well as the purchase.
- A benefit with no baseline. The commonest defect I see: a "before" number that is a guess and an "after" that is a demo. Capture steward hours per thousand records, duplicate and survivorship error rates, and the downstream cost of one bad record — before a vendor touches it.
Where to Go Deeper
Whose work I would actually read:
- Ivan Fellegi and Alan Sunter, A Theory for Record Linkage (Journal of the American Statistical Association, 1969) [8]. The foundational treatment, and still the clearest account of why matching is a decision under uncertainty with two error types you trade between.
- William Winkler, Overview of Record Linkage and Current Research Directions (U.S. Census Bureau, 2006) [7]. Free, and the most useful thirty pages in the subject: the Fellegi–Sunter rule in implementable form.
- Peter Christen, Data Matching (Springer, 2012) [15]. The standard practitioner reference; read his treatment of evaluation before writing a matching requirement.
- Teofili, Firmani, Koudas, Martello, Merialdo and Srivastava, Effective Explanations for Entity Resolution Models (2022) [14]. The best pointer for making explainability a specification you can hold someone to.
- The cloud vendors' own accuracy-measurement guidance, of which the AWS Entity Resolution write-up is a fair representative [6] — read it for the discipline.
What I would not point you at is an authority on the economics of this, because I do not think an honest one exists yet. The techniques are well documented; the cost models are barely two years old and the tier ladders move weekly. Measure it yourself, and put a date on your price list.
The bottom line
The good news first, because it is real: the MDM business case is genuinely stronger in the age of AI than it was five years ago, and work that was structurally unaffordable got cheap enough to attempt.
Then one correction, because every price in this piece quietly understates. A unit price is not a bill. The cheap classify call rarely stays a classify call: add the confidence band, the second pass, the explanation, the routing, and one request becomes five. Cost the workflow you end up with, not the call you start with — I have just spent a section telling you to add every one of those.
Then the lines themselves. Here are the ones I would put in the model first, gathered so you do not have to reassemble them from the sections above: the platform build; the temporary architecture; the wait for provisioning; defining the output standard; continuous source access; the labelled ground-truth set; and the clerical review queue. Seven — and explainability is deliberately not the eighth. It is not a line you add; it is a property of how you build the platform build line, which is why the answer to what it costs was a zero with a condition on it. Budget it there or you will pay for it somewhere you cannot see. The rule I have applied to everyone else here applies to this list too — nobody gets a number I have not measured — and I have sized exactly one of them, in records rather than money.
My co-author is standing one up now, and when it has run the follow-up is scheduled: these same seven lines with our own figures, assumptions and volumes against them, within the quarter. If it has not arrived by then, you are owed an explanation.