Weaving Intelligence
Show Me the Receipts: Decision Intelligence, Governed Master Data, and a Maturity Nobody Has Earned (original)
The 2026-09-03 original, preserved unedited for comparison.
The brief asked for evidence from organizations leading in AI-augmented analytics. I went and looked for it. What is actually on the shelf is a fifty-year-old parent discipline with better records than its own rebrand — and one charge in this piece I had to run against my own last article before I was allowed to print it.
The assignment I was handed contained a phrase I could not get past. It asked for evidence from organizations leading in AI-augmented analytics — as though the only open question were which of them to profile. So I went looking for that evidence before writing a word about it, on the theory that a brief which quietly assumes its own conclusion is worth checking rather than obeying.
What I found is the article, and it is not a list of leaders. It is a discipline with a fifty-year-old parent that kept careful records, a rebrand that has not started keeping any, and a serious charge from my co-author that I had to test on real pages before I could print it — because the charge is unattributed borrowing, and there is no faster way to lose a reader than to level that one while doing it. So: where the words below are his, they are marked as his; where a frame is borrowed, I name the work I took it from at the point I use it. Once, the honest answer was that I could only name a field, and I say so.
What the field says about itself, and what is actually on the shelf
Start with the strongest version of the received view, because it is a serious one. Gartner named decision intelligence a top data-and-analytics trend in 2020, describing it as a discipline that brings together several disciplines, including decision management and decision support
and supplies a framework to help data and analytics leaders design, model, align, execute, monitor and tune decision models and processes in the context of business outcomes and behavior
. The same entry made a dated, checkable forecast: By 2023, more than 33% of large organizations will have analysts practicing decision intelligence, including decision modeling
[1]. Hold on to that sentence; we come back to it.
Google formalized its own version and gave it a name around the same time. Cassie Kozyrkov, the company’s first chief decision officer, had by 2018 trained seventeen thousand employees in what Fast Company reported as augmenting data science with psychology, neuroscience, economics, and managerial science
— on the argument that decision science covers how humans decide but not the engineering, and data science covers the engineering but not how humans decide [2]. That is a real gap and a reasonable thing to build a practice around. Nothing below says otherwise.
So who is leading? The closest thing to a nationally representative measurement I could find is the U.S. Census Bureau’s Business Trends and Outlook Survey, which asks a rotating sample of businesses whether they used artificial intelligence in any business function in the past two weeks. As of the collection period ending May 3, 2026, the national rate was 19.8 per cent — 39.7 in information, 33.9 in finance and insurance, about 14 in retail trade, and 37 among firms with at least 250 employees [3]. It is independent, biweekly and large, and it has no decision-intelligence column. Nobody is running a nationally representative survey of the discipline, which means the brief’s phrase resolves to a set nobody has enumerated, measured against a benchmark nobody has published. There are plenty of case studies. Every one I found on the first several pages was published by a company selling the software, or in a journal I would not put in front of a client.
Then I ran the test that mattered. If decision intelligence is a maturing discipline, somewhere there should be a multi-decade body of outcome evidence attributing measured enterprise value to it, as distinct from the older things it is assembled out of. I searched for that. I did not find it.
What I found instead was the parent’s filing cabinet, and it is embarrassingly well kept. Decision analysis has a published applications literature surveyed across 1970 to 1989 [7], surveyed again across 1990 to 2001 in the first issue of its own journal [8], and reviewed at the fifty-year mark in the European Journal of Operational Research this year [10]. It has at least one company-level accounting of what the practice was worth over a decade, published under the title “The Value of Decision Analysis at Eastman Kodak Company, 1990–1999” [9]. I could not open that paper — the publisher’s copy and the author’s posted copy both returned nothing to me — so I cite it for what its bibliographic record establishes: a decade-long, single-company study of the value of the method exists and passed review. I quote none of its figures, because I have not read them.
That is the shape of it. The parent discipline has receipts. The rebrand has a conference track.
The first myth: the discipline is mature
I asked my co-author which claim in this space he would most like taken apart. He did not go after any of the techniques.
The ‘maturity’ of the discipline. Decision Intelligence and other ‘modern’ disciplines love to espouse their collective wisdom and produce volumes of detailed analytical texts. What they’re all missing is a multi-decade track record of empirical evidence that their wisdom produces value.
Read that carefully, because it is narrower and harder to dismiss than it first sounds. He is not saying the methods do not work. He is saying the standing being claimed for them has not been earned, and standing is a different property from usefulness. He drew the line himself, unprompted: I’m not saying they don’t have value to contribute—mostly that they’re just restating lessons learned in other disciplines over the last century.
I should say plainly that this is a stance, not a reaction to one field. He has said a version of it to me about governance councils, stewardship models and maturity assessments — all correct, all oversold — and about packaged domain wisdom generally. Third time I have heard it, and staging each instance as a fresh discovery would be dishonest.
Held against the shelf, the claim survives. In place of a fifty-year applications review, decision intelligence has a forecast — and we will come to what happened to it.
The charge I had to run before I could print it
The second half of his answer is the part that put me to work:
Human psychology and decision-making has been studied for centuries and most of their conclusions are restatements of that knowledge domain—usually without attribution.
Most. Usually. Those are quantifiers, and a quantifier with no instances attached is an accusation rather than a finding. So I traced one lineage properly.
Decision engineering — the strand that became one of the two popular accounts of decision intelligence — was set out by Lorien Pratt and Mark Zangari in a Quantellia white paper in December 2008. It is a good paper, and its claim is explicit: existing practices have not been integrated into a single process, nor have they been recognized as constituting an engineering discipline
, and the result creates, for the first time, a standardized conceptual framework
for complex decisions. Its central deliverables are decision maps, models and world diagrams — a visual formalism for reasoning about decisions under uncertainty. It credits analytic techniques that have a long, proven history in other disciplines
[4].
Which disciplines? The paper has eight footnotes. They cite the authors’ own prior work twice, a telecoms industry website, two essays on intelligence analysis, a clinical laboratory guideline, Gerd Gigerenzer on intuition, and a Pentagon memoir. Not one of them points at decision analysis. Not Howard, not Matheson, not Raiffa, not Keeney, not Simon [4].
By 2008 that literature was a quarter-century deep on precisely this construct. Howard and Matheson’s “Influence Diagrams” had, in the account of an artificial-intelligence researcher writing a retrospective on it, crystallized the state of the art in graphical models for decision making at that time
— decision nodes, uncertainty nodes, a value node, arrows carrying probabilistic and informational dependence — and had become a standard modeling tool for decision making under uncertainty within AI
[5][6].
Now the part that surprised me, and the reason this is a finding rather than a hit piece. Pratt has since gone looking for her own antecedents in public. In a 2019 post she notices that her diagram template is a template for operant conditioning as well
— antecedent, behaviour, consequence — maps her vocabulary onto it directly, calls the correspondence decades-old knowledge, and files the piece under both behavioural psychology and decision analysis. She reads it as more evidence that… they are a universal archetype
. The citation she offers for the framing is a dog-training podcast [11].
That is not plagiarism and I will not call it that. It is something more ordinary and more interesting: the prior art arrives eleven years late, and arrives as corroboration of an invention rather than as its source.
The same pattern shows up in the other popular account, and it is not a criticism of Kozyrkov either, because she attributes constantly — psychology, neuroscience, economics, managerial science, social science, decision theory [2]. Named disciplines, every one. Not a named researcher, not a citation, nothing a reader can follow to a page and check.
So the charge does not survive as written, and it survives in a form I find more damning. The borrowing is not hidden. It is that the field attributes to disciplines rather than to works — and a discipline is not a citation. You cannot follow it to a page, check what was actually claimed, or learn what was already known and by whom. The narrowed version forbids something specific: find me, in the freely readable canonical material, a named source cited beside the construct it is the antecedent of. I looked in three places and did not find one. If a reader finds one, I would like to see it, and this paragraph is where it should have been.
Which obliges me to file my own
I do not get to make that argument and skip the audit. So, the ledger for this piece.
In August I built an entire article on my co-author’s claim that evidence rarely moves anyone — that a small number gets embraced by whoever was already leaning that way, because it is not the number doing the work, it is the inclination
, and that a large one, at most, reopens a question it does not settle [12]. I presented that as a practitioner’s contrarian read on the field. It is. It is also, almost exactly, step one of the framework Google published in 2018: before you look at any data, work out what you would decide with none, because — in Kozyrkov’s words — We pretend that we don’t have a preference, but we’re really lying to ourselves
[2]. Two people arrived at the same observation from opposite ends of a career. One published it eight years before we did, and I did not say so at the time. I am saying it now, which is late, and the right response to being late is a correction rather than an explanation.
Second entry, and it is the weak kind. Later in this piece I borrow a phrase from software engineering, where writing the test before the code is an old and well-documented practice. I am attributing that one to a discipline rather than to a work, because I have not opened the source I would need to do it properly — the precise weaker form I just spent five paragraphs on, flagged rather than laundered. The concrete material at the end of the article is neither: it is elicited, abstracted at capture, and my co-author’s. And since a publication that prosecutes this charge has to be willing to wear it — if you catch this masthead using a frame with prior art and not naming it, the correction belongs on the page, not in a reply.
The second myth: you start from the model
The buying pattern I see most often runs platform, then pilot, then instrumentation, then — some months in, usually in a steering meeting — the question of what any of it was supposed to improve. Both popular accounts say the same thing about that ordering, so the prior art gets its paragraph before my co-author gets his.
Google’s published framework, in 2018, runs three steps. Decide what you would do with no additional information. Define what evidence would change that, and to what threshold — the actual cut-off, the number itself. Then check whether you can get that evidence, and if you cannot, decide which mistakes you are willing to make, because you will make them at machine scale. Kozyrkov is scathing about the middle step, rightly: That step is actually skipped a lot in industry, where people will use fuzzy concepts, they’ll never make them concrete, and they won’t own up to the fact that they’re using them. They’ll think that putting a bunch of mathematics near it fixes it
[2].
Here is what my co-author said, asked the same question:
I’d start with identifying the success and failure metrics before I’d focus on tracking anything—then build based upon proximity to those metrics. This would be similar to a test-driven design approach to decision-making, focused on output first, then back tracking to how the output was generated.
The overlap is real and I am not going to pretend it is coincidence. What he adds is not the sequence. It is two things sitting inside it.
The first is the word and. Success and failure metrics, defined together, before anything gets instrumented. In the published framework the failure side arrives third and conditionally — you reach it when the data you wanted turns out to be unavailable, and it takes the form of errors you have decided in advance to tolerate. In his version failure is a first-class artifact named at the same moment as success. That is a harder ask and a better one, because the programmes I have watched fail did not lack a definition of winning. They lacked an agreed description of what losing would look like, so nobody could say whether it was happening.
The second is the transposition. Test-driven design — his word, and I am leaving it alone rather than normalizing it to the commoner one, because designing the test and developing against it are not the same activity and the slip would be mine, not his. Writing the assertion first and building backwards to whatever produces it is a software discipline. Pointing it at a decision rather than a function is his move, and a good one: it makes the decision the unit under test, which is the whole rhetorical claim of decision intelligence, arrived at without any of the apparatus.
I owe you one honest gap. His proximity to those metrics
is undefined — proximity in what? He did not say, and I am not going to invent a definition and hand it back as his method. The reading I would use, and this one is mine, is joins: how many systems and transformations sit between the thing you are instrumenting and the metric that defines success or failure. Instrument nearest first. That is a heuristic, labelled as one.
And then he declined to endorse the outcome, which is the part a drafter is most tempted to lose:
‘Leaders’ in the field focus on a full-spectrum embrace of the paradigm, but I’m still not convinced that it produces true value to the company that can’t be achieved through other means.
The scare quotes are his. He answered the question about what leaders do differently, then refused to sign the receipt. Printing the sequence as a recommended method while quietly dropping that sentence would publish a framework its own source does not believe in, so it stays.
The third myth: you can tell the real thing by looking at it
I asked for a tell — five signs your programme is genuine, the article everybody writes. He would not give me one, and assigned a default instead:
I would characterize all ‘decision intelligence’ implementations as rebranding, BI team or Enterprise-wide, until they produced significant benefits—then I would ask how their achievements differ from those made by embracing other optimization paradigms with lower overhead and implementation costs.
Refusing the tell is the stronger move, and it is worth being clear why. A test can be dressed for. Publish five signs and within a quarter you will be shown five signs, staged, by a programme that has learned what you are looking at. A default cannot be dressed for, because it does not tell anybody what to perform. It simply declines to move until something arrives.
It also cannot be settled by observation, and I will not pretend otherwise. The phrase significant benefits
has no floor in it, so no result can falsify the position. That does not make it worthless; it makes it a prior rather than a test, which is a respectable thing for an experienced person to hold and a dishonest thing to dress as a measurement. Take it as what it is: where the burden of proof sits before anybody has shown you anything.
The second clause is where the bite is, and it is the one that gets dropped in the retelling. The bar is not did it work. Anything works given enough runway and a sympathetic sponsor. The bar is how the achievements differ from what a lower-overhead approach would have produced — an opportunity-cost bar, and almost nothing clears one of those. Operationally that is a single sentence you can put to a programme: name the cheaper alternative you considered and rejected, and the evidence you rejected it on. It forbids the commonest form of decision-intelligence success story, which is a benefit that a spreadsheet, a definition everybody agreed to, and a standing Thursday meeting would have produced for a fraction of the cost. The founding white paper takes the spreadsheet seriously as the incumbent and spends two pages on why it is a poor modelling tool [4]. It is a fair argument. It is also the comparison the programmes built on that paper are now obliged to keep making, and mostly do not.
Which brings back the forecast: a third of large organizations, practising the discipline, by 2023 [1]. That date is three years past and I could find no published audit of whether it happened. A discipline confident of its own maturity would have run that check itself, loudly, because it would have been the best marketing available to it. Three years past due with no audit is a burden of proof going unmet in public — and it is exactly the observation his default is built to survive.
What actually changed a decision
All of which would be easier to file under scepticism if he did not have two examples of the thing working. He offered them when I asked what governed data had ever surfaced that changed something. They are deliberately abstracted — no client, no sector, no scale — and I have added nothing to them.
The first was when governed data identified significant contract and software licensing redundancies across the Enterprise when various departments expense systems were tied together through the MDM system and identified no less than three resellers providing the same software packages to the company.
Master data management, if the abbreviation is new to you: the layer that decides what an entity is and keeps one agreed version of it across systems that would otherwise each keep their own. Now read what happened, because it is easy to misfile. Nobody’s data was wrong. Each department’s expense system was individually correct — right vendor, right contract, right spend, right approvals. The redundancy did not live in any of those systems. It lived in the space between them, and became visible at the moment they were joined on a common notion of who a supplier is. That is not a data-quality story; a data-quality story has an error in it. This one has three defensible answers nobody had ever put on the same page.
The second, when I asked:
The second was when governed data identified regional supplier redundancies beyond the company’s plans—including a regional supplier that was in the wrong region.
Nine words, and you do not need the rest of the sentence.
What the savings came to is not in this article, and I would rather be direct about why than let you assume the number was left out because it was awkward. It is left out because a figure attached to a real engagement is how an abstracted story stops being abstracted. Take it as directional and unquantified, which is the honest form it can take, or discount it accordingly. That is a fair response, and I would rather you did that than trusted a number I had rounded into safety.
One more thing, because this looks like a reversal and is not. I have asked him for war stories before and been turned down on structural grounds: he arrives after the decision, brought in to correct it, by which time the people who made it have moved on. That refusal was about retrospective blame. These two are about present discovery — not what somebody got wrong, but what became visible. Different question, and both answers are honest.
Now the part that matters for an argument about artificial intelligence, and it is why this piece is not simply a complaint about naming. In neither story was the model the mechanism. Nothing here was produced by a language model, an agent or a decision-modelling canvas. What produced it was a join — departmental systems tied together on a governed definition of an entity, so that a comparison nobody had been able to make became one anybody could make.
That is the real relationship between AI and governed master data, and it runs opposite to the sales deck. The master data is not the boring prerequisite that the interesting AI sits on top of. It is where the new information came from. What AI legitimately adds is reach: it will run that comparison across more entities, more often, at a cost per pass low enough that questions never worth a project become worth asking. That is a considerable gain, and it is a gain in economics rather than in insight. Point the same model at ungoverned data and it will not surface the redundancy at all, because the redundancy is in none of the records. It is in the disagreement between them, and a model that cannot tell that two supplier records are the same supplier will reproduce the invisibility faster and with better prose.
What I Would Watch For
Practitioner layer — the curator’s read on the consensus above.
The failure mode I’d watch hardest
A programme that reports the join as an achievement of the model. It happens because the two ship together: the master-data work lands, the analytics layer lands on top of it, and the first genuinely new finding arrives after both. Whichever is more expensive tends to get credited, and the more expensive one is almost never the join. The consequence is not a bruised ego — it is that next year’s budget goes to more model while the governance work quietly loses its funding, at which point the findings stop and nobody can explain why. Ask of every early win whether it would have appeared with the same governed data and no model at all. Ask while it is still cheap to ask.
The trade-off that usually bites
Defining failure metrics up front is the right practice and it is politically expensive in a way defining success is not. A success metric is a promise. A failure metric is a pre-agreed description of the circumstances under which the sponsor will have been wrong, written down, in advance, by the sponsor. That is a real cost and I will not pretend it away. It is also the difference between a programme that can be stopped and one that can only be renamed. If you can afford only one of those conversations, have the second — the success side will get defined for you anyway, by whoever writes the steering-committee slide.
The claim I’d be sceptical of
“Decision intelligence gives you a common language for decisions.” It gives you a common notation, which is a different property, and the founding white paper is unusually candid that the elements were already in the building — its claim is integration and recognition, not invention of the parts [4]. Notation is worth having; just price it as notation. The related claim I would push on harder is any assertion that a decision-modelling exercise surfaced something. Ask which system the something came out of, and how many systems had to be joined before anyone could see it.
Where to Go Deeper
Read the parent discipline rather than the rebrand. Keefer, Kirkwood and Corner’s survey of decision-analysis applications from 1990 to 2001 [8] is what a mature applications literature looks like, and its comparison against the 1970–1989 survey [7] is the thing decision intelligence has no version of; the fifty-year review in the European Journal of Operational Research [10] is the current front door. For the graphical construct, Howard and Matheson [6], with Boutilier’s retrospective [5] as the readable way in — three pages, openly available, which the original is not. On the discipline itself, go to the primaries rather than to anybody’s summary, mine included: the 2008 Quantellia white paper [4] is the founding document and is better written than its reputation, Pratt’s own blog [11] is more candid about her sources than her critics assume, and Kozyrkov by way of the 2018 Fast Company account [2] remains the least hype-prone entry point. And for who is actually using any of this, the Census Bureau’s survey [3] rather than a vendor’s adoption chart. It will not tell you about decision intelligence. That is the finding.
Back to the bandstand
I should be careful not to leave the wrong impression, because I am an enthusiast and this reads colder than I feel. The techniques are mostly good techniques. Framing the decision before the analysis, naming the threshold before the data arrives, treating a decision as a designed object — I would rather work with a team that does those things than one that does not, and it hardly matters what they call it.
What has not been earned is the standing. A discipline claims maturity by producing a record: here is what we did, here is what it was worth, here is how you would check. That is not a scandal — it is early, which is a fine thing to be, as long as nobody prices it as late.
So keep the question taped to the wall — did the decision actually get better, or just faster and more confident? — and add a second underneath it, which this article cost me something to learn. When somebody hands you a frame, ask whose it was first. If they can name a discipline but not a work, you are being handed folklore in a lab coat. And if you ask that of a publication and it cannot answer, the same conclusion applies to the publication. AI is a brilliant sideman. It makes a terrible bandleader. But even a sideman is expected to know whose tune he is playing, and to say the name before he takes the solo.