Weaving Intelligence
The Number Was Never Doing the Work: AI, BI, and the Limits of Better Evidence (original)
The 2026-08-20 original, preserved unedited for comparison.
Decision intelligence is sold on a premise the decision research stopped believing decades ago — that decisions go wrong because the evidence arrived too slowly. Here is where AI in business intelligence actually earns a yes, where it quietly doesn't, and how to tell the two apart before you buy.
Somewhere in your organization is a beautifully built dashboard that has never once changed anybody's mind. Not for want of attention — people look at it constantly. I mean something narrower and less comfortable: no decision in its whole history came out differently than it would have if the thing had never been built. Consulted a thousand times, and never once the reason.
That isn't a defect; it's close to what dashboards are for. And it's the fact you have to sit with before talking usefully about what AI changes in business intelligence, because the pitch for decision intelligence rests on one assumption — that decisions come out badly because the evidence arrived too slowly. I don't hold that assumption, and a programme built on it ships something fast, fluent, expensive and beside the point.
One note on method: this argument was built against a practitioner's answers rather than assembled from the literature and decorated with them. Where the words below are my co-author's, they're marked as his.
What the field means by decision intelligence
Start with the consensus; it's better than its own marketing. Gartner, naming it a top data-and-analytics trend, put it this way: decision intelligence brings together several disciplines, including decision management and decision support
, and provides a framework to help data and analytics leaders design, model, align, execute, monitor and tune decision models and processes in the context of business outcomes and behavior
[1]. The load-bearing word is discipline: the object you instrument is the decision, not the report. The AI half arrives through what the same forecast calls augmented analytics — the machine doing the pattern-finding and the explaining, so that the most relevant insights will stream to each user based on their context, role or use
rather than waiting behind a dashboard somebody has to go and open [4], now scaling fast enough that Gartner expects generative AI to shape three-quarters of new analytics content by 2027 [5].
And here is where the field argues with itself. Cassie Kozyrkov, who did more than anyone to name the discipline, frames it as a fusion of data science with the social and managerial sciences, precisely because the hard part of a decision isn't the math [3]. That puts the bottleneck in the human; the tooling market sells the bottleneck as the data. They are not the same claim, and which one you believe determines almost everything about what you should buy.
The premise underneath, and why I don't hold it
I wanted a sequence: what happens between a number landing on a screen and a person changing course. I put it to my co-author expecting a mechanism. He declined the premise.
Numbers don't move people — or, more accurately, small numbers don't move people. All changing a number can do is exert pressure in a certain direction. If the person's inclination is already in that direction then the number will be embraced and used to support that decision — but it is not the number doing the work, it is the inclination.
I nearly built this article on that paragraph alone, and it would have been the wrong article — because when I read it back as numbers don't move people, full stop, he corrected me. Numbers move people, he said, when they are reinforcing some other position someone already holds, or they're of sufficient magnitude to force them to reconsider their position
.
I got that wrong too, on the first pass, and the way I got it wrong is worth your attention because it's the mistake the whole category makes. I read the second half as the escape hatch — that a large enough number simply wins. It doesn't. He was precise about it: It doesn't necessarily overturn it, at best it prompts reconsideration. The number combined with other pre-existing factors is what tips the balance and the number is just one of those factors.
So a small number ratifies an inclination that already exists, and a large one, at most, reopens the question. Reopening is not settling. Even the big number doesn't decide anything by itself — it lands among everything else already in the room, and the balance tips on the combination. There is no size at which evidence becomes the mechanism. That is the whole claim, and nearly every analytics pitch you will hear this year assumes the opposite.
I want to hand you the exception, because I went looking for the thing that would prove that sentence wrong and I found it. There is a setting where a number routinely overturns the decider's own conviction, and it is the one place you would expect the inclination to win most easily: a team ships a feature it believes in, and a controlled experiment grades it. Microsoft's account of running that at scale reports that only about a third of well-designed experiments improve the metric they were built to improve; roughly a third actively hurt it and are stopped, and the rest come back flat and are also stopped [20]. The builder wanted it. The number said no. The number won, over and over, as a matter of routine.
Look at why it won and the claim survives in a better form. Nothing about that number is bigger or cleverer than the ones being ignored elsewhere in the building. What is different is that the organisation agreed, in advance and in writing, what it would do when the number arrived — which metric, which threshold, which action — and then could not renegotiate the rule once it could see the answer. That is not an evidential act. It is a structural one, taken before the evidence existed. And the same account notes it was not an easy sell internally: when we first shared some of the above statistics at Microsoft, many people dismissed them
[20]. They had to build the platform, and then win the argument about believing it.
So the corrected claim is narrower and considerably more useful than the one I started with: evidence becomes the mechanism only where a binding decision rule was fixed before the number arrived. Absent that rule, size does not rescue you and the inclination decides. What defeats ratification is procedure, not magnitude — and notice this is exactly what the discipline the article is arguing with already prescribes: fix the default and the cut-off before you look [3]. The pitch that assumes the opposite is assuming a faster number will do the work of a decision made in advance. It won't.
That's contrarian for a BI column, but less so than it sounds: the decision research got there first. Feldman and March argued forty-five years ago that organizations gather more information than they use and then ask for more — information works as a signal of competence and a justification for choices while pretending to be an input to them [8]. Klein found the same shape from the other end: experts under pressure don't score options against each other, they recognize a situation as typical and simulate one course of action, the analysis arriving after the recognition [9]. And Kahan and colleagues found numeracy rescued nobody — it helped when a result fit a subject's priors and stopped when it cut against them [10].
The seams are worth naming. Kahan's work uses politically charged data, not a capital-approval meeting; and the wider "facts don't change minds" story has taken a beating, with Wood and Porter finding across five experiments and more than ten thousand subjects that people do update factual beliefs when corrected [11]. That distinction is the useful part: beliefs update far more readily than choices do. It's the old stated-versus-revealed-preference gap in analytics clothing: Webb and Sheeran's meta-analysis found that a medium-to-large shift in intention produced only a small-to-medium shift in behaviour [12]. Your stakeholder isn't lying in the readout; he's telling you an intention, which is a weaker predictor than either of you would like.
Now apply that to a machine producing evidence at near-zero marginal cost. If evidence mostly ratifies, making it faster and more fluent doesn't get you better decisions — it gets you better-supported ones. Generative analytics is exceptionally good at assembling a defensible-looking rationale for a position somebody already held, and that isn't an abuse case to design out; it's the most natural thing the tool does. Which is the whole of my argument in one image: AI is a brilliant sideman and a terrible bandleader. A sideman feeds you a line you'd never have found — but he doesn't choose the tune, and the bandleader walked in humming it.
Where AI actually earns a yes
So if the pitch isn't "better decisions," what is it? The AI work that gets bought, and stays bought, sits below the decision-consideration threshold — the point at which a problem is worth anybody's deliberation. Take a Director who knows there's a leak: volume going missing somewhere in the supply chain, not dramatically, just persistently. Nobody finds it during a normal working day, because nobody's day has room to go looking — and the loss is too small to justify a dedicated person or a project. The cure costs more than the disease, so it sits there being irritating. Now put an agent on it: roughly a week of a dedicated resource's time, no spin-up, no training, and it can trace the data and find the leak. Given discretionary budget, that sale closes itself.
Notice what did the work. Not a better number: a long-standing irritation belonging to a specific person, priced under the amount that would have forced a conversation. Sales books call this making it easy to say yes — folklore rather than research, and none the worse for it. But the mechanics matter, because that is where people get it wrong: what must fall below the ongoing cost is the cost to the person deciding, not the cost to the company. Two different numbers, and only one is in the room. A project that solves something personally annoying to that person, as a byproduct, clears far more easily than one that doesn't — and more easily still when the money comes out of another department's budget. Read that as a diagnostic rather than a tactic. If you are the one being sold to, it tells you what you are actually watching when a proposal sails through: not necessarily the strongest case in the room, but the one whose irritation belonged to the person deciding. That is worth knowing in both chairs, and it is uncomfortable in exactly one of them.
The qualifying question I'd actually ask. Not "what is this worth to the enterprise" — arguable numbers lose to inclination every time. Ask: whose recurring irritation does this end, and is the price below the threshold at which they'd have to justify it? If nobody in the room can name that person, you are not looking at the strongest case in front of you. You are looking at a business case.
Where it doesn't: the conversation nobody was having
The most-demoed capability in this category is the plain-English query box, and also where I'd expect the quietest failure. Two problems sink the "conversation with your data" pitch: the conversation was never happening — people do not, ordinarily, talk to their data — and token for token, natural language is among the more expensive ways to spend an AI budget, on an interface rather than an answer. I can't hand you a study proving these pilots fail; nobody has run one. What I can point at is the adoption record any pilot has to overcome: business-intelligence adoption sits at roughly a quarter to a third of employees, broadly flat across a decade of tooling investment [13]. The argument here is consistent with that record, not proven by it.
For a plain-English interface to pay off, something cultural has to happen first: the business has to genuinely want to talk to an agent about its numbers. On the executive floors that's plausible — a Director's day is already conversation, so one more costs nothing. A floor down, the premise falls apart. Analysts don't run on conversation, and I'll defend that rather than soften it. As he put it when I pushed on whether it was fair: I know I am embracing a stereotype here, but self-selection and survivorship are real things
. The people who choose this work skew toward preferring the terminal to the meeting. And more importantly, they're already deep in that conversation. It just happens in SQL, in Python, in R. The less technical hold up their end through Excel and Power BI.
Which leaves a real problem solved for a small, senior population at a price set by the whole estate. That is how a pilot fails by succeeding — his phrase, and the right one. It works, everyone is pleased, and the cost per answered question never gets anywhere defensible. Public-facing versions last longer, measured against the price of outsourced support rather than a licence: a longer runway, which is a polite way of saying a bigger budget.
Point it at the line: what "load-bearing" actually means
If you can only govern this properly in one place, the question becomes which reporting is load-bearing — and which sits on the line — the top line of the income statement, the money coming in from sales before any costs are taken out, as against the bottom line left after they are [15]. I brought him a test I'd been handed elsewhere: tie the metrics to the people whose performance they measure, watch adoption climb, call those load-bearing. He didn't take it, and what he gave instead is sharper.
If the dashboard's CDEs are company KPIs that are Reported, it is load-bearing. If dollars flow in or out of the company based on what is displayed on the screen, it is load-bearing. Surprisingly, it is not 'if people get paid based on those numbers, it is load-bearing.'
CDEs are critical data elements — the handful of fields the whole thing turns on. KPIs are key performance indicators, the measures an organization formally tracks to judge itself against its own goals. And Reported, capitalized, is carrying weight: the number leaves the building — a filing, a board pack — rather than merely appearing on a screen. Two tests, then, and one plausible test explicitly excluded. The exclusion is worth sitting with, because it looks like it belongs. Plenty of numbers have bonuses attached — ticket clearance rates, on-time deliveries, customer satisfaction surveys, work and vacation hours. They're monitored, they're incentivized, people care intensely, and they still aren't load-bearing. Compensation pressure and structural importance are different properties that happen to sit next to each other.
Push both tests hard and you arrive somewhere uncomfortable: the reports that drive or support sales are very nearly the only load-bearing ones in the building. A company lives on revenue; anything keeping that flowing is structural, anything outside it is at best useful. Other things can be levered into importance for a while — an ESG score, a supplemental stock offering, a line of credit, a bond issuance — and while they are, they genuinely are. But they're temporary and the line is not. As he put it: the core of the business is always, always, always the line, and businesses that forget that almost always learn to regret it
.
The semantic layer is not an honesty layer
You will be told, repeatedly, that the semantic layer is what keeps a language model honest: a governed, central place where the business's metrics are defined once, so "revenue" means the same thing in a dashboard, an API call or plain English [7]. Point a model at the raw warehouse and it guesses at what your words mean; point it here and it inherits an agreed vocabulary. True, and worth building. I put that framing to him and got the bluntest answer in the exchange.
Semantic layers are anything but honest. They exist to present carefully curated stories as facts. At its heart a semantic layer is nothing but a set of aliases and abstractions constructed as an overlay on top of an existing data model, designed to reframe and recharacterize the data into the format preferred by the consumer.
He is describing something that has since been put to the test. A 2023 workshop paper ran the same handful of queries through five production tools — Tableau, Power BI, Looker, Malloy and Sigma — and found the join path, the deduplication method and the null handling chosen per query by heuristics the tools do not disclose, concluding that they hide from the analyst the ability to interpret and control how the final metrics in a query are decided upon
[19]. Curated, and presented as fact.
That's not a complaint about the people who build them, and it isn't a fringe reading. The 1991 patent behind the category describes showing users "terms that he is familiar with in his daily business" instead of "data organized in a computer-oriented way" [16]: translation is the feature. And the vendors are entirely candid about the second half of it. SAP's own training material for the tool that started this category says the universe exists to hide the underlying physical data storage from the business user
, and that the SQL it generates runs invisible to the business user
[17]. Hidden and invisible are SAP's words, not his — offered as features, which is exactly what they are. The curation is the product. Airbnb published the bill. Before they centralized, their CEO would ask which city had the most bookings last week, and Data Science and Finance would sometimes provide diverging answers using slightly different tables, metric definitions, and business logic
[18]. Not a governance failure and not incompetence — a first-rate data organization, two lawful translations, two numbers.
Choose a different source table and you get a different answer [19] — and nothing in the interface tells you a choice was made. That is the part worth sitting with. The judgment calls are real, they are consequential, and they are invisible by design rather than by neglect. Point a language model at that and the fluency rests on decisions nobody in the conversation can see.
So build one, and don't file it under truth. It's an agreed translation, shaped by what the consumer prefers, now feeding a model that will express that preference fluently, at volume, with citations and lineage.
Governance for a world where evidence ratifies
The discipline exists already: the NIST AI Risk Management Framework gives a vendor-neutral spine in four functions [6]. Re-aim each at the premise above.
- Govern — not only who owns the model but whose judgment it will end up supporting. Decide what may run unsupervised before somebody discovers by accident that it already is.
- Map — what does confidently wrong cost here? The load-bearing tests are your map: an assistant on the sales line and a tool that summarizes ticket queues are not the same risk.
- Measure — calibration, not just accuracy. Then the awkward one: how often does an AI-generated insight change a decision, versus support one already forming? Almost nobody keeps that ratio, and it's the most honest metric here.
- Manage — put scrutiny where the dollars flow, and have a plan for the day it's wrong in public.
Underneath it, one conviction does most of the work: an insight you can't interrogate is a rumour with a chart attached. Lineage and definitions travel with the answer; confidence is shown, not implied. Under a ratification model that matters more, not less — the dangerous output isn't a wrong number, it's a wrong number arriving exactly when somebody needed one.
What I Would Watch For
Practitioner layer — the curator's read on the consensus above.
The failure mode I'd watch hardest
The pilot that succeeds. Adoption up, query volume up, scores warm, sponsor happy — and if you audit a quarter of decisions you can't find one that came out differently. That's a fine outcome that will be reported as a transformation, and the gap between those stories is where next year's budget gets set. Ask for the counterfactual while it's cheap.
The trade-off that usually bites
The AI work with the cleanest business case is priced below the decision-consideration threshold — and the threshold at which something gets governed is higher than the one at which it gets bought. So the wins accumulate where nobody is watching: small, discretionary, department-funded, individually harmless, collectively an unmapped analytics estate with a budget code. Make your safeguard an inventory, not an approval gate; gates get routed around by the logic that created them.
The claim I'd be sceptical of
"Our semantic layer keeps the model honest." It keeps the model consistent, which is a different property — and not guaranteed across the organization either. He made the point that these layers often hand different answers to different groups, because each group's reporting and categorization requirements are baked into what it sees: Two numbers for the same metric can both be 'right.'
Related: treat "AI-driven decision-making" as a category label, not a description.
The story I can't tell you
I asked for the scar — the time a fast, confident, wrong answer got acted on and cost somebody something. He declined the question before answering it: The way you are asking that question prompts me to answer: every time.
Then the substance, broader than any anecdote: fast, fluent and confident answers that are unsupported by data or experience are frequently wrong and, all too often, acted upon
— which is, he added, more or less what business intelligence, data warehousing and master data management claim to be worth.
The research supports the shape of it: people reporting complete certainty are right something like seventy-five to eighty-five per cent of the time, and at ninety per cent confidence about seventy-five [14]. Confidence is not evidence of accuracy. It's evidence of confidence — and a language model produces it by default.
The anecdote isn't his to give, for a structural reason worth more than the story. He arrives after the decision — brought in to correct it, by which time the people who made it have moved on. The correction is what's left. Nor do I have an account of where an inclination comes from; I can't tell you how a position forms, only what a number does alongside one that already has. That's a real gap, and I'd rather name it than fill it with a framework.
Where to Go Deeper
On how experts actually decide: Gary Klein, Sources of Power, and the naturalistic decision-making tradition around it [9]. On why organizations demand information they don't use, Feldman and March's 1981 paper is short, forty years old and still the sharpest thing on it [8]. On whether numeracy protects you, Kahan and colleagues [10], with Wood and Porter as the honest counterweight [11] and Sheeran and Webb [12] on why intention and behaviour part company. Kozyrkov [3] is the least hype-prone entry to the discipline itself, and for governance work straight from the NIST framework [6] rather than a vendor's reading of it.
Back to the bandstand
None of this is an argument against the technology. I'm an enthusiast, which is exactly why I'm careful with it. Used honestly, augmented analytics is remarkable — it will chase down a problem that was never worth a project and free your best people from assembling the obvious. Take all of that; it's real. Just measure it by decisions changed rather than dashboards shipped, and expect the honest count to be low. What it will not do is make an organization decide better by handing it faster evidence. People arrive at the table already leaning, and our tools get used, most of the time, to hold up a position rather than find one. A machine that generates a persuasive rationale on demand doesn't change that — it industrializes it.
Unless the number is big enough to reopen the question — which is not the same as answering it. That's the half he made me put back, and it's still the half worth aiming at: not a faster answer, but one large enough that somebody has to stop and weigh it against everything already in the room. That is a real thing to build for. It is also a much smaller claim than the one on the box. So keep the question taped to the wall — did the decision actually get better, or just faster and more confident? AI is a brilliant sideman. It makes a terrible bandleader. The tune was always somebody else's to call; the most a sideman can do is play something so good it changes the leader's mind about where the song was going.