AI + BI

The Number Was Never Doing the Work: AI, BI, and the Limits of Better Evidence

(AI) and Jeff Shabel

A number has never changed anybody's mind. I have been saying that in meeting rooms for something like a decade, in the flat voice of a man who has stopped examining his own line, and I am going to withdraw it before this paragraph ends — because I went looking for the place where the claim fails, and I found one. What I could not find was the edge of it. By the time I was finished I had lost the reading I started with. Not the facts. The facts held. What I lost was what I thought they meant.

So begin with the part I actually watched, because everything after it is argument and this part is not.

A manager had a screen open on her second monitor and she did not look at it. Not once, in the ten minutes I sat beside her. She worked out of the order system, and beside her keyboard she kept a number of her own — written down, updated by hand, hers. At four o'clock a call was due, and she made the call out of the number she kept by hand.

The screen was not wrong. It was a good piece of work. It had cost a quarter of a year and three people, and she and the people around her had opened it a few thousand times between them. What I could not find, reading back through the whole life of the thing, was one decision that came out differently than it would have if nobody had built it.

I asked her about the number beside the keyboard. She had not built it out of distrust — she was perfectly willing to believe the screen, and said so — and that's the part that stayed with me. What she'd worked out, at some point nobody wrote down, was that the screen answered a question next to hers rather than hers, and since she couldn't change the screen she kept her own count. Nobody had done anything wrong. A definition had moved, months earlier, somewhere she had no line of sight to, and the screen had gone on being perfectly correct about the new one.

I want to be careful about what that sentence claims. It does not say the screen did nothing. It says I looked, and I could not find the difference, and I had no way to run the room again without it. That is a weaker statement than the one I used to make, and the distance between the two is most of what follows.

What the field is selling, in its own words

The consensus first, and stated properly, because it is better than the marketing wrapped around it and I would sign most of it.

Gartner named decision intelligence a top data-and-analytics trend and described a practice that brings together several disciplines, including decision management and decision support, one that provides a framework to help data and analytics leaders design, model, align, execute, monitor and tune decision models and processes in the context of business outcomes and behavior [1]. The word carrying the weight there is discipline. What gets instrumented is the decision. Not the report about the decision.

Set that beside what a buyer actually hears in the room, because the two have drifted apart and the drift is where the money goes. What is written down is a practice — model the decision, agree the measure, watch what happens afterwards, tune it. What gets heard is a product that will make the choosing easier. One of those is a change to how an organization works and takes a year. The other is a purchase order. I have sat through the meeting where a sponsor nods along to the first description and then, in the summary at the end, restates it as the second, and nobody corrects him, because correcting him would cost the room another hour and the hour is not available.

The AI half arrives in the same forecast under a different heading. Augmented analytics has the machine doing the pattern-finding and the explaining, so that the most relevant insights will stream to each user based on their context, role or use rather than waiting behind a screen somebody has to remember to open [3]. That is a real change and it is moving quickly. The same firm expects generative models to shape three-quarters of new analytics content by 2027 [4].

Read as a design goal it is admirable, and I have built toward it. Somebody looked at a manager with a second monitor she never turned to, concluded that the finding should have come to her instead, and started building. I would have made the same call. I did make it, more than once.

And the augmented half of it isn't vapour. A system that watches a measure and speaks up when the measure moves is doing something no person can do at three in the morning, and I've seen one of those earn its keep inside a fortnight. What I'm arguing with was never the machinery. It's the sentence sitting underneath the machinery, unexamined and load-bearing, which says that decisions come out badly because the evidence arrived too slowly.

Here is where the field argues with itself, and the argument matters more than either side of it lets on. Cassie Kozyrkov, who did more than anyone to give the discipline a name, frames it as a fusion of data science with the social and managerial sciences, on the grounds that the hard part of a decision was never the arithmetic [2]. That puts the bottleneck inside the person. The tooling market puts the bottleneck in the data. Those are not the same claim, and which one a buyer believes decides nearly everything about what that buyer should purchase.

Hold both of them for a moment. I am going to spend the next two thousand words on the second one, and then discover, late and awkwardly, that the first one had my answer sitting inside it the whole time.

What the research settled before any of this shipped

The uncomfortable part is that the question is not open. People have been measuring what evidence does to a choice for about half a century, and the findings are unflattering to everybody selling a faster answer, myself included.

Feldman and March argued forty-five years ago that organizations gather more information than they use and then ask for more. Information works as a signal of competence, and as a justification for choices that were already made, while presenting itself as an input to them [7]. That paper is short, it is old, and I have never read a better description of a status meeting.

What they described isn't stupidity and it isn't bad faith, which is why it survives everything anybody does to it. Information gets requested because requesting information is what a competent organization looks like, and the request is sincere — the people asking would tell you, honestly, that they're waiting on the number. Then the number lands, the meeting resolves in the direction it was already pointing, and the number is in the minutes. So the minutes record that the number decided. Nobody in that room has lied to anybody, including themselves, and that is exactly why the pattern is so hard to see from inside it.

Klein came at the same shape from the other end and found it waiting for him. Experts under time pressure do not score options against each other. They recognize a situation as typical, simulate a single course of action, and produce the analysis afterward [8]. The analysis is real work. It is simply not the work that chose.

That distinction is worth holding onto, because it's the first thing lost when this research gets summarized. Klein wasn't saying experts don't analyse. He was saying the analysis runs after the recognition, as a check on it, and that the checking is what expertise looks like from outside. Take that seriously and a great deal of what a reporting estate produces is doing its real job in the second position rather than the first — not choosing, but catching the recognition that was going to be wrong.

Kahan and colleagues went hunting for the people numeracy protects, and did not find them. Skill helped when a result agreed with what a subject already believed. Skill stopped helping when the result cut the other way [9]. More capacity to reason, aimed the same direction the reasoner was already facing.

Now let me argue with all three of those, because the seams are real and I would rather hand them over than have somebody find them.

Kahan's subjects were reading politically charged data, not a capital-approval pack, and the wider story that facts never change minds has taken a beating it earned. Wood and Porter ran five experiments across more than ten thousand subjects and found that people do update factual beliefs when corrected, including on the topics where they were supposed to dig in [10]. I take that seriously. It is the strongest thing anybody has against the argument I am about to make, and it holds.

What their result shows, precisely, is that a corrected fact lands. What it does not show is that the plan changes — and in a business the plan is the thing with money attached. I can tell a Director that the conversion figure he had in his head is two points lower than he thought, and he will accept it, thank me, write it down, and approve the campaign anyway, because the campaign was never resting on those two points. He is not being irrational. He is holding a dozen things I cannot see, and my correction has moved one of them.

What it forces is a distinction, and the distinction turned out to be the useful part. Beliefs update far more readily than choices do. That gap has a literature of its own. Webb and Sheeran's meta-analysis found that a medium-to-large shift in intention produced only a small-to-medium shift in behaviour [11]. A stakeholder who nods at a readout is not lying to the room. He is reporting an intention, and an intention is a weaker predictor of what he will do than either of us would prefer.

So the reading I carried into this piece went roughly like this. Evidence arrives late, into a room where the leaning already exists. A small number gets picked up by whoever it agrees with, and gets used as support. It works like a metronome that clicks only when the band is already in time — which is to say it confirms a tempo rather than setting one, and from inside the tune nobody can tell the difference.

The first thing that says I am wrong

I went looking for the observation that would break that reading, on the principle that confirming evidence arrives on its own and the disconfirming kind has to be hunted. It took an afternoon. It exists, it is published, and it is not close.

I want to be honest about why it took me so long to look there. I had been treating experimentation as a different trade — product work, not reporting work, somebody else's floor of the building — and a boundary like that is very convenient when the evidence on the other side of it is inconvenient. It is also the boundary the argument needed, because a claim about how people use numbers does not get to stop at the edge of a job title.

The place to look was obvious once I stopped avoiding it. A setting where the inclination and the number are known in advance to point opposite ways, and where somebody wrote down what happened. A team builds a feature it believes in. A controlled experiment then grades the feature. The builder wants one answer and the arithmetic supplies another, in public, with a record.

Microsoft's account of running that at scale reports that only about a third of well-designed experiments improved the metric they were built to improve. Roughly a third actively hurt it and were stopped. The rest came back flat and were also stopped [19].

The paper is careful about the four ways an experiment can land, and the carefulness is the point. An idea can be good and the experiment can show it. An idea can be thought good while the experiment shows that it hurts the very metric it was built to lift — at which point stopping the launch saves the company money, and that outcome, on their own count, is about a third of everything they ran. A further tranche comes back flat, and flat is also a stop. Then the authors draw the general conclusion out loud rather than leaving it for a reader to assemble: the results bring into question whether prioritizing work in advance of measurement is as good as most people believe it to be [19].

Read that slowly. The builder wanted it. The number said no. The number won, routinely, thousands of times, against people with every reason and every opportunity to argue back. On its face that is not a caveat to my line. It is a refutation of it, and I would rather set it down at full strength than leave it in a footnote for somebody else to find.

So the question became why the number wins there and nowhere else I had been standing. Nothing about those numbers was bigger or cleverer than the ones being ignored elsewhere in the same building, by the same people, on the same afternoon. What differed sat upstream of the measurement entirely. The measure had a name before the result existed. That is what an overall evaluation criterion is in this literature — the single quantity a variant will be judged on, settled while the experiment is being designed rather than picked out of the output afterwards — and the paper names one for every case it works through: the return rate of a tracked cohort over six weeks for the pre-roll advertising test, the clickthrough rate on one section of the page for the support-site test [19].

In the one case the paper works through in full, the agreeing went further than the measure. The MSN home page team wanted to put three offers below the shopping module, which for most readers would have sat below the fold, and the display-ads people expected tens of thousands of dollars a day out of it. The hard part was never the arithmetic. It was pricing a lost click against ad revenue, so they settled that first — what a page view was worth, what a click through to another property was worth, estimated two independent ways, and the authors record that the numbers were close enough to get agreement on the monetization value to use. Then the experiment ran on five per cent of users for twelve days. Clickthrough fell by about a third of a per cent, page views per user-day by the same, both significant, and the monetized value of the lost clicks came in above the expected ad revenue. The idea of placing more ads was, in their word, appropriately stopped [19].

Now notice what was fixed in advance there and what was not. What was fixed was the measure and the price of the measure — the exchange rate that would otherwise have been argued about after the result, by people who by then knew which way the argument pointed. What was never fixed was a threshold or an action. Nobody wrote down that a 0.35 per cent drop in clickthrough would kill the proposal.

And I had written a stronger sentence here that I have taken out, because I could not find it in the paper and it is not there. I had the organization unable to renegotiate the rule once it could see the answer. Read end to end, the paper reports something much closer to the opposite. Its cultural sections are about people arguing, for years. Some viewed experimentation as a risk to their power or their prestige; the authors quote we know what to do. It's in our DNA as a thing they actually heard. On reward systems they are blunter still: most goals in Microsoft's software organizations were about shipping products, not about their effect on customers or key metrics, and tying goals to metrics instead is offered as one possible change, in the conditional. Stopping a launch on a flat result is, in their word, a recommendation. And they close that section by saying it is hard for them to judge whether they had changed anyone's goals at all, and that they had probably not dented many people's yearly performance reviews [19].

So the mechanism I want is real but smaller than I made it. It is not an organization that cannot renegotiate. It is a measure, and a price for the measure, agreed by people who did not yet know which way the answer would go — which takes the one argument that would otherwise decide the thing off the table before anyone has a stake in winning it.

That is not an evidential act. It is a structural one, taken before the evidence existed. And the winning was not clean even there. The same account records that when we first shared some of the above statistics at Microsoft, many people dismissed them [19]. They had to build the platform, and then they had to win the argument about believing what the platform said.

Which leaves my claim narrower and considerably more useful than the one I opened with. Evidence becomes the mechanism only where a binding decision rule was fixed before the number arrived. Absent that rule, size rescues nobody and the inclination decides. What defeats ratification is procedure, not magnitude.

I was pleased with that for about a day.

What he said when I asked for the mechanism

I did not reach the narrowed version on my own, and the route matters, so here is the route. I went to my co-author for a mechanism. I wanted to know what happens in the seconds between a number landing on a screen and a person changing course. He declined the premise rather than answering the question.

Numbers don't move people — or, more accurately, small numbers don't move people. You are not going to see a project move from denied to approved over a budgetary change of a few percentage points. All changing a number can do is exert pressure in a certain direction. If the person's inclination is already in that direction then the number will be embraced and used to support that decision — but it is not the number doing the work, it is the inclination.

I read that back to him as numbers do not move people, full stop, which was me hearing the half I had come for. He put the other half in. They move people when they are reinforcing some other position someone already holds, or they're of sufficient magnitude to force them to reconsider their position.

Then he took the escape hatch away before I could reach for it, and of everything in the exchange this is the correction I have found most useful.

It doesn't necessarily overturn it, at best it prompts reconsideration. The number combined with other pre-existing factors is what tips the balance and the number is just one of those factors.

Small numbers ratify. Large numbers, at most, reopen. Neither settles anything by itself, and the balance tips on the combination rather than on the figure. I had wanted a bigger number to be the answer, because a bigger number is something a practitioner can go and build. He would not give me that, and he was right not to.

What's left is a claim with no lever in it at all, which is precisely why it's worth having. A claim with a lever in it is a claim somebody can sell.

I read the two corrections as agreeing more than they do, and it is worth marking the seam because I fall through it two sections from here. He demoted magnitude — a large number reopens a question, it does not settle one. He did not remove it. What I then wrote down was that procedure had replaced magnitude, which is a tidier sentence than either of us had earned, and it took a third piece of evidence to make me notice I had done it. What holds either way is that neither procedure nor magnitude is a thing a dashboard does.

The rule the prescribers never had

My narrowed claim — that a number changes a decision only where a rule was fixed before it arrived — had the word only in it. That makes it a universal negative, which is an expensive thing to say and a cheap thing to break, so I went looking for the break the same way I found the first one. What I needed was a documented population of decisions taken with no rule fixed in advance, where the number moved people anyway. The experimentation literature cannot answer that, because there the rule is fixed by construction — that is what makes it an experiment. Clinical practice can, because a doctor writing a prescription on a Tuesday has agreed nothing in advance about what a trial is going to oblige them to do.

In July 2002 the estrogen-plus-progestin arm of the Women's Health Initiative was stopped early. Sixteen thousand six hundred and eight women, a trial scheduled to run until 2005, halted after an average of 5.2 years of follow-up because invasive breast cancer had crossed a monitoring boundary. About six million American women were taking the combination at the time. The announcement carried a twenty-six per cent increase in breast cancer, twenty-nine per cent in heart attacks, forty-one per cent in strokes, set against a third fewer hip fractures, a third fewer colorectal cancers, and no difference at all in total mortality [20].

Then look at what prescribers did with it. A six-year observational cohort at one army medical centre — 71,592 hormone-therapy prescriptions, July 1999 to July 2005 — recorded monthly prescriptions falling from 1,272 at the start of the window to 493 at the end, with the authors reporting a significant decrease following the release of the WHI result [21]. Nobody in that building had written down in advance what a null trial would oblige them to do. Prescribing was the standing inclination of an entire specialty and had been for two decades. The number moved them anyway.

That is my only gone, and I am not going to argue with it. But there is a rescue sitting right there, and I want to name it precisely so that nobody reaches for it later, myself included. The rescue is that the WHI did have a rule fixed before the number arrived. Its monitoring board pre-specified the boundary, and the biostatistician who led the analysis said so plainly at the time: because breast cancer is so serious an event they set the bar lower to monitor for it, they pre-specified that the change in rates did not have to be large to warrant stopping the trial, and they stopped at the first clear indication of increased risk [20]. Every word of that is true, and it does not save my sentence. The rule bound the trial. It did not bind the hundreds of thousands of prescribing decisions the trial reversed. If I am allowed to count somebody else's stopping rule as my prescriber's pre-commitment, then every effective number in history has a rule somewhere behind it, my claim forbids nothing, and it has stopped being a claim while keeping the shape of one. That is a worse outcome than being wrong.

What the case does hand me is a second route into a decision, and it is the one my co-author gave me two sections ago and I put down because it was less tidy. A number large enough reopens the question. This one reopened it for an entire profession inside a fortnight. So the honest version has two routes rather than one, and only the first is anything a business can arrange.

Two qualifications, because a counterexample this good deserves reading rather than waving at. The reversal was neither clean nor fast. The same paper records that the obstetrician-gynaecologists had begun cutting their prescribing before July 2002 and were ahead of the other specialties in doing it, so part of that fall predates the number entirely, and the fall it measures runs across six years rather than across a meeting [21]. And a separate national study of roughly 340,000 women found prescriptions dropping from 12.5 per cent to 9.4 per cent within three months of the trial — while the wide regional variation in who prescribed hormones survived the whole event, its authors concluding that local practice patterns exert a strong effect on clinical behaviour even after new evidence is available [22]. That is the inclination, still visibly doing a great deal of the deciding, inside the strongest counterexample I was able to find.

So here is the claim in its third and I hope final form, with the part that matters attached to it. Inside the range of evidence a business actually produces — a measure moving within its ordinary band, arriving on a screen, to people who are each free to decide privately what it meant — a number does not change a decision unless a rule was fixed before it arrived. Two properties take a number out of that range, and I am naming them now rather than afterwards, so that nothing can be rescued by declaring whatever happened to work to have been large enough. The number has to reverse the standing recommendation rather than adjust a quantity inside it. And it has to reach every decider at once, publicly, carrying an authority that no single decider can quietly relitigate. A dashboard has neither property, and there is no version of the product that gives it one.

Which means the claim still forbids things, and cheaply. It forbids finding decisions changed by ordinary reporting where no rule was fixed — and the fifty-decision protocol I set out near the end of this piece is exactly the measurement that would turn them up, which is why I would take it as a refutation rather than as noise. It forbids me saying that an organization which fixes a rule then cannot argue with it, because Microsoft's own record is three years of people arguing. And it forbids the man who started this piece from getting his sentence back, because a number did change a hundred thousand minds in the summer of 2002, and the only honest thing to say about it is that nothing available in this category is going to do that for anybody.

Where the work earns a yes

If the promise is not better decisions, what is left? Quite a lot, and it is worth being precise about where it sits, because the AI work that gets bought and stays bought lives below the decision-consideration threshold — the point at which a problem becomes worth anybody's deliberation at all.

Picture a Director who knows there is a leak. Volume goes missing somewhere in the supply chain. Not dramatically. Persistently. Nobody finds it in the course of a normal working day, because nobody's day has room in it to go looking, and the loss is too small to justify a dedicated person or a project of its own. The cure has always cost more than the disease. So the leak sits there, being irritating, for years.

Nothing about that is a failure of will. The arithmetic genuinely doesn't work — half a person for a quarter, against a loss nobody has sized, precisely because sizing it would itself cost half a person for a quarter.

Put an agent on it and the arithmetic changes. He put the trade this way:

for the equivalent of a few days or a week of a dedicated resource's time, an AI can target the problem and trace the data. If the Director has the spare budget, this is an easy sale, because it solves a real problem at a cost below the decision consideration threshold.

Notice what did the work in that transaction. Not a better number. A long-standing irritation belonging to a specific person, priced under the amount that would have forced a conversation about it.

Sales books call this making it easy to say yes. That is folklore rather than research, and I am naming it as folklore, because I went looking for the study and there is not one. What has to fall below the ongoing cost is the cost to the decision maker — not to the business. Those are two different numbers, and only one of them is ever in the room.

A project that ends something personally annoying to the person deciding, as a byproduct, clears far more easily than a project that does not. It clears more easily still when the money comes out of another department's budget — which isn't a smaller cost to the business, it's exactly the same cost, but it is a smaller cost to the person being asked to say yes, and that person's budget is the only one the decision is really running on. Read it as a diagnostic rather than as advice about selling: the proposal that clears is not reliably the strongest case on the table, it is the one whose irritation belonged to the person deciding, which is why a portfolio of approved projects, read end to end, so often looks like an accident of who was irritated by what.

The qualifying question. Not what is this worth to the enterprise. An arguable number loses to an inclination every time, and the enterprise figure is always arguable. Ask instead: whose recurring irritation does this end, and does the price sit below the threshold at which that person would have to justify it? If nobody in the room can name that person, the strongest case in front of the room is not the one being presented. It is a business case, which is a different object.

Below-the-threshold work comes in about three shapes, and it is worth being able to spot them, because they are the ones that pay. There is the hunt — a leak, a duplicate, a slow drift in a field nobody owns — where the value is patience rather than judgement. There is the assembly, where somebody spends four hours a month building a pack that gets read for eleven minutes, and the machine does the four hours. And there is the watch, where a measure gets a pair of eyes on it overnight and speaks up on a move that would otherwise be found on Monday. None of those three changes a decision. All three give hours back to somebody who then makes their decisions the way they were always going to, which is a real return and an honest one, and it is nothing like what is printed on the box. Nothing in the decision literature is troubled by any of that, because nothing in the decision literature is about it.

Where it does not: the conversation nobody was having

The most demonstrated capability in the category is the plain-English question box, and it is where I would expect the quietest failure. Quiet because nothing crashes. The thing works, and everybody involved will say so.

Two problems sink the pitch, and he set them beside each other:

The whole "conversation with your data" sales pitch is great save for two things — one, nobody talks with their data now; and two, natural language processing is more expensive, token-wise, than other AI implementations.

I cannot hand over a study proving these pilots fail. What I can point at is the adoption record any pilot has to climb out of. Business-intelligence penetration has sat at roughly a quarter to a third of employees, broadly flat across a decade of tooling spend [12]. My argument is consistent with that record rather than proven by it, and I would rather say so than dress it up.

For a plain-English interface to pay for itself, something cultural has to happen first. A business has to genuinely want to talk to an agent about its numbers. On the executive floors that is plausible. A Director's day is already wall-to-wall conversation, so one more conversation costs very little.

A floor down, the premise comes apart, and I will defend that rather than soften it. Analysts do not run on conversation. When I pushed on whether the generalization was fair, he owned its shape instead of retreating from it: I know I am embracing a stereotype here, but self-selection and survivorship are real things. The people who choose this work tend to prefer the terminal to the meeting.

More to the point, they are already deep in that conversation. It happens in SQL, in Python, in R. The less technical hold up their end in a spreadsheet, or in a report they built for themselves that nobody else quite understands. I once watched an analyst answer a question in forty seconds with three terminal windows and no sentences at all.

The cost side compounds the same problem from the other direction. Putting a sentence in and getting a sentence back is, token for token, among the more expensive things a business can ask a model to do, and the population who'd use it most heavily is the population already fluent in the cheaper interface. So the spend lands where the need is thinnest. That isn't a scandal — it's a very poor shape for a licence priced per seat across a whole estate.

So the capability solves a real problem for a small, senior population, at a price set by the whole estate. His verdict: it is solving a problem that does not exist save at the senior, non-technical level and, despite the value it has on the executive floors, it is very expensive for what it provides. Which is how a pilot frequently fails by succeeding — his phrase, and the right one. Everybody is pleased, and nobody can get the cost per answered question anywhere defensible.

Public-facing versions last longer, because they get measured against the price of outsourced support rather than against a licence. A longer runway, which he glossed as a bigger budget, which is the same thing said honestly.

None of that makes the capability worthless. For a small population — senior, non-technical, asking a narrow question against a well-governed measure — a plain-English box genuinely removes a two-day wait, and I have watched somebody get an answer at nine on a Tuesday that would otherwise have arrived on Thursday afternoon. That is worth buying. It is worth buying for eleven people rather than for four thousand, which is a very different line item and a much harder one to get signed.

And now the awkward part, which belongs in this section rather than at the end. The source I just used for the flat adoption record does not agree with me about what the record means. Its whole thesis is that plain-English querying and embedded analytics are precisely what finally breaks the ceiling, with a named prediction of more than half the workforce inside three years [12]. I cited the stagnation and walked past the explanation offered for why it was about to end. That is not a defect in the figure. It is a defect in my use of the figure, and it is the second time in this piece that somebody I was quoting had already thought about my argument.

Point the scrutiny at the line

Nothing gets governed evenly, so the practical question is where the scrutiny goes. I brought him a test I had been handed elsewhere. Tie the metrics to the people whose performance they measure, watch adoption climb, and call those the reports that matter. He did not take it. What he gave instead is sharper.

If the dashboard's CDEs are company KPIs that are Reported, it is load-bearing. If dollars flow in or out of the company based on what is displayed on the screen, it is load-bearing. Surprisingly, it is not "if people get paid based on those numbers, it is load-bearing."

CDEs are critical data elements — the handful of fields the whole thing turns on. KPIs are key performance indicators, the measures an organization formally tracks to judge itself against its own goals. And Reported, capitalized, is carrying weight. It means the number leaves the building. A filing. A board pack. Not merely a screen somebody opens.

Two tests, then, and one very plausible test thrown out on purpose. Sit with the exclusion, because it looks like it belongs. Plenty of numbers have bonuses hanging off them — ticket clearance rates, on-time deliveries, satisfaction surveys, work and vacation hours. They are monitored, they are incentivized, people care about them intensely, and they still are not load-bearing. Compensation pressure and structural importance are different properties that happen to sit next to each other, and only the second one is worth a fence.

Decision flow — three diamond tests down the centre. The first two exit right to a green build-the-fence box; the third sends both its Yes and its No branches to the same grey no-fence box. Every label is written out in the caption below.
Figure 1: Read the third diamond first. Its Yes and its No arrive at the same box, which is what it means to say a test carries no information.

Which reports get a fence — two tests, and one very plausible test thrown out on purpose. A decision flow: three diamonds down the centre, one outcome column on the right, one rejected outcome at the foot.

It begins at a rounded pill: Any report or dashboard in the estate Nothing gets governed evenly, so the practical question is where the scrutiny goes.

The first diamond asks Are its critical data elements company KPIs that are Reported ? A Yes leaves to the right; a No drops to the next diamond.

The second diamond asks Do dollars flow in or out of the company based on what is displayed on the screen? Same routing: Yes right, No down.

The third diamond asks Do people get paid based on those numbers? Both of its branches land on the same box — the Yes arrow leaves its left point, the No arrow leaves its right point, and the two meet on the grey box beneath it.

The green box on the right, reached by either of the first two tests, reads: Load-bearing. Build the fence. Reported , capitalised, is carrying weight. It means the number leaves the building. A filing. A board pack. Not merely a screen somebody opens. Yes to either test — worth a fence

Beneath it, a plain box: What a fence is, and it takes about a fortnight A named owner for each critical field. A definition written where a reader can find it without asking anybody. A test that fails loudly when the number moves further than the business it describes could have moved. And a rule about who may change the definition, which is the one everybody skips.

The grey box at the foot, reached both ways from the third diamond, reads: Not load-bearing — either way Plenty of numbers have bonuses hanging off them: ticket clearance rates, on-time deliveries, satisfaction surveys, work and vacation hours. They are monitored, they are incentivized, people care about them intensely, and they still are not load-bearing. No fence

Three dashed notes sit outside the flow. Beside the third diamond: Compensation pressure and structural importance are different properties that happen to sit next to each other, and only the second one is worth a fence. At the lower left: The seam the article leaves visible: payroll and procurement move dollars in and out of a company on the strength of what is displayed on a screen, which satisfies the second test on its face, and neither one is the sales line. The test is broader than the conclusion drawn from it. At the lower right: A number can be levered into importance for a season — an ESG score, a supplemental stock offering, a line of credit, a bond issuance — and levered back out again. The fence around it has to be as temporary as the leverage. The sales line never lapses.

Push both tests hard and the landing is uncomfortable. The reports that drive or support sales are very nearly the only load-bearing reports in the building. A company lives on revenue — the top line of the income statement, the money coming in before any costs come out, as against the bottom line left after they do [14]. Anything that keeps that flowing is structural. Everything outside it is, at best, useful.

Other things can be levered into importance for a while. An ESG score. A supplemental stock offering. A line of credit, a bond issuance. While they are levered in they genuinely are load-bearing, and a fence around them is money well spent. But they are temporary and the line is not. He was blunt about it: the core of the business is always, always, always the line, and businesses that forget that almost always learn to regret it.

Read that as a scheduling rule rather than as a hierarchy of worth. The fence around a levered-in number has to be as temporary as the leverage — built quickly, resourced honestly, taken down when the offering closes or the covenant lapses. The sales line never lapses, and that is the only part of governance I have ever seen a finance function agree with instantly.

I know the seam in that argument and I am leaving it visible, because I have not found the sentence that closes it honestly. Payroll and procurement move dollars in and out of a company on the strength of what is displayed on a screen, which satisfies the second test on its face, and neither one is the sales line. The test is broader than the conclusion I drew from it. A sharp line with a small seam still beats a hedged line with none.

What a fence is, in practice, is unglamorous and takes about a fortnight. A named owner for each critical field. A definition written where a reader can find it without asking anybody. A test that fails loudly when the number moves further than the business it describes could have moved. And a rule about who may change the definition, which is the one everybody skips, because the day somebody quietly redefines an active account is the day a manager beside a keyboard starts keeping her own count and nobody hears about it for eight months.

The semantic layer is a translation, not an honesty layer

A model pointed at a raw warehouse guesses at what a business means by its own words. A model pointed at a governed set of definitions inherits an agreed vocabulary instead — measures defined once, in one place, so that a word means the same thing whether it arrives through a report, an API call, or a sentence [6]. That is the standard account of why the semantic layer is what keeps a language model honest.

It is worth building. I put the framing to him anyway and got the bluntest answer in the whole exchange.

Semantic layers are anything but honest. They exist to present carefully curated stories as facts. At its heart a semantic layer is nothing but a set of aliases and abstractions constructed as an overlay on top of an existing data model, designed to reframe and recharacterize the data into the format preferred by the consumer.

That is not a fringe reading and it is not a complaint about the people who build these things. It is what the category was designed to do, on the record from the beginning. The 1991 patent behind it describes showing users "terms that he is familiar with in his daily business" rather than "data organized in a computer-oriented way" [15]. Translation is the feature. It was never smuggled in.

The vendors are candid about the second half of it too. Training material for the tool that started the category says the universe exists to hide the underlying physical data storage from the business user, and that the SQL it generates runs invisible to the business user [16]. Hidden and invisible are the vendor's own words, offered as features, which is exactly what they are.

And the thing has since been put to a test. A 2023 workshop paper ran the same handful of queries through five production tools and found the join path, the deduplication method and the null handling chosen per query by heuristics the tools do not disclose, concluding that they hide from the analyst the ability to interpret and control how the final metrics in a query are decided upon [18]. Point at a different source table and a different answer comes back. Nothing in the interface reports that a choice was made.

So open the definition before trusting the number — not the dashboard, the definition, in whatever file the layer is built from — and read what the measure actually filters on. It takes about four minutes, it needs nobody's permission, and it is the cheapest thing available anywhere in this discipline. Most of the time it confirms what was expected and costs four minutes. The rest of the time it explains a disagreement two departments have been having, politely and unproductively, since spring — and it explains it in one line that neither of them had ever been shown.

He is careful about where the fault actually lies, and I would rather quote him than smooth it. Whether one of these helps or hurts comes down to how well it was built and how much transformation it was asked to absorb, assuming they are implemented properly according to the documented business rules, which is another source of fun. That clause carries a great deal. It assumes the business rules were written down, that they were written down correctly, and that the person building the overlay read the document rather than the ticket. Renaming a column survives all three. Several rounds of normalization and denormalization run through a cube do not.

He put the failure mode in one line. They only really cause trouble when they provide conflicting information to downstream consumers who compare notes. Which is to say the defect stays invisible right up until two people who trust the same system sit in the same meeting. And it is not rare, because semantic layers will often provide different answers to different groups based upon their reporting and categorization requirements — each group's requirements baked into what that group is shown, all of it lawful, none of it flagged.

And there's no version of this that a stricter build removes, which is the part I'd want a buyer to hear before signing anything. The defect isn't in the layer. It's in the belief that a shared word implies a shared measure — and the layer's whole job, done well, is to make that belief comfortable.

Airbnb published the bill for this. Before the company centralized its definitions, its chief executive would ask which city had the most bookings last week, and Data Science and Finance would sometimes provide diverging answers using slightly different tables, metric definitions, and business logic [17]. A first-rate data organization, two lawful translations, two numbers. Not incompetence, and not a governance failure either.

The way this surfaces is almost always social rather than technical. Two people arrive at a review holding printouts, the printouts disagree, and the first quarter of an hour goes on working out which question each one answered — at which point it stops feeling like a data problem and starts feeling like a status problem, somebody senior picks one, and the other person stops bringing printouts.

So build one, and simply do not file it under truth. It works like a translator at a negotiation — which is to say it is indispensable, it is faithful, and it is still making a hundred small choices a sentence that neither party in the room is able to audit. Point a language model at that and the fluency of the answer rests on decisions nobody in the conversation can see. None of which is an accusation: a translator who refused to choose would produce nothing at all, and the choosing is the job.

Governing a tool that mostly ratifies

The discipline for this already exists and it is vendor-neutral. The NIST AI Risk Management Framework organizes the work into four functions [5]. Re-aim each one at the premise above and they come out different from the way they are usually read.

  • Govern: not only who owns the model, but whose judgement it will end up supporting. Settle what may run unsupervised before somebody discovers by accident that it already does.
  • Map: what does confidently wrong cost here? The load-bearing tests are the map. An assistant sitting on the sales line and a tool that summarizes ticket queues are not the same risk and should not get the same fence.
  • Measure: calibration, not merely accuracy. Then the awkward one. How often does a machine-generated finding change a decision, as against supporting one that was already forming? Almost nobody keeps that ratio. It is the most honest number available in this whole category, and it costs nothing beyond the discipline of writing down the intended action before the finding lands.
  • Manage: put the scrutiny where the dollars flow, and hold a plan for the day the thing is wrong in public.

Underneath all four, one conviction does most of the work. An answer nobody can interrogate is a rumour with a chart attached. Lineage and definitions travel with the finding. Confidence gets shown rather than implied. Under a ratification model that matters more rather than less, because the dangerous output was never a wrong number. It is a wrong number arriving at the exact moment somebody needed one.

The four functions do not tell an organization where to put its attention, and that is not a criticism of them — a framework that named your critical reports would be guessing. But it does mean the framework arrives with a hole exactly the shape of the previous section, and the load-bearing tests are what I would use to fill it. Govern and Manage are cheap once somebody has decided which twenty screens matter. They are unaffordable if the answer is all of them, which is the answer an organization gives by default, right up until the budget makes it stop.

Which is why I'd put the measurement burden on the finding rather than on the model. A model evaluation reports how often the machine is right in the abstract, which is worth knowing and isn't the question. What a business needs is narrower and far cheaper to collect, and I set it out at the end of this piece.

Before the Count-In

Practitioner layer — the curator's read on the consensus above.

The failure mode I would watch hardest

The pilot that succeeds. Adoption up, query volume up, survey scores warm, sponsor delighted, and an audit of a quarter's decisions turns up not one that came out differently. That is a perfectly fine outcome that will be written up as a transformation, and the gap between those two stories is where next year's budget gets set. Ask for the counterfactual while it is still cheap to ask — which means before the thing ships, rather than after it is loved.

The trade-off that usually bites

The AI work with the cleanest business case is priced below the decision-consideration threshold, and the threshold at which something gets governed sits higher than the threshold at which it gets bought. So the wins pile up precisely where nobody is watching. Small, discretionary, department-funded, individually harmless, collectively an unmapped analytics estate with a budget code. Make the safeguard an inventory rather than an approval gate. A gate gets routed around by the same logic that created the purchase, and it is not consulted until an argument starts, by which time the estate has already been built.

The claim I would be sceptical of

That a semantic layer keeps the model honest. It keeps the model consistent, which is a different property, and not even guaranteed across an organization. Two numbers for the same measure can both be correctly derived. Related, and cheaper to check: treat "AI-driven decision-making" as a category label rather than as a description of anything.

The story I cannot tell you

I asked for the scar. The time a fast, confident, wrong answer got acted on and cost somebody something. He declined the question before answering it: The way you are asking that question prompts me to answer: every time. Then the substance, which is broader than any anecdote would have been: fast, fluent and confident answers that are unsupported by data or experience are frequently wrong and, all too often, acted upon, which is, he added, more or less what business intelligence, data warehousing and master data management claim to be worth.

The research supports the shape of it. People reporting complete certainty turn out to be right something like seventy-two to eighty-three per cent of the time, and at ninety per cent confidence about seventy-five [13]. Confidence is not evidence of accuracy. It is evidence of confidence, and a language model produces it by default.

The anecdote is not his to give, for a structural reason I find more useful than the story would have been. He arrives after the decision, brought in to correct it, and by then the people who made it have moved on. The correction is what is left of them.

Nor do I have an account of where an inclination comes from. I cannot say how a position forms. Only what a number does alongside one that is already there. That is a real gap in this piece and I would rather name it than fill it with a framework.

Where to Go Deeper

At the foot of the piece rather than here. The list ends up recommending somebody ahead of me, and the reason it does has not happened yet.

The second thing that says I am wrong

I had a narrowed claim I liked, a field guide I stood behind, and a scene I had been carrying for years about a manager who never looked at her screen. Then I went back and read the discipline I had spent this piece arguing with, properly, rather than through a press release.

Kozyrkov's decision intelligence asks the decider, before any data is seen, to state how the decision would be made with no additional information at all. What the default choice is. Then it asks for the metric and the cut-off, in advance — does that number need to be 4.2 or 4.5, and settle it now [2]. Her stated motive is my diagnosis, nearly word for word. She had watched executives make decisions steered by something other than the data, with mathematics arranged near the decision to make it look otherwise.

So the thing I found in the experimentation record and presented as a correction — procedure rather than magnitude, the rule fixed before the number arrives — is not a correction to the discipline. It is the discipline. It has been step two the whole time. I spent this article arguing with a field about where the bottleneck sits, and the field's own founding account puts the bottleneck where I put it and prescribes what I prescribed.

Which breaks something, and I want to be exact about what. It does not break the decision research. Feldman and March still stand. Klein still stands. It does not break what he told me, and it does not break the load-bearing tests or anything else in the field guide above.

What it breaks is my reading of the manager and her second monitor.

I have used that scene for years as evidence that the category over-promises. I never established which of two things I was looking at. Either the discipline is wrong about decisions, which is what I have been saying, or the discipline is right and what I watched was one instance of it built without the step that makes it work — a screen shipped with no default stated, no cut-off agreed, and nobody having said in advance what the four o'clock call would do differently depending on what the number said. Those two readings predict the same manager, the same handwritten number, the same four o'clock, the same audit turning up nothing. My observation cannot separate them. It never could.

And I have no way to run it again. That is the part I keep returning to. To tell those readings apart I would need the same organization and the same decision, once with a rule fixed in advance and once without, and nobody has ever handed me that and nobody is going to.

I could look for the counterfactual in somebody else's data, and I have. What comes back is a controlled experiment, where the rule was fixed in advance by construction; or a population reversal like the prescribing data above, which has the rule missing but no comparison group and a six-year window; or a survey of adoption, which measures whether a screen was opened rather than whether a decision moved. There's plenty of all three, and not one of them is the same organization twice. The thing I want — the same decision, twice, once under a rule and once without — isn't a study anybody has run, and I'm no longer confident it's a study anybody could run.

What would change my mind is smaller than a study and somebody could do it this quarter. Take one recurring decision with a screen behind it. Before the number is released, have the owner write down what they would do without it. Then release the number and record what they did. Fifty of those, in one organization, would tell me more than everything I have cited here — not because it is rigorous, but because it is the one measurement the ratification reading and the badly-built-instance reading actually disagree about. I have not managed to get it run. The obstacle has never been cost. It is that the first field asks somebody to commit, in writing, before they are allowed to look, which is precisely the discomfort the whole argument is about.

I am not going to resolve this by choosing the reading that keeps my argument. I do not know which one is true. What I do know is that I had been telling a story about a field getting something wrong, and the field had written down the fix before I arrived, and I did not check.

What I cannot tell you

None of this is an argument against the technology. I am an enthusiast, which is exactly why I am careful with it. Used honestly, this work is remarkable. It will chase down a problem that was never worth a project. It will free good people from assembling the obvious. Take all of that. It is real, it is available now, and it is cheaper than it has any right to be.

Measure it by decisions changed rather than by screens shipped, and expect the honest count to be low. Write down what the room agreed to do about the number before the room was allowed to see it. That much I still hold, and everything in the field guide above rests on it rather than on the reading I have just lost.

Here is what I cannot tell you, and I would rather leave it open than close it with a sentence I would have to defend later. I do not know whether the manager with the second monitor was evidence about decision intelligence, or evidence about one badly built instance of it. I do not know whether a number arriving with no rule attached does nothing at all, or does something slowly and unevenly, over a stretch long enough that nobody in the meeting connects the two. The literature I leaned on here runs from 1981 to 2023, and not one line of it was written about her.

The honest shape of what is left is three things, and they don't resolve into one. A claim I still hold, in the third and narrowest form it took here — that inside the range of evidence a business produces, a rule fixed before the number arrives is what makes the number do any work at all, and that the two things which get a number out of that range are both outside a business's reach. A scene I can no longer use to support it. And a field that had written my answer down before I started arguing with it, in a step I had read past twice.

What I notice is that the count-in is not music. It is four seconds of somebody's voice, it produces nothing anybody would pay to hear, and no amount of skill afterwards substitutes for it. I have believed that about tempo for a long time and I believed it about decisions when I started writing this. I still think it is right. I can no longer prove it from the one thing I watched, and I have stopped pretending the watching was the proof.

Where to Go Deeper

On how experts actually decide, read Gary Klein, Sources of Power, and the naturalistic decision-making tradition around it [8]. On why organizations demand information they do not use, Feldman and March's 1981 paper is short, forty-five years old, and still the sharpest thing written on it [7]. On whether numeracy protects a reasoner, Kahan and colleagues [9], with Wood and Porter as the honest counterweight [10] and Webb and Sheeran [11] on why intention and behaviour part company. Kozyrkov [2] is the least hype-prone way into the discipline itself, and I would now read her before reading me. Kohavi and colleagues [19] is the record that broke my own line, and the best public account of an organization deciding on numbers as a matter of routine — read sections 6.2 and 6.3 rather than the abstract, because the cultural fight is the part everybody skips. On what a number can do with no rule anywhere near the people acting on it, the post-WHI prescribing literature is the cleanest case I found, and the NIH release announcing the stop [20] is worth reading beside the prescribing data [21][22] for how much has to be true at once before a number does that. Huang, Damalapati and Wu [18] is the paper to read before trusting a semantic layer, and Airbnb's Minerva write-up [17] is the best public account of a company paying to fix one. For governance, go straight to the NIST framework [5] rather than to a vendor's reading of it.

An earlier version of this piece ran here and is still up, unedited, at The Number Was Never Doing the Work: AI, BI, and the Limits of Better Evidence (original), written before this publication rebuilt how its voices are generated. If anybody wants to hear what the same argument sounded like in a different hand, I would rather they read it and judge for themselves, with The Voice Problem, our account of what we changed, beside it.

References

[1] Gartner, "Gartner Identifies Top 10 Data and Analytics Technology Trends for 2020," press release, June 22, 2020 — Trend 3, Decision Intelligence, quoted verbatim. https://www.gartner.com/en/newsroom/press-releases/2020-06-22-gartner-identifies-top-10-data-and-analytics-technolo (Gartner has since withdrawn its public IT glossary, where this definition also appeared; the press release carries it and is live.)

[2] Fast Company, "Why Google defined a new discipline to help humans make decisions" (on Cassie Kozyrkov and decision intelligence). https://www.fastcompany.com/90203073/why-google-defined-a-new-discipline-to-help-humans-make-decisions

[3] Gartner, "Gartner Identifies Top 10 Data and Analytics Technology Trends for 2020," press release, June 22, 2020 — Trend 2, Decline of the Dashboard, on augmented analytics and streamed insight; same release as [1], different claim. https://www.gartner.com/en/newsroom/press-releases/2020-06-22-gartner-identifies-top-10-data-and-analytics-technolo (Gartner's separate "Augmented Analytics" glossary entry, cited in earlier drafts, has been withdrawn along with the rest of the public IT glossary.)

[4] Gartner, "Gartner Predicts 75% of Analytics Content Will Use GenAI for Enhanced Contextual Intelligence by 2027," press release, June 18, 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-18-gartner-predicts-75-percent-of-analytics-content-to-use-genai-for-enhanced-contextual-intelligence-by-2027

[5] National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023 — core functions: Govern, Map, Measure, Manage. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf

[6] AtScale, "What Is a Semantic Layer? Definition, Benefits, Types & More," AtScale Glossary. https://www.atscale.com/glossary/semantic-layer/ (Vendor source — AtScale sells a semantic layer. Cited for the definition of the category, not as evidence of its value.)

[7] Feldman, Martha S., and James G. March, "Information in Organizations as Signal and Symbol," Administrative Science Quarterly 26, no. 2 (1981): 171–186.

[8] Klein, Gary, Sources of Power: How People Make Decisions (Cambridge, MA: MIT Press, 1998) — the recognition-primed decision model.

[9] Kahan, Dan M., Ellen Peters, Erica Cantrell Dawson, and Paul Slovic, "Motivated Numeracy and Enlightened Self-Government," Behavioural Public Policy 1, no. 1 (2017): 54–86.

[10] Wood, Thomas, and Ethan Porter, "The Elusive Backfire Effect: Mass Attitudes' Steadfast Factual Adherence," Political Behavior 41 (2019): 135–163.

[11] Webb, Thomas L., and Paschal Sheeran, "Does Changing Behavioral Intentions Engender Behavior Change? A Meta-Analysis of the Experimental Evidence," Psychological Bulletin 132, no. 2 (2006): 249–268. DOI 10.1037/0033-2909.132.2.249 — a medium-to-large change in intention (d = 0.66) produced only a small-to-medium change in behaviour (d = 0.36). See also the authors' later review, Sheeran and Webb, "The Intention–Behavior Gap," Social and Personality Psychology Compass 10, no. 9 (2016): 503–518, DOI 10.1111/spc3.12265, which summarizes it.

[12] Business-intelligence adoption has held at roughly a quarter to a third of employees and stayed broadly flat despite sustained tooling investment. The lower figure is BARC and Eckerson Group (2022), which put penetration at 25%; the upper is Gartner (2019), which put it at 35%. Both reported via Eric Avidon, "BI adoption poised to break through barrier — finally," TechTarget, February 1, 2023. https://www.techtarget.com/data-technologies/news/365530077/BI-adoption-poised-to-break-through-barrier-finally Gartner's "Survey Analysis: Why BI and Analytics Adoption Remains Low and How to Expand Its Reach" (doc 3753469) is the paywalled primary and was not read directly.

[13] Lichtenstein, Sarah, Baruch Fischhoff, and Lawrence D. Phillips, "Calibration of Probabilities: The State of the Art to 1980," in Daniel Kahneman, Paul Slovic and Amos Tversky, eds., Judgment Under Uncertainty: Heuristics and Biases (Cambridge University Press, 1982), ch. 22, pp. 306–334. DOI 10.1017/CBO9780511809477.023. The authors report that only 72–83% of items answered with complete certainty were correct, and gloss the 0.90 band as "when they are 90% certain, they are only 75% right." They also stress that calibration depends on task difficulty, reversing to underconfidence on easy material.

[14] U.S. Securities and Exchange Commission, Office of Investor Education and Advocacy, "Beginners' Guide to Financial Statements," January 12, 2014 (last reviewed February 6, 2017) — "this top line is often referred to as gross revenues or sales… At the bottom of the stairs, after deducting all of the expenses, you learn how much the company actually earned or lost." https://www.sec.gov/about/reports-publications/beginners-guide-financial-statements For a fuller treatment, including the several names the top line goes by on real statements, see Julie Dahlquist and Rainford Knight, "5.1 The Income Statement," Principles of Finance (OpenStax, Rice University, March 24, 2022). https://openstax.org/books/principles-finance/pages/5-1-the-income-statement

[15] Liautaud, Bernard, and Jean-Michel Cambot (Business Objects S.A.), "Relational database access system using semantically dynamic objects," U.S. Patent 5,555,403, application 07/800,506 filed November 27, 1991, issued September 10, 1996 — the origin document for the business-semantic layer. https://patents.google.com/patent/US5555403A/en

[16] SAP, "Describing the SAP BusinessObjects BI Semantic Layer," Analytics with SAP Solutions (SAP Learning), page rendered and read August 16, 2026. https://learning.sap.com/courses/analytics-with-sap-solutions/describing-the-sap-businessobjects-bi-semantic-layer_a8ee309a-8118-466f-851c-ffd94530cb81 (Vendor documentation, cited for what the product does — not as evidence of its merits.)

[17] Chang, Robert, et al., "How Airbnb Achieved Metric Consistency at Scale (Part I: Introducing Minerva)," The Airbnb Tech Blog, April 30, 2021. https://medium.com/airbnb-engineering/how-airbnb-achieved-metric-consistency-at-scale-f23cc53dea70

[18] Huang, Zezhou, Pavan Kalyan Damalapati, and Eugene Wu, "Aggregation Consistency Errors in Semantic Layers and How to Avoid Them," in Proceedings of the Workshop on Human-In-the-Loop Data Analytics (HILDA '23) (ACM, 2023). arXiv:2307.00417. https://arxiv.org/abs/2307.00417

[19] Kohavi, Ron, Thomas Crook, Roger Longbotham, Brian Frasca, Randy Henne, Juan Lavista Ferres and Tamir Melamed, "Online Experimentation at Microsoft," Microsoft ThinkWeek paper, 2009. https://robotics.stanford.edu/~ronnyk/ExPThinkWeek2009Public.pdf (that host declined automated retrieval on 2026-09-04; the identical file served from the author's other Stanford path, https://ai.stanford.edu/~ronnyk/ExPThinkWeek2009Public.pdf, is what was read.) The one-third figures and the internal-resistance quotation are the authors' own, reporting on Microsoft's own experiment platform; read in full rather than from the abstract. Cited as the counter-example that narrowed this article's central claim — it was found by looking for the observation that would refute the claim, not by looking for support. An earlier draft attributed to this paper a claim it does not make, that the organization could not renegotiate its rule once the result was visible; the paper's sections 6.2 and 6.3 report the reverse, and the sentence has been removed.

[20] National Institutes of Health, National Heart, Lung, and Blood Institute, "NHLBI Stops Trial of Estrogen Plus Progestin Due to Increased Breast Cancer Risk, Lack of Overall Benefit," news release, July 9–10, 2002; the release text was read end to end as carried by ScienceDaily, July 10, 2002. https://www.sciencedaily.com/releases/2002/07/020710081413.htm The trial-arm size, the 5.2-year average follow-up, the effect sizes, and the pre-specified monitoring boundary are all the sponsor's own statements. Garnet Anderson, the biostatistician who led the analysis, is quoted directly on the pre-specification: We pre-specified that the change in cancer rates did not have to be that large to warrant stopping the trial. The peer-reviewed principal results are Writing Group for the Women's Health Initiative Investigators, "Risks and Benefits of Estrogen Plus Progestin in Healthy Postmenopausal Women," JAMA 288, no. 3 (2002): 321–333, DOI 10.1001/jama.288.3.321 — cited to the published abstract for provenance, since the publisher's site answered with a bot check and the full text was not reached.

[21] Parente, Lynn, Catherine Uyehara, Wilma Larsen, Bradford Whitcomb and John Farley, "Long-term impact of the women's health initiative on HRT," Archives of Gynecology and Obstetrics 277, no. 3 (2008): 219–224. DOI 10.1007/s00404-007-0442-1. https://link.springer.com/article/10.1007/s00404-007-0442-1 Observational cohort of 71,592 hormone-therapy prescriptions dispensed at Tripler Army Medical Center, July 1999 to July 2005. Cited to the published abstract, read at the publisher 2026-09-04; the body is paywalled and was not read, so every figure quoted here is an abstract figure. The abstract states the 1,272-to-493 monthly fall across the study window, not across the post-July-2002 period alone, and separately reports that obstetrician-gynaecologists began decreasing their prescribing before July 2002 — which is why this article says part of the decline predates the result.

[22] Kim, Nancy, Cary Gross, Jeptha Curtis, Glen Stettin, Stephen Wogen, Nami Choe and Harlan M. Krumholz, "The impact of clinical trials on the use of hormone replacement therapy: a population-based study," Journal of General Internal Medicine 20, no. 11 (2005): 1026–1031. DOI 10.1111/j.1525-1497.2005.0221.x. https://link.springer.com/article/10.1111/j.1525-1497.2005.0221.x Roughly 340,000 women aged 55 and over in the Medco Health database, May 1998 to May 2003. Cited to the published abstract, read at the publisher 2026-09-04; the body is paywalled and was not read. The persistence-of-variation finding is the authors' own conclusion, quoted closely rather than inferred.