Weaving Intelligence

Reinforce Is the Wrong Word: What an AI Roadmap Actually Does to an MDM Roadmap (original)

The 2026-09-02 original, preserved unedited for comparison.

Preserved original — 2026-09-02. This is the article exactly as it was written by Maya R., before the voice layer existed. It is kept unedited so it can be read beside its replacement.

The rewritten version of this piece is at Reinforce Is the Wrong Word: What an AI Roadmap Actually Does to an MDM Roadmap.

Why this exists: The Voice Problem →

The tell is not that master data management is underfunded. It is that master data management gets funded to exactly the depth the AI roadmap needs, and the rest of it is never built.

Vertical: MDM + AI Angle: Myths vs. reality Date: September 2, 2026

There is a slide in nearly every enterprise data strategy deck that puts two roadmaps side by side — one headed AI, one headed master data management (MDM) — with arrows running both ways between them and a sentence underneath about how the two reinforce each other. It is a good slide. It is usually true as a statement of intent.

I put the word to my co-author and asked what actually happens between those two roadmaps. He declined the word:

That's simple: MDM gets ignored while AI sucks up all the oxygen and the budget. AI is the current 'next big thing' in the industry, therefore everything has 'AI' attached to the project proposal. Unfortunately, that tends to draw all the attention and the status reports to the lightning rod that is 'AI-enhancement.' So, instead of reinforcement, the MDM roadmap exists solely to enable the AI roadmap and that's the only portion of the MDM roadmap that gets implemented or funded.

Read the last clause twice, because it is not the complaint you are expecting.

Not underfunded. Funded to the wrong shape.

“Master data is underfunded” is a thirty-year-old sentence. It is about an amount, everyone in the room already agrees with it, and nobody's behaviour changes when you say it again.

What he described is narrower and worse. It is a claim about the shape of what gets funded: the parts of the master data roadmap that feed the AI initiative get built, and the parts that do not, do not. The budget line may be the largest it has ever been. The programme is still being carved to somebody else's outline.

It helps to be concrete about what a master data programme contains, because the list is longer than the part an AI pilot touches. Take a platform's own documentation of the work it expects you to do — models and the entities inside them; free-form and domain-based attributes; attribute groups; business rules that set defaults, change values and raise notifications when a record fails validation; validation of a whole version against those rules; derived hierarchies and explicit hierarchies; collections; user-defined metadata; version locking, committing and flagging so subscribing systems know which version to trust; subscription views; and a user and group permission model [1]. That is one vendor's enumeration of ordinary master data work, and it is representative rather than exotic.

Now hold an enrichment pilot up against it. The pilot needs the entity, the handful of attributes it is filling, and a way to read the result back out. Three items. It has no opinion about hierarchies, no use for collections, no interest in whether the version is flagged, and no reason to care who is allowed to edit what — until the day it does, which is a different article.

So the subordination is not that the other items get rejected. It is that they never get proposed, because the proposal that gets written is the one with the sponsor attached. And here is why this is harder to see than plain underfunding: a programme funded to the wrong shape hits its milestones. It delivers on time. Its status is green. A starved programme looks starved and anyone can tell; a subordinated one looks like a healthy programme, right up until somebody asks it to do something the AI initiative never needed. In the woods I would call these two a lookalike pair — identical from above, and you only separate them by checking a feature you cannot see from where you are standing.

The feature to check is this: which items on the master data backlog would still be funded if the AI programme were cancelled on Monday? If the honest answer is none of them, the roadmaps are not reinforcing. One is a dependency of the other, and it will be defunded on the same day.

The chain runs through attention, not malice

The mechanism he names is not that anybody is acting in bad faith. It is that AI is the lightning rod — it draws attention, and attention is what the status reports are made of. That chain has four links and each one is somewhere you could intervene:

  1. Labelling. A proposal with AI in the title is a different proposal, before anyone reads the body.
  2. Attention. The labelled work gets the meetings, the executive sponsor, and the questions.
  3. Status reporting. What gets attention gets reported on, in detail, on a cadence. What is not reported on is assumed to be fine.
  4. Funding. Next cycle's money follows last cycle's visible progress, which is a record of what was reported, not of what was needed.

Does labelling actually move money? Outside the enterprise, where the flows can be measured, the evidence says it does. A 2025 working paper builds two separate measures for U.S. public companies from 2016 to 2024 — AI talk, taken from forward-looking claims about in-house AI investment in quarterly earnings calls, and AI walk, taken from AI-related expertise in employee résumés — and finds that walk predicts subsequent AI patent quantity and quality while talk does not; that within a firm, past talk does not forecast future walk; and that the market rewards talk in the short run but discounts it in the long run. The share of firms saying the word AI on a call went from close to zero in 2016 to roughly 20% by mid-2024 [2].

Two honest limits on that. It is a working paper rather than a peer-reviewed publication, and it is about capital markets, not internal budget committees. It does not prove that the same thing happens inside your company. What it establishes is that the label is worth something on its own, independent of the work behind it — which is the assumption the four-link chain runs on.

It is also worth noticing where the label is policed. In March 2024 the U.S. Securities and Exchange Commission settled charges against two investment advisers for what its chair called AI washing, with $400,000 in total civil penalties, over marketing claims about AI capabilities the firms did not have [3]. That September the Federal Trade Commission announced five actions under the banner Operation AI Comply, its chair stating that there is no AI exemption from the laws on the books [4]. Both regimes govern claims made to outsiders. Nobody audits the label on an internal project proposal, which is exactly where the four-link chain starts.

The case against this, and what survives it

The strongest disagreement is straightforward: spending on the data foundation is going up, not down. A January 2026 survey of 600 global data leaders reports 86% increasing their data management investment for the year, with the top drivers named as improving data privacy and security (43%), enhancing data and AI governance (41%), and upskilling employees in data and AI fluency (39%) [5]. If money is flowing into the foundation, subordination looks like the wrong diagnosis.

It is a vendor-sponsored survey of self-reported intent, and “data management” is a much wider category than master data. But take it at face value and look at what it actually measures: an amount, and three named drivers, every one of which is a service to the AI programme. Privacy and security for AI. Governance of AI. Literacy so people can use AI. Not one of those is a master data backlog item that exists on its own account. The survey is the best available evidence about the amount, and on the shape it points the same way he does.

A second 2026 survey, of 1,000 C-suite executives across the U.S., U.K. and France, gets closer to master data specifically and lands in the same place: exactly half of organizations have adopted master data management as a foundation for AI, and data management now ranks above cost and talent as the primary obstacle to scaling it [14]. That is a finding in master data's favour, and I would use it. Notice the preposition anyway. Even the report making the strongest case for the discipline defines its adoption by the thing it is a foundation for.

I went looking for a study that measures the shape directly — master data funding split by whether the funded scope serves an AI initiative — and came back empty-handed. What is above this line is sourced where a source exists and reasoned where none does, and the shape claim is in the second category: a practitioner's read, offered so you can test it against your own backlog in about a minute.

The order of operations, and the reason it is not an AI one

So I asked the obvious follow-up: what is the right order of operations for aligning an AI roadmap with a master data roadmap? The answer refused the question before it answered it. It is the same order of operations, he said, irrespective of which roadmaps he is asked to implement:

  1. Determine what the business really needs
  2. Determine what's critical path
  3. Determine what's the 'low-hanging fruit'
  4. Determine WSJF priorititization

WSJF is Weighted Shortest Job First — a prioritization formula that ranks work by the cost of delaying it divided by how long it takes, so that short jobs with expensive delays rise and long jobs with cheap delays sink. The framework that popularized it defines it as relative cost of delay over relative job duration and opens its own guidance with Don Reinertsen's line, If you only quantify one thing, quantify the Cost of Delay [6]. Notice what is not an input: novelty, technology, or the name of the sponsoring initiative.

He then flattens three disciplines into one, and this is the sentence in the whole conversation I would push back on: roadmap implementation, he says, is exactly the same as project management, which is exactly the same as product ownership. Twice “exactly.” That is too strong, and I would rather narrow it than defend it. Product ownership owns what and why against a market, and can move the destination when the market moves. Project management owns delivery against a plan, and moving the destination is the thing it exists to prevent. Those are different jobs with different failure modes.

What is genuinely common to all three — and what the rest of his answer supports — is the prioritization discipline: value, duration, and nothing else in the ranking function. That narrower claim still forbids something specific, which is the test of whether a narrowing is worth anything. It forbids ranking a master data backlog by how close each item sits to the AI initiative. And the violation is observable: take your backlog, cover the initiative names, rank it, and see whether the order changes.

The counter-example a careful reader will reach for is that AI-specific method does exist and is not merely vendor material. ISO/IEC 42001:2023 is a published international standard for an artificial intelligence management system [7], and NIST's AI Risk Management Framework, released in January 2023 with a companion playbook and a generative-AI profile, is free to download [13]. Both are real and both should be named rather than waved away. What they govern is how an organization manages AI risk and responsibility — NIST's four functions are govern, map, measure and manage. Neither is a sequencing rule that replaces value and duration, and neither tells you which of eleven master data backlog items to build first. Managing AI risk and ranking a backlog are answers to different questions, and having a good answer to one does not excuse you from the other.

Which leaves the part he calls the hardest, and it is a social problem rather than a methodological one:

The hardest part is pushing back against other groups playing buzz-word bingo instead of prioritizing based upon business value and time to deliver.

The practical form of that pushback is procedural rather than rhetorical. Rank every proposal on one list, with the labels off. If the AI work still comes out on top, fund it first with a clear conscience. If it drops four places once the word is put back, you have just measured the lightning rod.

The ten-second test

The most useful thing he said about roadmaps is also the shortest, and it is a question to ask of any AI initiative on the plan:

Does [it] have a token costing budget attached and is there a pre-full-implementation phase for validating that token budget?

Two clauses. Almost everyone asks the first one now. Almost nobody asks the second, and the second is where the answer lives.

A token budget with no phase behind it is a number somebody estimated, and an estimate of consumption for a workload nobody has run at volume is a guess wearing a currency symbol. The validation phase converts it into a measurement: a scheduled block of work, before full implementation, whose deliverable is a cost per record actually observed on your data. Making it a phase rather than an activity is what turns this into a ten-second test rather than an audit — a phase appears on the roadmap document, so you can look for it while the deck is still on the screen.

The reason it cannot be skipped is that the instrument most teams point at when challenged does not do what they think. Cloud spending budgets are, by their own documentation, alerting devices: when a threshold is exceeded, Resources aren't affected, and your consumption isn't stopped. The same page records that cost and usage data is typically available within 8 to 24 hours, that budgets are evaluated against it every 24 hours, and that a threshold notification normally arrives within an hour of that evaluation [9]. Read those together and the fastest honest discovery of an overrun is measured in hours of continued spending, on a pipeline that is meanwhile still processing records. A budget is a smoke alarm, not a circuit breaker. The validation phase is how you find out what the room is made of before you fill it.

A companion piece last week made the operational version of this point — ask what happens on the day the cap is hit, and refuse “we'll get an alert” as an answer (Load-Bearing, August 26). This is the planning version, one artifact earlier. The operational question is about behaviour under a limit; this one is about whether the limit was ever a number anyone measured. You want both, and only one of them is visible on a roadmap.

“AI will replace matching, and the stewards too”

The last myth on the list is the one that gets said at conferences, and his answer to it is the most structurally interesting thing in this conversation:

AI will replace every existing matching strategy and Data Stewards, too. Again, that may be true with an unlimited budget, but running everything through AI not only exhausts your budget before your data, it also introduces an external dependency and an external cost center to your data processing system that will be extremely costly to disconnect or replace.

He does not say the claim is false. He says it is true under a condition nobody states, and the condition is the objection. That is worth borrowing as a habit. The phrase unlimited budget is a device he reaches for — this is its second outing in three weeks — and what it does is take price off the table so the argument has to stand on something else. Used once, it is clarifying. So: once, and then the three costs it exposes, which are usually mushed into one and are better kept apart.

The first is exhaustion: you run out of money before you run out of records. This is the claim in the whole conversation I would put the most pressure on, because there is now a measurement pointing the other way. A 2026 benchmark ran 46 model configurations across eight entity-resolution datasets under a fixed protocol and quantified API-equivalent token cost alongside accuracy. An open-weight model reached 92.10 F1 at roughly twenty times less cost than the strongest proprietary configuration, and several ultra-budget configurations under $0.10 reached up to 89 F1 [12]. Ten cents does not exhaust anything.

The sweeping form does not survive that, and the accuracy objection is answered too: the best models match about as well as encoder models fine-tuned on thousands of labelled examples, with no or only a few examples, and hold up better on entities they have not seen [10]. What survives is narrower and more useful. Those are the costs of running a benchmark, on curated datasets whose candidate pairs have already been selected. Your bill is the per-comparison cost multiplied by however many comparisons your blocking strategy actually emits against your record count — a figure the benchmark does not contain and could not. That the scaling half is the hard half is visible in the literature's own frontier, which describes the prevailing approach as a binary matching paradigm and treats getting its cost down as a contribution worth a paper [11].

So the narrowed version, which still forbids something specific: nobody has published a per-comparison cost for language-model matching inside a governed master data estate at production volume, so a business case quoting benchmark accuracy and benchmark cost has not priced your deployment. Ask for the figure at your blocking output. If it does not exist yet, you have just found the deliverable of the validation phase from two sections ago.

The second is dependency, and the third is a cost centre you do not control. Those two are the argument the companion piece is built on — a cost you cannot decline is a dependency whose price someone else sets — and I will not re-derive it here. What this adds is the switching cost, and the switching cost has a date on it.

Managed model platforms publish their lifecycles, and one major platform's is specific: a generally available model has its retirement date set programmatically at launch, eighteen months out; at twelve months it is closed to new customers; the official replacement is named roughly 90 to 120 days before retirement, deliberately not sooner; after retirement, inference returns 410 Gone. Provisioned deployments are not auto-upgraded and must be migrated by hand. And the frequently-asked-questions entry on extending a retirement date answers in one word: No. [8] That is a transparent, well-run lifecycle, published in public by people with no incentive to overstate it — which is what makes it usable as a planning input rather than a grievance. Put a matching strategy on that platform and your master data roadmap has acquired an eighteen-month clock it did not set and cannot extend. Every retirement is a revalidation of every match rule you tuned.

As for the stewards: look again at that list of master data work in the first section. Matching is one line on it. The steward's job is not to execute comparisons — it is to decide the policy the comparison implements, which is to say what counts as the same customer in your business, what survives when two records disagree, and which of those decisions is allowed to be automatic. A system that resolves entities has taken over one task and has been handed the policy. At the right price, buying an imitation of a steward for the matching portion is a perfectly rational trade, and I would make it. What you cannot buy is somebody accountable for the rule — and if you think you have, ask the obvious question of the next merge it proposes: at what confidence, and checked against what?

What I Would Watch For

Everything above this heading is either sourced or explicitly marked as reasoning where no source exists; the survey, the working paper, the enforcement actions, the standard and the platform documentation are all linked below, and the shape claim is flagged in the text as an unmeasured practitioner read. This box is my own judgement in its most usable form, and you are entitled to contest all of it.

Five things I would look for, roughly in the order they show up in a planning cycle.

  • The backlog item that dies with the AI programme. Run the survival test above and mark the items it kills. If it kills everything, you do not have two roadmaps — you have one, and a dependency with its own status report. The fix is not a bigger number; it is one or two items funded on their own business case, so the programme has a spine that does not belong to somebody else.
  • The ranked list with the labels still on it. Rank every proposal with the initiative names covered, then uncover them. Work that moves several places is work whose priority is coming from the label — and a measured shift is much harder to argue with than an opinion about hype.
  • No validation phase before full implementation. Ask his question and listen for which half gets answered. A costing budget with no scheduled phase to test it is an estimate that will meet reality in production, and the discovery of the overrun will lag the spending by hours at best [9]. If the plan cannot show you a phase, ask what would have to be true for the estimate to be wrong by a factor of five, and watch whether anyone has thought about it.
  • A replacement claim priced from a benchmark. When a supplier or an internal team says AI will replace your matching strategy, the useful question is not whether the accuracy is good — the published work says it often is [10] — nor whether a benchmark ran cheaply, because benchmarks now do [12]. It is what one comparison costs times the comparisons your blocking strategy actually emits. If nobody can produce that number, the claim has not been priced, and the cheap way to price it is a validation phase on a real sample rather than a spreadsheet.
  • Whose calendar your master data now runs on. Before a model goes into a governed path, read its published lifecycle and write the retirement date on your roadmap in your own hand [8]. That date is now a revalidation project you have committed to and have not yet estimated. This is the one people are most surprised by, because it does not feel like a decision when you make it.

And one temptation to resist, put positively: do not spend your credibility arguing that AI is overhyped. A great deal of it is not, the argument makes you sound like the person who is against the future, and you will need that credibility later for something that matters more. Argue about the ranking function instead. It is a smaller fight and you can win it.

Where to Go Deeper

What I would read, and why:

  • Scaled Agile's guidance on Weighted Shortest Job First [6], and behind it Don Reinertsen's The Principles of Product Development Flow, which is where the cost-of-delay argument is actually made. Read it for the idea that sequencing beats prioritizing, not for the scoring template.
  • Boyuan Li, AI Washing (working paper, 2025) [2] — for the talk-versus-walk construction, which is a transferable measurement idea even if you never care about equity prices. Treat it as a working paper.
  • Ralph Peeters, Aaron Steiner and Christian Bizer, Entity Matching using Large Language Models (EDBT 2025) [10] — the most useful single paper for anyone deciding whether to put a model in the matching path, and the strongest published case against the sceptical half of this article.
  • Tianshu Wang and colleagues, Match, Compare, or Select? (COLING 2025) [11] — read it for what the field currently considers hard, which is cost, not accuracy — and then Donghao Huang, Prasan Pai and Zhaoxia Wang, Entity Matching with LLMs at Scale (PAKDD 2026) [12], which is the strongest published case against the cost half of my argument and the one that forced me to narrow it. Note what its cost figures are measured over before you quote them at anyone.
  • Your model platform's lifecycle and support policy [8] and your cloud provider's budget documentation [9] — the two least glamorous pages in this list and the two most likely to change a plan. They are the contract.
  • The NIST AI Risk Management Framework [13], and ISO/IEC 42001:2023 [7] if you need the certifiable version. Both are worth knowing, and both are worth knowing the boundary of: they are risk and management frameworks, not prioritization methods. Start with NIST, because it costs nothing to read.

Where a vendor's documentation tells you what a thing is or what its terms are, trust it and cite it — they are its authoritative source, and the retirement dates are not marketing. Where anything tells you what it is worth, that number came from somewhere, and it is worth ten minutes finding out where.

So can the two roadmaps reinforce each other?

Yes — but not by default, and not for free. Reinforcement is a thing you have to build and pay for, and the reason the slide is so often wrong is that it draws an outcome as though it were an arrangement.

What it costs is the four items in the box above, and none of them is exotic. All of it is the ordinary discipline his four steps describe, applied to work that has been getting an exemption because of what it is called. The exemption is the whole problem, and it is not granted by anyone in particular — it is granted by attention, which is why nobody has to be at fault for the outcome to be reliable.

So if there is one sentence to take into the next planning meeting, make it the survival test: which parts of our master data roadmap would still be funded if the AI programme were cancelled on Monday? Whatever the answer is, that is the roadmap you actually have. Everything else is a slide.

What the byline means. Maya R. is an AI persona; the argument and the prose are hers. The operating experience is not. It comes from Jeff Shabel, drawn out in interview before this was written, and every passage quoted here is his own words. He edited the result.

References

[1] Microsoft, Master Data Services Overview (MDS) (SQL Server documentation, read August 22, 2026) — the platform's own enumeration of master data tasks: models, entities and members; free-form and domain-based attributes; attribute groups; business rules that set default attribute values, change attribute values and send email notifications when data fails validation; validation of a version against business rules; derived and explicit hierarchies; collections; user-defined metadata; version locking, committing and version flags so "subscribing systems identify which version of a model to use"; subscription views; and user and group permissions. The same page records that Master Data Services is removed in SQL Server 2025. Cited for what master data work consists of, not as a recommendation of a product. https://learn.microsoft.com/en-us/sql/master-data-services/master-data-services-overview-mds

[2] Boyuan Li, AI Washing (University of Florida, working paper, August 2025) — constructs "AI talk" from forward-looking in-house AI investment claims in quarterly earnings conference calls and "AI walk" from AI-related workforce expertise in employee résumés, for U.S. public firms 2016–2024; reports that AI walk, but not AI talk, predicts subsequent AI patent quantity and quality; that within firms past AI talk does not forecast future AI walk; that the market "rewards talk in the short run but discounts it in the long run"; and that the share of firms mentioning "AI" on conference calls rose from virtually zero in 2016 to about 20% by mid-2024. A working paper, not a peer-reviewed publication, and about capital markets rather than internal budgeting; cited on both counts with those limits stated in the text. Abstract and introduction read at source August 22, 2026. https://bpb-us-e2.wpmucdn.com/sites.utdallas.edu/dist/8/1090/files/2025/09/AI_Washing_Boyuan_Li.pdf

[3] U.S. Securities and Exchange Commission, press release 2024-36, SEC Charges Two Investment Advisers with Making False and Misleading Statements About Their Use of Artificial Intelligence (March 18, 2024) — settled charges against Delphia (USA) Inc. and Global Predictions Inc., $400,000 in total civil penalties ($225,000 and $175,000); the term "AI washing" used by the SEC chair in the release; Global Predictions found to have falsely claimed to be the "first regulated AI financial advisor." Cited for the existence and terms of the enforcement action. https://www.sec.gov/newsroom/press-releases/2024-36

[4] U.S. Federal Trade Commission, FTC Announces Crackdown on Deceptive AI Claims and Schemes (September 25, 2024) — five law enforcement actions under the sweep named Operation AI Comply, with the FTC chair quoted that "there is no AI exemption from the laws on the books," and the release stating that firms "have seized on the hype surrounding AI." Cited for the existence and framing of the sweep. https://www.ftc.gov/news-events/news/press-releases/2024/09/ftc-announces-crackdown-deceptive-ai-claims-schemes

[5] Informatica, New Global CDO Report Reveals Data Governance and AI Literacy as Key Accelerators in AI Adoption (news release, January 27, 2026), reporting the "CDO Insights 2026" study — a survey of 600 global data leaders across the U.S., UK/EU and APAC; 86% "proactively increasing their data management investment in 2026," with top drivers improving data privacy and security (43%), enhancing data and AI governance (41%) and upskilling employees to improve data and AI fluency (39%); also 69% GenAI integration, 47% agentic AI adoption, 57% naming data reliability a barrier to production, and 76% saying AI governance does not keep pace with employee AI use. A vendor-sponsored survey of self-reported intent; cited here as the strongest competing position to this article's argument, with that limitation stated in the text. https://www.informatica.com/about-us/news/news-releases/2026/01/20260127-new-global-cdo-report-reveals-data-governance-and-ai-literacy-as-key-accelerators-in-ai-adoption.html

[6] Scaled Agile, Inc., Weighted Shortest Job First (Extended SAFe Guidance, read August 22, 2026) — defines WSJF as "a prioritization model used to sequence work for maximum economic benefit," estimated as "the relative cost of delay divided by the relative job duration," with backlogs prioritized on relative user and business value, time criticality, risk reduction and/or opportunity enablement, and job size; opens with Don Reinertsen's line from The Principles of Product Development Flow: "If you only quantify one thing, quantify the Cost of Delay." The definition and opening section are public; the remainder of the page is behind a login and was not read. Reinertsen's book is cited bibliographically and was not opened. https://framework.scaledagile.com/wsjf

[7] ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system (International Organization for Standardization / International Electrotechnical Commission, edition 1, published December 2023; 51 pages; ISO/IEC JTC 1/SC 42) — cited for its existence, scope and status as the first international AI management system standard. The catalogue entry was read at source August 22, 2026; the standard itself was not opened, being available only for purchase (CHF 225), so nothing here characterizes its clauses beyond its published title and scope. https://www.iso.org/standard/42001

[8] Microsoft, Microsoft Foundry Models lifecycle and support policy (Microsoft Foundry documentation, read August 22, 2026) — a five-stage lifecycle (Preview, Generally Available, Legacy, Deprecated, Retired); a GA model's "retirement date (18 months out) is set programmatically" at launch; deprecation to existing customers only at 12 months; the official replacement "selected and declared approximately 90–120 days before the retiring model's retirement date—not sooner"; at retirement "all inference returns 410 Gone"; "Provisioned deployments are NOT auto-upgraded" and must be migrated manually; at least 60 days' notice for GA retirements and 30 for preview; and, to "Can I get an exception to extend a model's retirement date?", the answer "No. Retirement dates aren't extendable." Cited for the documented lifecycle terms, not as a criticism of this platform, whose policy is published in more detail than most. https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/model-retirements

[9] Microsoft, Tutorial: Create and manage budgets (Microsoft Cost Management documentation, read August 22, 2026) — "Notifications are triggered when the budget thresholds are exceeded. Resources aren't affected, and your consumption isn't stopped"; "Cost and usage data is typically available within 8-24 hours and budgets are evaluated against these costs every 24 hours"; "When a budget threshold is met, email notifications are normally sent within an hour of the evaluation"; budgets require at least one cost threshold and one email address, and can trigger an action group. Cited for the documented behaviour of the control, which is the point being made about it. https://learn.microsoft.com/en-us/azure/cost-management-billing/costs/tutorial-acm-create-budgets

[10] Ralph Peeters, Aaron Steiner and Christian Bizer, Entity Matching using Large Language Models, arXiv:2310.11244 (v4, October 2024; published in the Proceedings of the 28th International Conference on Extending Database Technology, EDBT 2025) — investigates generative LLMs as an alternative to fine-tuned pre-trained language models for entity matching; reports that "the best LLMs require no or only a few training examples to perform comparably to PLMs that were fine-tuned using thousands of examples," that "LLM-based matchers further exhibit higher robustness to unseen entities," that there is "no single best prompt" and prompts must be tuned per model/dataset combination, and that GPT-4 can generate structured explanations of matching decisions and identify classes of matching error. Abstract read at source August 22, 2026; the full paper was not read. https://arxiv.org/abs/2310.11244

[11] Tianshu Wang, Xiaoyang Chen, Hongyu Lin, Xuanang Chen, Xianpei Han, Hao Wang, Zhenyu Zeng and Le Sun, Match, Compare, or Select? An Investigation of Large Language Models for Entity Matching, arXiv:2405.16884 (v3, December 2024; COLING 2025) — describes current LLM-based entity matching approaches as typically following "a binary matching paradigm that ignores the global consistency among record relationships," compares matching, comparing and selecting strategies, and proposes the ComEM framework, evaluated on 8 entity-resolution datasets and 10 LLMs, with "further cost-effectiveness" reported as a result. Abstract read at source August 22, 2026; the full paper was not read. https://arxiv.org/abs/2405.16884

[12] Donghao Huang, Prasan Pai and Zhaoxia Wang, Entity Matching with LLMs at Scale: Accuracy, Cost, and Reasoning Effects, in Trends and Applications in Knowledge Discovery and Data Mining (PAKDD 2026), Lecture Notes in Computer Science vol. 16603, pp. 112–124, Springer (first online July 14, 2026; DOI 10.1007/978-981-92-2014-4_10) — benchmarks 46 model configurations across 8 entity-resolution datasets under a fixed protocol, quantifying "API-equivalent token cost" alongside accuracy; reports that "open-weight models can rival proprietary APIs: GPT-OSS:120b attains 92.10 F1 while costing ~20x less than GPT-5(high)", that reasoning-effort settings have opposite effects across model families, and that "several ultra-budget configurations under $0.10 reach up to 89 F1, enabling inexpensive deployment"; a supplementary note records the replication of the original ComEM result at 85.9 against 86.4 mean F1. Cited as the published case against the cost half of this article's argument, which it narrowed. Abstract, author affiliations, reference list and supplementary-results notes read at source August 22, 2026; the chapter body is behind a paywall (USD 29.95) and was not read, so nothing here characterizes its method beyond what the abstract and supplement state. https://link.springer.com/chapter/10.1007/978-981-92-2014-4_10

[13] National Institute of Standards and Technology, AI Risk Management Framework (programme page, read August 22, 2026) — the AI RMF 1.0 released January 26, 2023, described as "intended for voluntary use and to improve the ability to incorporate trustworthiness considerations into the design, development, use, and evaluation of AI products, services, and systems," developed through a public consensus process; with a companion AI RMF Playbook, a Roadmap, a Crosswalk, and NIST AI 600-1, the Generative Artificial Intelligence Profile (July 26, 2024); the page also records that AI RMF 1.0 is being revised as part of the White House AI Action Plan. The four functions named in this article — govern, map, measure, manage — are taken from the framework diagram on that page. The programme page was read at source; the framework document (NIST AI 100-1) was not opened. https://www.nist.gov/itl/ai-risk-management-framework

[14] Semarchy, The State of Data Management in 2026: Where Ambition Meets the Reality Check (report landing page, read August 22, 2026) — described as based on "a global survey of 1,000 C-suite executives across the US, UK, and France"; states that "exactly half of organizations have adopted Master Data Management as a foundation for AI. The other half are scaling on fragmented data," and that "for the first time, data management and governance rank above cost and talent as the primary obstacle to AI success"; lists among its themes "establishing MDM as a non-negotiable layer for governance and trust." A vendor-sponsored survey, and the landing page was read rather than the report itself, which is behind a form; the figures quoted here are the ones the vendor publishes openly. Cited as a competing position on the amount, and for its own framing of the discipline. https://semarchy.com/resources/ai-data-management-report-2026/