Weaving Intelligence

The Voice Problem

Every voice in this publication sounded the same. What we found when we measured it, what we did about it, why the first fix did not hold, and what the pause has cost.

Every voice in this publication sounded the same. Not the same subject — the same person. Six writers with different regions, hobbies, temperaments and beats had produced fifteen articles a reader could not tell apart, and no check we had could see it. We stopped publishing to fix it properly, and we are still stopped — the first attempt did not work, and the second is still being tested. This page is what we found, what we did, what it cost, and where it stands.

Dated, because a page that says “a week” forever is a page nobody updated. Publishing paused 31 August 2026. The first fix shipped rewrites of ten published articles; the editor read the next four and found them still recognisably one writer. The cause turned out to be close to the opposite of what we assumed. Cadence resumes 14 September 2026, and if that slips this line moves with it.

How we found it

Not through any automated check. Through reading. Weaving Intelligence publishes five days a week, one voice per day, and a subscriber meets them one at a time — which is exactly the reading pattern that hides the problem. Jeff was editing them back to back. Read that way, in one sitting, they were obviously one writer in six costumes: the same argument shape, the same beats, the same closing move, the same rhythm.

The suspicion came first and the measurement came second, which is the right order but an uncomfortable one. Confirming it took a purpose-built comparison across the whole set, and building that instrument took longer than reading the articles did — a crude version counting only the small connecting words writers use unconsciously missed the rhythm and punctuation that carry real individual signal, and an early version fed the machine-assembled reference list at the foot of each article into the comparison along with the prose. That bibliography is roughly a fifth of the words and none of the voice; including it made the writers look further apart than they are. The settled instrument reads authored prose only, and the number below is its number.

What was actually wrong

10 of 15 Articles whose closest stylistic match was written by a different voice than the one on the byline.

Three more measurements, across the fifteen articles in the corpus — ten published between 17 and 28 August, five written and still in draft:

What we measuredWhat we found
Where the practitioner section sat in the article Between 47% and 72% of the way through — in all fifteen, mean 59%
How each article ended Thirteen of fifteen closed on a question the reader could ask at work
How far apart the voices were, statistically A separation ratio of 0.99× — no clustering at all

That last number is the honest one, and it is the worst of the three. A ratio of 1.0 means the difference between two articles by the same voice was exactly as large as the difference between two articles by different voices — which is another way of saying the byline carried no information. We had assumed the voices clustered weakly. Measured on the prose they actually wrote, they did not cluster.

Why it happened

Each voice had a detailed character document — where they are from, what they have lived through, how they argue, what they refuse to say. Those documents were good, and they were being used. But the instructions that actually shaped an article — how long it should be, what sections it needed, where the practitioner material went — lived in a single shared brief that was byte-for-byte identical for all fourteen writers, and it outweighed the character documents by roughly 1.4 to 1.

Worse than the ratio: the shared instructions were enforced and the character instructions were merely stated. Automated checks refused to publish an article with a missing citation, a missing section, or the wrong word count. Nothing checked whether it sounded like the person whose name was on it. So generation optimised for what could stop it, and the shape of the checks became the house voice.

Whatever you enforce hardest becomes what your output is — not merely what it satisfies. That is the transferable lesson here, and it is not really about AI. It is about what happens to any process when one set of rules has teeth and another has good intentions.

The blind spot underneath it

We had a dense set of quality checks: citation integrity, evidence gathered against our own claims, required sections, length, independent review. Every one of them examined one article at a time. Our independent reviewers read one article at a time. Nothing in the entire apparatus had ever compared two articles to each other.

So the sameness was invisible by construction, not by oversight. Every check was green, truthfully, all the way through. The only instrument that could see a defect living between articles was a person reading them consecutively.

If you are running AI-assisted production of anything at volume, that is the question worth stealing: what is the largest thing any of your checks ever looks at? Anything bigger than that is invisible to all of them at once.

What we tried first, and why it failed

The obvious fix is to ask the voices to be more distinctive. We did that properly: all fourteen were interviewed separately, in parallel, each with access only to its own character document and no sight of the others or of the published articles. Each was asked how it builds an argument.

They converged.

Asked independently, with no contactAgreed
Refused to structure an article as a numbered list14 of 14 — ten of them giving the same reason
Refused to put disagreement in its own section14 of 14
Said they open on a concrete object rather than a scene13 of 14
Said they end on a question the reader can ask at work13 of 14 — eight named a specific weekday, and seven of those said Monday

Four of them, independently, used the same word — weather — to describe the kind of opening they were rejecting.

This was the most useful failure of the week. Difference cannot be elicited. It has to be assigned. Ask a generator to diverge from nothing and it converges on its own average, then describes that average using the vocabulary of divergence. The fourteen answers about being distinctive were themselves indistinguishable.

What we did instead

We stopped asking and started allocating.

  1. We went and found real variety rather than inventing it. Writers have been diverging in public for centuries under competitive pressure, so the differences are observable rather than imaginable. We characterised over seventy distinct article structures drawn from long-form journalism, newspapers and columns, technical and engineering writing, the essay, scientific and medical papers, business and analyst writing, and older rhetorical traditions — measuring where each one turns, where it puts disagreement, how it ends, how it paces itself, how long it runs.
  2. We placed each voice at a distinct coordinate in that space. Five structures each, chosen to fit the person already described in their character document, and no two voices share one.
  3. We gave every voice a list of things it may not do. This is the half that actually produces difference. A move assigned to one writer is forbidden to the other thirteen — because variety comes from subtraction, not addition. Where all fourteen had refused something, somebody who could do it well was given it, or nobody stands where everybody stands.
  4. We moved shape out of the shared brief and into each character document, and cut the shared brief by roughly a third.
  5. We built the check that was missing — one that compares an article against every other article we have published, rather than examining it alone.

What changed, measurably

BeforeAfter
Where articles turnOne band, 47–72%Deliberately spread from 4% to 90%, plus several with no turn at all
How articles endOne kind, in 13 of 1527 distinct kinds assigned across the roster
Article lengthRoughly 5,000–9,000 words, everyone250 to 45,000, depending on the structure — a permitted range, not a plan; most sit well under 5,000
Shared instruction vs. character1.4 : 1 against characterRoughly 1 : 1

The honest caveat: those are properties of the instructions. The proof is in the articles, and the articles are being rewritten now. Each one, as it republishes, links to the version it replaces, and both sit side by side below.

What the measurement says

Every rewritten article is measured against the version it replaces. That comparison is unusually clean, and not by our cleverness: the standard problem in measuring writing style is that you end up measuring the subject instead, because different writers write about different things. Here both halves of every pair are the same article on the same topic in the same slot. The subject is not merely matched. It is identical.

12 of the 15 can be measured against the version they replace. Everything below describes those 12. The rest are not in here in any form — not averaged in, not estimated, not projected. (How many rewrites exist is a different count, and the pair list further down is the one that reports it: a rewrite can be finished and still be waiting on the measurement that needs both halves.)

Voice Axes on target Separation margin Banned words Topic held
Jordan M. 8/18 → 8/18 -0.0046 → 0.0306 0 → 0 0.51
Jordan M. 7/18 → 8/18 -0.0294 → 0.0508 1 → 0 0.70
Jordan M. 6/18 → 10/18 -0.0229 → 0.0692 0 → 0 0.70
Catherine D. 8/18 → 11/18 0.0000 → 0.0000 7 → 0 0.61
Marcus B. 9/18 → 10/18 0.0000 → 0.0000 11 → 1 0.44
Maya R. 5/18 → 11/18 0.0550 → 0.0676 2 → 1 0.58
Maya R. 7/18 → 10/18 0.0512 → 0.0473 3 → 0 0.64
Maya R. 5/18 → 12/18 0.0534 → 0.0470 0 → 0 0.60
Elias K. 9/18 → 12/18 0.0258 → 0.0396 0 → 0 0.80
Elias K. 7/18 → 11/18 0.0142 → 0.0143 0 → 0 0.68
Elias K. 6/18 → 11/18 -0.0078 → 0.0126 1 → 0 0.73
Isabel B. 4/18 → 10/18 -0.0198 → 0.0507 3 → 0 0.47
Isabel B. 4/18 → 12/18 -0.0139 → 0.0079 3 → 1 0.60
Isabel B. 6/18 → 11/18 -0.0288 → 0.0076 6 → 0 0.67

The instructions are being followed. Each voice has a measured coordinate on 18 axes — sentence length and its variance, how often a sentence runs past forty words, monosyllable share, Latinate share, semicolons per ten thousand words, and so on — each with its own tolerance derived from how strong a habit it is rather than a single invented margin. The rewrites landed inside 10.5 of those 18 on average, against 6.5 before, and the average miss shrank from 1.80 to 1.22 standard deviations. 13 of the 14 moved more axes inside, 1 was unchanged, and none went backwards. Separately, the forbidden-word lists went from 37 violations to 3 across the set.

And here is the instrument that disagrees. Everything in the paragraph above is partly circular: those coordinates were derived with our own corpus in view, so "the articles moved toward the targets" is a weaker claim than it sounds. The independent measure is a stylometric distance — function words, sentence rhythm, punctuation — that never sees the coordinates at all. It asks whether an article sits closer to other articles by its own voice than to articles by the other thirteen. That margin improved in 10 of 14, by 0.0266 on average, against a between-voice distance in this archive of 0.1170. Two went backwards.

The part of that worth stopping on is the sign. A negative margin does not mean a voice was weakly distinguished. It means the article was closer to other people's articles than to its own author's — separated the wrong way round. Before the rewrite, 5 of these 12 were on the right side of zero. After, 12 of 12 are. On the strictest version of the test — is the single nearest article in the archive by the same voice? — there is no reliable movement yet.

So the honest reading is: the layer is doing what it was told, and that has not yet turned into robust separability. Those are different claims and only the first is established.

Two controls, because the numbers above are worthless without them. The prompts contain sample passages in each voice, so an article could score well simply by copying their phrasing — a real effect, measured elsewhere at 28 percentage points of inflation. Across all 14 rewrites there are 4 shared eight-word runs with the samples, 0.07% of any article at most. Nothing here is copying. And content overlap between each pair holds at 0.62, so what changed is the writing and not the article.

What we can and cannot claim, and the evidence against us. An earlier version of this page said a signed-rank test at this sample size could not reach significance no matter what the data say. That was wrong twice over, and it was wrong in our own favour's opposite direction: the smallest attainable p-value here is 0.0005, and the test was run. On the separation margin, across 12 pairs, it returns p = 0.0049 with a matched-pairs rank-biserial correlation of 0.87. That is a real result and declining to report it was not scrupulousness, it was an error. The strictest test is the one that stays negative: asking whether an article's single nearest neighbour is by its own voice gives p = 0.2852 over 9 usable pairs — no movement we can stand behind. So one instrument says the separation improved and the other says the ranking has not, and both belong on this page. More uncomfortably, the closest published work to what we are doing — handing a generator a written style profile — found it gained essentially nothing against a trained authorship model, and concluded the bottleneck is not instruction quality. That result measures generated text against a real human author, which is a harder target than ours; the same work finds machine-versus-machine separation is measurable. So the claim this publication makes is the narrow one: that these voices are becoming distinguishable from each other. Not that any of them has become a person.

For scale, the archive as it stood before any of this: across 15 articles the nearest neighbour shared its author's voice 5 times, and one single article was the closest match to 5 of the others. The raw data behind every figure on this page, including the runs that went the wrong way, is generated by the same tool that produced the table.

Why we stopped publishing

Because the alternative was to publish more of the problem while describing the fix.

Weaving Intelligence exists to argue that trust in data is engineered rather than assumed, and that standards which are not enforced are not standards. A publication making that argument in fourteen indistinguishable voices is refuting itself in the medium. Shipping five more articles on the old pattern to protect a streak would have cost more than the streak was worth.

The cost, updated 7 September 2026, is two full publishing weeks — ten articles that will arrive late, and a rewrite of every one already out. It was one week when this paragraph was written. The first fix did not hold, the editor read the next four articles and still heard one writer, and a second week was the honest price of not shipping the same problem twice.

It is still the smallest this problem will ever be. The archive is fifteen articles rather than five hundred, and every week we wait multiplies a rewrite we have now done once and would rather not do a third time.

What we would tell someone doing the same thing

  • Check between the units, not just inside them. Whatever your quality checks examine, defects larger than that are invisible to every one of them simultaneously.
  • Enforce the thing you actually care about. Anything merely encouraged loses to anything enforced, and you will not notice because the enforced checks stay green.
  • Do not ask for variety. Allocate it. And write down what each voice may not do — the forbidden list does more work than the description.
  • Read your output consecutively, in one sitting, on purpose. It takes an afternoon, and no tool we owned would have found this first.

Before and after

Every article published before this change is preserved exactly as it was, and every rewritten article links back to it. Nothing has been quietly edited or removed. The pairs appear here as each one is completed.

14 of 15 rewritten so far. Every original stays live and unedited until its replacement is ready, and each pair appears here as it is completed.

Before · 2026-08-17

We've Made This Case Before: How the Enterprise MDM Business Case Grew Up

Elias K. · 7,252 words. Published on this date, preserved unedited.

After · rewritten

The Case Was Already Assembled

Elias K. · 4,537 words, rewritten through the voice layer.

Before · 2026-08-18

The Business Case We Built From the Demo

Isabel B. · 6,587 words. Published on this date, preserved unedited.

After · rewritten

The Business Case, Priced in Days

Isabel B. · 3,847 words, rewritten through the voice layer.

Before · 2026-08-19

Costing the Promise: The Parts of an AI-MDM Business Case You Can Actually Check

Maya R. · 6,726 words. Published on this date, preserved unedited.

After · rewritten

Costing the Promise: The Parts of an AI-MDM Business Case You Can Actually Check

Maya R. · 5,195 words, rewritten through the voice layer.

Before · 2026-08-20

The Number Was Never Doing the Work: AI, BI, and the Limits of Better Evidence

Jordan M. · 5,635 words. Published on this date, preserved unedited.

After · rewritten

The Number Was Never Doing the Work: AI, BI, and the Limits of Better Evidence

Jordan M. · 9,486 words, rewritten through the voice layer.

Before · 2026-08-21

The Strategy Document Is Newer Than the Argument It Settles

Marcus B. · 9,027 words. Published on this date, preserved unedited.

After · rewritten

Does Being Named the Foundation Get a Master Data Programme Funded?

Marcus B. · 3,595 words, rewritten through the voice layer.

Before · 2026-08-24

One Sponsor Is a Single Point of Failure: Sustaining Executive Support for an MDM Program

Elias K. · 8,805 words. Published on this date, preserved unedited.

After · rewritten

One Sponsor Is a Single Point of Failure: Sustaining Executive Support for an MDM Program

Elias K. · 7,308 words, rewritten through the voice layer.

Before · 2026-08-25

Nobody Wins a Sponsor in Flight

Isabel B. · 7,589 words. Published on this date, preserved unedited.

After · rewritten

Nobody Wins a Sponsor in Flight

Isabel B. · 4,821 words, rewritten through the voice layer.

Before · 2026-08-26

Load-Bearing: The Question to Ask Before AI-Enhanced MDM Becomes Your Differentiator

Maya R. · 6,771 words. Published on this date, preserved unedited.

After · rewritten

Filed Under Plants: Differentiator, Cost Centre, and the Question Neither Label Asks

Maya R. · 4,942 words, rewritten through the voice layer.

Before · 2026-08-27

Ideally, Nothing: What Agentic AI Actually Changes About Trusted Master Data

Jordan M. · 8,965 words. Published on this date, preserved unedited.

After · rewritten

Ideally, Nothing: What Agentic AI Actually Changes About Trusted Master Data

Jordan M. · 5,670 words, rewritten through the voice layer.

Before · 2026-08-28

Before It Had a Name: How Master Data Ended Up at the Bottom of Every Governance Framework

Catherine D. · 7,757 words. Published on this date, preserved unedited.

After · rewritten

Before It Had a Name: How Master Data Ended Up at the Bottom of Every Governance Framework

Catherine D. · 4,342 words, rewritten through the voice layer.

Before · 2026-08-31

Master Data Doesn't Transform Anything: A Playbook for Serving Somebody Else's Transformation

Elias K. · 6,824 words. Never published — held back when the cadence paused, and put up here unedited so the comparison is real.

After · rewritten

Master Data Doesn't Transform Anything: A Playbook for Serving Somebody Else's Transformation

Elias K. · 4,109 words, rewritten through the voice layer.

Before · 2026-09-01

Profisee Decides What. Other Systems Decide How.

Isabel B. · 7,509 words. Never published — held back when the cadence paused, and put up here unedited so the comparison is real.

After · rewritten

Profisee Decides What. Other Systems Decide How.

Isabel B. · 4,841 words, rewritten through the voice layer.

Before · 2026-09-02

Reinforce Is the Wrong Word: What an AI Roadmap Actually Does to an MDM Roadmap

Maya R. · 6,573 words. Never published — held back when the cadence paused, and put up here unedited so the comparison is real.

After · rewritten

Reinforce Is the Wrong Word: What an AI Roadmap Actually Does to an MDM Roadmap

Maya R. · 4,476 words, rewritten through the voice layer.

Before · 2026-09-03

Show Me the Receipts: Decision Intelligence, Governed Master Data, and a Maturity Nobody Has Earned

Jordan M. · 5,853 words. Never published — held back when the cadence paused, and put up here unedited so the comparison is real.

After · rewritten

Show Me the Receipts: Decision Intelligence, Governed Master Data, and a Maturity Nobody Has Earned

Jordan M. · 3,593 words, rewritten through the voice layer.

Before · 2026-09-08

The Grade Is Not the Climb

Isabel B. · 6,212 words. Drafted, never published — held back when the cadence paused.

After · in progress

Not yet rewritten

This one is still being rebuilt. The original stays live and unedited until its replacement is ready.

What you can compare today are the voice profiles themselves. Each was a single paragraph — between 215 and 317 characters, which was the entire page — while a full background and a short Q&A sat written and unpublished. Those are live now, roughly twenty-four times longer. The version that stood until 26 August 2026 is preserved for each one and linked from the profile, so that comparison is available immediately rather than promised.


Every article in this publication is written by an AI voice and reviewed by Jeff Shabel before it publishes. This page was written the same way. How Weaving Intelligence is written →