Why AI writing sounds the same

You noticed the output sounds like everything else, and every explanation you can find says the same thing: the model does not know your brand, so give it more context. That is a description of the symptom being sold as a diagnosis.

There is an actual mechanism, it is published, and it is more interesting than the folk version. The flattening is not inherited from what the model read. It is introduced later, when the model is aligned to human preferences, and it has been measured on real preference data.

This sits under the AI writing tells. If you want the practical end rather than the mechanism, eleven questions produce the constraints that survive it.

Is there an actual mechanism, or is this a vibe?

There is a mechanism, and it has a name: typicality bias in preference data. The argument is that when people rate one model output against another, they systematically prefer the more familiar-sounding one, and that preference is baked into the reward model the alignment process optimises against.

The consequence is uncomfortable. If the bias sits in the data rather than in the algorithm, then a perfect reward model and perfect optimisation would still encode it. It is structural, not a bug in the training run.

What has actually been measured?

Enough to take it seriously, from several independent directions. These are other people's numbers; this site runs no experiments and publishes none.

Published findings on where the flattening comes from
FindingDetail
Typicality bias is real and significantFitted to a preference dataset over 6,874 correctness-matched response pairs, giving a bias term of about 0.57 and 0.65, both at p below 10 to the minus 14
Alignment introduces it, pretraining does notDiversity drops after the supervised and preference-optimisation stages, and the aligned model still retains the base model's diversity underneath
Instruction tuning is the step that mattersWriting with an instruction-tuned model produced a statistically significant diversity reduction; the same task with the base model did not
The loss is the model's, not the writer'sIn that study the reduction was attributable to the model's contributions rather than to what the humans wrote
It is a known trade-off, not an accidentPreference optimisation reduces output diversity relative to supervised fine-tuning while generalising better to new inputs
Bigger models do not fix itStylistic deviation from human writing persists with scale and is larger for instruction-tuned models than base models

What is not the cause?

This section exists because the cluster is full of confident explanations that the evidence does not support, and saying so is more useful than adding another one.

Why does more prompting not fix it?

Because you are asking the model to move against a pull that was optimised into it, using the same channel the pull acts on. More words about how you want it to sound do not change the distribution those words are sampled from.

There is a published prompting technique that does move the needle, and the shape of it is instructive: instead of asking for one answer, you ask the model to lay out a spread of possible answers and their likelihoods. Reported diversity gains of roughly 1.6 to 2.1 times over asking directly, on creative writing tasks.

Note what that is not. It is not a voice. It is a way of getting more variety out of a model that has been trained towards the middle. Variety and your voice are different problems, and only one of them is solved by writing your rules down.

What does a constraint have to look like to survive the pull?

Specific, mechanical, and checkable. The pull is towards the typical, so a constraint phrased as a preference is a suggestion the model can satisfy by being slightly less typical and calling it done.

What we cannot explain

Why particular words rise and fall. The vocabulary turns over between model generations and no published account explains the specific choices, only the general pull towards the familiar.

How much of any given draft is the flattening rather than the prompt. Nobody has separated those cleanly, and this site will not pretend to: we publish no measurements of model behaviour and every number on this page belongs to somebody else.

Where these come from

Every claim above is quoted from one of these, and each was read on the date beside it. If one of them has changed since, the page is wrong and we would like to know.

The pull is towards the middle, so the fix is a constraint specific enough to resist it: answer eleven questions and get one.