What a voice file cannot fix

Everything else written about brand voice files is upside. This page is the other half, because a tool that cannot say what it fails at is asking you to take the rest on faith.

Two of the refusals below cost this site something. There is no voice score, which is the single most requested feature in the category, and the file does not make writing undetectable, which is the highest-volume adjacent search there is. Both refusals are made for reasons you can check.

It hangs off the verification page, which is where the check you can run yourself lives. The questionnaire itself is eleven questions and about five minutes.

Which instructions do models reliably ignore?

At what length does adherence start to degrade?

Earlier than the character limits suggest, and the vendors say so themselves rather than this being an outside claim. Anthropic targets under 200 lines per memory file and states directly that longer files consume more context and reduce adherence, adding that files are loaded in full regardless of length so nothing truncates: the rules simply get lost.

That is why the output here caps at fifteen binding rules. The cap is enforced by the renderer rather than left to you, because the failure it prevents is invisible: a longer file does not error, it just quietly works less well. The documented caps are collected here.

What happens to a rule that contradicts the model's own default?

It needs restating more explicitly than feels necessary, and it is the class of rule most likely to erode over a long conversation. A default the model holds strongly and your file mentions once will reassert itself as the instruction gets further away in the context.

The practical mitigation is to write it as a prohibition rather than a preference, and to put it in the highest-authority slot the tool offers rather than the most convenient one.

Which claims about brand voice are unfalsifiable?

Most of the ones the category is sold on. A claim is unfalsifiable here if no observation would settle it: that a voice is now consistent, that a brand sounds more like itself, that alignment improved. None of those name a thing you could measure and disagree about.

The falsifiable version is narrow and boring, which is why it is rarely advertised. Did the output use British spelling. Did it use contractions. Did it contain a banned word. Those can be counted, and counting them is what the verification check does.

Why there is no voice score on this site

Because there is no honest one to give, and this is the most requested feature in the category.

Within a single author writing in one register, readability varies from paragraph to paragraph by several grade levels. Any per-paragraph score would therefore report a spread several times larger than the difference it claims to detect between brands. A number that moves that much for reasons unrelated to what it measures is not a measurement, however good it looks on a dashboard.

A score would be straightforward to build and would probably increase how much people trust the output, which is precisely why not building it is worth stating out loud.

What this does not do to AI detectors

Nothing, and any tool telling you otherwise is selling something that does not exist. Stylistic fingerprints survive rewriting at high attribution accuracy, so a voice file changes how writing sounds without making its origin undetectable.

This site does not sell a humaniser, does not claim to defeat detection, and would not be able to prove the claim if it made it. What a voice file does is make output sound like you rather than like the default, which is a different and more useful thing.

Detector vendors themselves concede on their own marketing pages that no detector is fully accurate and that results should not be the sole basis for punitive action, which is worth knowing whichever side of that conversation you are on.

What we would need to measure to say more

A fixed prompt set, a fixed voice file, several models at named versions, a sample large enough to separate signal from run-to-run variation, and a date. Then adherence per instruction shape becomes a number that means something and can be repeated by somebody else.

That harness is not built yet, so this site publishes no adherence rate. When it is, the numbers will carry the model version, the sample size and the date, and the ones that are unflattering will be published too.

There is one more thing that cannot be measured from here at all, and it is a direct cost of the privacy promise. Nothing is stored and nothing is transmitted, so whether your file worked in your tools is invisible to us. We can measure adherence in our own harness and say so. We cannot measure yours and will not imply we can. The check you can run yourself exists because of that gap, not despite it.

Where these come from

Every claim above is quoted from one of these, and each was read on the date beside it. If one of them has changed since, the page is wrong and we would like to know.

Everything above is what it will not do. What it does do takes about five minutes: make the file and check it yourself.