Detection is the wrong control
Detection sits at the wrong end of the process. It runs after the writing exists, it produces a probability rather than a fact, and the cost of it being wrong falls entirely on the person accused.
You do not have to take that from us. Almost every source on this page is a detector vendor's own site, and in several cases the disclaimer sits on the same page as the accuracy claim it undercuts.
This sits under the AI writing tells, and the alternative it argues for is the ordinary one: decide how you write, then hold the tools to it.
What do the vendors claim, and who measured it?
The headline numbers are high and they come from the companies selling the product. That is not automatically wrong, and it is worth knowing before the numbers are quoted back at you as though they were independent.
GPTZero leads with 99% accuracy, citing its own large-scale testing alongside a university research lab and a third-party benchmark. Turnitin, in March 2023, published a less than 1% false positive rate, with its chief product officer writing that if we say there is AI writing, we are very sure there is.
What do the same pages concede?
This is the part worth reading, and it is all on the vendors' own sites.
- GPTZero, on the page carrying the 99% figure: no AI detector is 100% accurate, AI itself is changing constantly, and results should not be used to punish students.
- GPTZero's FAQ, on scope: the majority of its training data is English prose written by adults, so accuracy is dataset-dependent by its own account.
- GPTZero's FAQ, on what the number means: the high figures apply within a high-confidence band, not to every result the tool returns.
- GPTZero's FAQ, on granularity: sentence-level classification should not be used on its own to conclude that an essay contains AI writing.
- Turnitin, in the same March 2023 post as the 1% claim: it does not make a determination of misconduct, and because the false positive rate is not zero, the instructor has to apply judgement.
Read together, the offer is a probability with a disclaimer attached, sold into a setting where it will be treated as a verdict.
What happened when the lab numbers met the real world?
Eight weeks. That is the gap between Turnitin's confident March 2023 post and its May 2023 update, and the update is the most useful document in this entire subject because a vendor wrote it against its own interest.
It states that the model was tested in a controlled lab setting before release, and that after release they discovered the real-world behaviour differed. They re-tested by running eight hundred thousand academic writing samples written before ChatGPT existed through the detector, which is a genuinely good experimental design: anything flagged is a false positive by construction.
The conclusion they published from it was that results below a threshold share of detected AI writing in a document are not reliable enough to act on, and they changed the product to say so.
Nothing about that story is disgraceful. Lab-to-deployment gaps are ordinary. It is only damning if you are quoting the March number in 2026 without the May correction, which is what most pages ranking for detector accuracy do.
Who pays when it is wrong?
Not the vendor, and not the person running the check. The cost lands on whoever was flagged, and it does not land evenly.
Published research found detectors are biased against non-native English writers, misclassifying their work as machine-generated at substantially higher rates than writing by native speakers. That is a systematic error aimed at the population least able to contest it.
A false positive here is not a mild inconvenience. It is an accusation of dishonesty that the accused cannot disprove, because there is no evidence that would satisfy it: you cannot produce the absence of a chatbot.
What have institutions actually done?
Backed away from it. Vanderbilt disabled detection in its learning platform and published its reasoning. MIT Sloan's teaching and learning group published guidance under the plain heading that AI detectors do not work.
OpenAI is the most telling case, because it had every commercial reason to succeed and the best possible access to the models being detected. It withdrew its own classifier in July 2023, citing low accuracy.
When the organisation that built the generator cannot build a reliable detector for it, the problem is unlikely to be effort.
What does a constraint at writing time do instead?
It moves the control from after to during, and from probability to fact. Whether a sentence uses British spelling, contains a banned word, or runs past your sentence-length rule is not a judgement call. A checker returns the same answer twice.
That is a smaller claim than detection makes and it is the reason it holds. You are not asking a tool to infer the origin of a text. You are asking whether the text obeys rules you wrote down, which is a question with an answer.
Which rules are actually checkable covers the line between the two, and the verification check is the ninety-second version you can run on your own output.
What this site will not claim
- That a voice file makes writing undetectable. It does not, we have not tested whether it does, and anyone promising it is selling something that does not exist.
- That evading detection is a goal worth having. If disclosure is required where you are, a tool that helps you hide it is a liability rather than a feature.
- That we will build a humaniser. We will not, at any traffic level, and this page exists partly to say so in public where it costs us the search term.
- That detectors are worthless. They are evidence, weakly, in aggregate. They are not proof about one person's document, which is the only way anybody actually uses them.
The output this tool produces is marked as AI-assisted in its own metadata, deliberately, because the honest response to a disclosure obligation is disclosure. The rest of what this cannot fix is listed here.
Where these come from
Every claim above is quoted from one of these, and each was read on the date beside it. If one of them has changed since, the page is wrong and we would like to know.
- GPTZero homepage, accuracy claim and disclaimer checked 2026-08-17
- GPTZero FAQ, scope and confidence bands checked 2026-08-17
- Turnitin, March 2023: understanding false positives checked 2026-08-17
- Turnitin, May 2023: the update after release checked 2026-08-17
- Liang et al., detectors are biased against non-native English writers checked 2026-08-17
- Vanderbilt: guidance on AI detection, and why it was disabled checked 2026-08-17
- MIT Sloan: AI detectors don't work checked 2026-08-17
Constraint at writing time beats inference afterwards, because it answers a question that has an answer: write the rules down and check against them.