← Back to blog

A 1% False Positive Rate Is Not Small

August 11, 2026 · PT Technologies · 8 min read

A vendor tells you their AI writing detector has a 1% false positive rate. That sounds like a rounding error. It is the single most misleading number in this entire product category, and not because it is untrue.

There are two separate problems with reading it the way it invites you to. The first is scale. The second is that it answers a question nobody asked.

Problem one: you mark more papers than you think

Vanderbilt University did this arithmetic publicly when it disabled Turnitin's AI detector in August 2023. It had submitted 75,000 papers to Turnitin the previous year. At the vendor's own claimed 1% false-positive rate, that is roughly 750 papers wrongly flagged in a single year at a single university.

Scale it down to yourself. Three sections of 40, three times a year, is 360 submissions. At 1%, you will wrongly flag between three and four students annually - in a world where not one of them ever touched a language model. At the 2% that independent evaluations more commonly find, it is seven.

You will not experience those as statistics. You will experience them as four conversations a year with a student who is upset, and who is right.

Problem two: the score answers the wrong question

This is the part that survives even if the vendor's number is perfectly accurate, and it is worth slowing down for.

A false positive rate tells you: given that a text is human-written, how often does the tool flag it? What you need to know standing in front of a flagged essay is the reverse: given that this text was flagged, how likely is it that AI wrote it?

Those are different quantities, and the gap between them depends entirely on how common AI use actually is in the population you are testing. It is the same reasoning that governs medical screening, where a highly accurate test for a rare condition still produces mostly false alarms.

Here is a cohort of 120 essays, a detector that catches 80% of genuine AI use, and the vendor's 1% false positive rate - generous assumptions throughout. The first column is how many students actually used AI; the last is the share of the flags on your desk that land on someone innocent.

Real AI useCorrect flagsFalse flagsFalse share
30%28.80.83%
10%9.61.110%
3%2.91.229%
1%1.01.255%

Read the bottom row again. With a detector performing exactly as advertised, in a class where one student in a hundred used AI undeclared, the majority of the students you flag did nothing wrong. At a 2% false positive rate the 3% row rises from 29% to 45%.

The uncomfortable implication is structural: the better your teaching, assignment design and classroom culture work, the less trustworthy each individual flag becomes. Success at deterrence lowers the prevalence, and lowering the prevalence is exactly what tips the ratio against you. A tool that becomes least reliable precisely when you are doing your job well cannot be the thing that decides an allegation.

What the vendors themselves say

It is worth knowing how carefully the tools are worded, because the caution tends to live in documentation while the percentage lives on the screen.

Turnitin's own guidance holds that its score should not be the sole basis for an accusation, and states plainly that its reports may misidentify both human and AI-generated text. It requires a minimum of 300 words before it will score a document, having raised the threshold from 150 because accuracy improves with length. It publishes a sentence-level false positive rate of around 4% - four times the document-level figure - which matters because the highlighted sentences are what most people actually look at. And for documents scoring under 20% AI, it now shows an asterisk instead of a number, because it does not consider those percentages reliable enough to report.

OpenAI, which had every advantage in building one, shut its own detector down in July 2023 after it managed 26% detection at a 9% false positive rate.

We build one of these tools, so let us be equally specific about ours. When we introduced length-aware thresholds, false flags on short human writing fell from 8.75% to 2.75%, and on short writing by non-native English speakers from 10.2% to 2.0%. It cost us real detection ability - recall on short AI text fell from 30.5% to 19.5% - and we took the trade deliberately, because being wrong about a machine costs nothing and being wrong about a person costs a great deal. Every one of those numbers is published, along with how the engine actually works.

Who absorbs the errors

False positives are not distributed randomly across your class, and this is the part with real equity consequences.

Stanford researchers tested seven commercial detectors on 91 TOEFL essays written by non-native English speakers under supervised exam conditions. The tools misclassified them as AI-generated at an average rate of 61.3%, while handling essays by US eighth-graders near-perfectly. Enriching the vocabulary of the same essays dropped the misclassification rate to 11.6% - the detectors were substantially measuring linguistic range and calling the lower end of it artificial. (Liang et al., Patterns, 2023)

Commercial tools have improved since. The underlying mechanism has not changed, because it cannot: writing in an acquired language produces simpler vocabulary, more regular structure and heavier use of taught connectives, and that is a statistical description of the AI signal. Any detector that measures predictability inherits this.

The same logic catches other groups. Students who follow the essay template you taught them. Anyone writing in a formulaic genre - methods sections, lab reports, clinical write-ups - where uniformity is the genre working correctly. Students who revise heavily, because polishing removes exactly the irregularities that read as human. Your most conscientious students are over-represented in your false positives.

Australian Catholic University offers the fully-worked example. Nearly 6,000 misconduct referrals in 2024, around 90% AI-related, roughly a quarter dismissed after investigation, and the detector dropped in March 2025. Students described months-long investigations in which results were withheld while cases resolved. One was cleared after six months, by which point the delay had done its own damage.

What a defensible process looks like

None of this argues for ignoring AI misuse. It argues for treating a score as what it is: a reason to look, never a finding in itself.

Decide the policy before the term, and put it in the syllabus. What uses of AI are permitted, what must be disclosed, and how disclosure is made. Most disputes are really disagreements about a rule nobody stated. A student using a grammar checker and a student generating an essay both currently say "I used AI" and mean entirely different things.

Never let a score be the whole case. Vanderbilt's guidance to its own instructors is a good short list: compare the submission to other work by the same student, check whether cited sources actually exist and say what they are claimed to say, and talk to the student. ACU's own position ended up in the same place - a case resting solely on the detector output was not a case.

Ask for process, not proof of innocence. Version history in Google Docs or Word with AutoSave shows a document accumulating over days. Outlines, notes, annotated readings. A student who can walk you through their draft's growth and discuss their sources has answered the question; one who cannot has not thereby confessed, but you have a better conversation to have.

Hallucinated citations remain the most reliable single signal available to you, and they require no tooling. A reference that does not exist, or exists and does not say what the essay claims, is concrete, checkable, and directly about the work.

Watch the pattern of your own flags. If the flagged names in your gradebook skew toward your international students, the tool is telling you about itself, not about them.

Move quickly. Whatever the outcome, a case that takes months imposes a heavy penalty before anything is established.

Where the score genuinely earns its place

Used as a triage signal on a large pile, a detector points you at the twelve submissions worth reading closely first. That is a real saving, and it is honest work for the tool as long as reading them is the next step.

Used on your own assignment design, a whole cohort scoring high tells you something about the task. Prompts that reward a formulaic answer produce formulaic answers, from people and models alike.

Used by students on their own drafts, before submission, it is straightforwardly useful - and considerably less adversarial than the same tool pointed at them afterwards. Long stretches of uniform, transition-heavy, predictable prose read as artificial partly because they are hard to read. That is worth fixing regardless of who is scanning.

If you want to see what one of these tools reports and how much it hedges, ours is free and needs no account: AI Detector. It highlights the passages that drove the score, explains the measurements behind each flagged sentence, reports how stable the result is under harmless rewording, and says on every page of every report that no detector is 100% accurate - including ours.

And if you are handing a flag to a student, Flagged as AI When You Wrote Every Word is written for the person on the other side of that email. It may be a fairer thing to send than a percentage.