← Back to blog

Why AI Detectors Flag Non-Native English Writers

September 8, 2026 · PT Technologies · 6 min read

If you write in English as a second language and an AI detector has flagged your work, the first thing worth knowing is that this is a documented, structural property of how these tools work — not a coincidence, and not a judgement about your writing.

It is also, in our view, the most serious unsolved problem in this product category, and the one most likely to produce a genuinely unjust outcome.

The mechanism

Detectors do not recognise machine text. They measure statistical properties and infer from them. The main ones are:

  • Predictability. How surprising is each word, given the words before it? Generated text tends to pick the likely next word more consistently than a person does.
  • Burstiness. How much does sentence length vary? Human writing tends to lurch — a long sentence, then a short one. Generated text is steadier.
  • Vocabulary spread and connective use. Generated text leans on a recognisable set of transitions and reaches for a narrower band of vocabulary.

Now describe competent second-language academic writing without mentioning AI at all. It tends toward a more controlled vocabulary, because you write the words you are confident about. It tends toward more regular sentence structure, because you have learned reliable patterns and you use them. It uses learned connectivesmoreover, furthermore, in addition, it is important to note — because those are explicitly taught as the way to signal structure in academic English.

Every one of those is a marker of careful, well-taught, competent writing. Every one of them also pushes the statistics toward the AI end of the scale.

The detector is not making a mistake about your writing. It is measuring what it measures, and what it measures does not distinguish "wrote this in a second language" from "did not write this".

What our own rate is

We publish this number because a detector that will not tell you its failure rate on the group most likely to be harmed by it is not being straight with you.

On a frozen evaluation set of non-native human writing that is never used in training, our current false-positive rate is 0.67%.

Two things about that figure matter more than its size.

It is measured, not fitted. The evaluation set is held out permanently. It is not part of the training data, it is not tuned against, and its only job is to tell us when a change we made for accuracy has cost fairness. That has happened: in one earlier version, adding hard negatives improved robustness against unseen generators and hurt the non-native rate, pushing it from 0.8% to 1.5%. We could see that only because the eval was frozen.

0.67% is not zero, and a percentage is a headcount. In a cohort of 3,000 international students submitting one essay each, a 0.67% rate is twenty people wrongly flagged. If your institution runs every assignment through a detector across a year, multiply again. The arithmetic that turns a small-sounding rate into real people is the same one we set out in a 1% false positive rate is not small.

We also cannot tell you what the equivalent number is for Turnitin, GPTZero or anyone else, because they do not publish it and the evaluation cannot be run from the outside. The published research that exists on this — the widely cited finding that a majority of TOEFL essays were misclassified as AI-generated by several detectors — suggests the problem is common and, in some tools, very large. Treat any vendor's headline accuracy figure as silent on this question unless they say otherwise.

Why this is worse than a normal error

Two things compound it.

The errors are not random, they are concentrated. A detector that is wrong 1% of the time uniformly is a nuisance. A detector that is wrong disproportionately about one group is something else: the same students face the same elevated risk on every assignment, all year. The unfairness accumulates on the same people.

The people it lands on are least able to contest it. A student writing in a second language, often far from home, often on a visa where a misconduct finding carries consequences well beyond a grade, is the person least equipped to push back against an authoritative-looking percentage produced by software the institution bought.

That combination is why we treat this as the central fairness problem rather than one limitation among many.

What to do if it is you

Ask what the score is being used for. A detector output should be one input to a conversation, never the conclusion of one. If a percentage is being presented as proof, that is the thing to challenge — not your writing.

Bring your drafts. Version history is the strongest evidence of authorship available to you, and it is far more persuasive than arguing about statistics. Google Docs and Word both keep it automatically. Start now, before you need it.

Point at what was flagged. Ask which passages drove the score. In our experience they are very often the most formulaic sections — the methods, the literature summary, the standard framing sentences — which are formulaic because that is what those sections are supposed to be.

Name the limitation explicitly. You are allowed to say: this tool has a known elevated error rate on writing by non-native English speakers, and I am a non-native English speaker. That is a documented property of the technology, not a special pleading.

If you want to see it before someone else does, our AI Detector reports the score with its margin of error, the sentences that drove it, and a note when the text sits outside what the model was calibrated on. Knowing which of your paragraphs read as formulaic is useful information regardless of what any detector says.

What institutions should take from this

If you set policy, three things follow directly.

A detector score should never on its own trigger a misconduct process. It is a prompt to look, not a finding.

Any process built on one should ask, explicitly and early, whether the student writes English as a second language — because the tool's error rate is not uniform and pretending it is produces systematically unfair outcomes.

And the question to put to a vendor is not "how accurate is it". It is "what is your false positive rate on non-native English writing, measured on data you did not train on". If they cannot answer, that is the answer.


Our position on our own tool is the same as on everyone else's: the score is a signal, not a verdict, and it is least reliable exactly where the stakes are highest — short work, edited drafts, and writing by people composing in a language that is not their first. That limitation is printed on the tool itself, not buried here. If you have already been flagged, flagged as AI when you wrote every word covers what to do next.