Why Detectors Flag Non-Native English Writers
An AI detector does not read your essay and decide whether you cheated. It measures one thing: how predictable your word choices are. Everything else — the percentage, the red highlighting, the “97% AI” verdict — is a translation of that single number into a language that sounds like an accusation. Once you understand what the number actually measures, the pattern that makes these tools flag non-native English writers stops looking like a glitch and starts looking like arithmetic. Second-language English is, on average, more predictable than native English. Detectors are built to punish predictability. The bias is not an accident in the code. It is the metric doing exactly what it was designed to do, applied to the people it was never calibrated for.
That distinction — mechanism, not motive — matters, because it changes what you do about it. A bug gets patched. A structural feature of how a measurement works does not go away with a software update; it just gets quieter, and the writers underneath it keep getting flagged.
What the detector is actually measuring
Most detectors rest on two statistical ideas: perplexity and burstiness. Perplexity is a measure of surprise. A language model reads your text word by word and asks, at each step, “how likely was that next word?” If your writing keeps choosing the most probable next word — common vocabulary, standard phrasing, safe collocations — perplexity is low. If you keep surprising the model with unusual words and turns, perplexity is high. Burstiness measures variation across the text: the mix of long and short sentences, the swings in rhythm and structure that human writing tends to have and that early machine output tended to lack.
The core assumption is simple and, on its face, reasonable: AI-generated text is low-perplexity and low-burstiness, because a model literally optimizes for the probable next token. So the tools learned a shortcut — low perplexity means machine. Modern products have moved past a raw perplexity score; GPTZero now advertises a seven-component proprietary model trained on student writing. But the foundation is the same. Predictable text scores as artificial. That is the load-bearing assumption, and it is where non-native writers fall through the floor.
Why second-language writing is predictable by design
Here is the fact the whole problem turns on, and it comes straight from linguistics rather than from any grievance about detectors. People writing in a second language use a narrower, more predictable slice of the language. Weixin Liang and colleagues at Stanford put it plainly: non-native English writers exhibit reduced linguistic variability in “lexical richness, lexical diversity, syntactic complexity, and grammatical complexity” (Liang et al., 2023, p. 2). A writer working in their third or fourth language reaches for the word they are sure of, not the rare one. They build sentences on patterns they have drilled. They avoid the idiomatic swerve that a native speaker throws in without thinking.
Every one of those habits lowers perplexity. Not because the writing is worse — a plain, correct, well-argued paragraph can be excellent — but because it is more probable. The detector cannot tell the difference between “this text is predictable because a model generated it” and “this text is predictable because its author was being careful in a language they learned as an adult.” Both produce the same low number. The tool reports the number. The number reads as guilt.
This is the mechanism, stated in one sentence: a detector flags non-native writing because low perplexity is its evidence for “AI,” and low perplexity is also the signature of L2 fluency. The two things it is trying to separate look identical through the only lens it has.
The numbers, with the conditions attached
The Stanford team measured exactly how bad this gets. They ran seven widely used GPT detectors over two piles of essays: 88 essays by US eighth-graders, and 91 TOEFL essays written by non-native English speakers. On the American students, the detectors were near-perfect. On the TOEFL essays, they posted an average false positive rate of 61.22% (Liang et al., 2023, p. 2). All seven detectors unanimously misclassified 18 of the 91 essays — 19.78% — as AI-authored, and 89 of the 91 (97.80%) were flagged by at least one detector (Liang et al., 2023, p. 2). The essays that all seven agreed were “AI” had significantly lower perplexity than the rest, with a p-value of 9.74×10⁻⁵ — statistical confirmation that the tools were tracking predictability, not authorship.
Then they proved the mechanism directly. They took genuine US student essays and prompted a model to “simplify word choices as if written by a non-native speaker.” Nothing about those essays became machine-written — a human still wrote every one; only the vocabulary got plainer. The misclassification rate for that real human writing jumped from 5.19% to 56.65% (Liang et al., 2023, p. 2). Flatten the vocabulary of a native speaker and the detector starts calling them a robot too. The trigger was never authorship. It was always predictability.
And this is not a quirk of student essays. The same team analyzed 1,574 accepted papers from ICLR 2023 and found that abstracts by non-native authors carried lower perplexity than those by native authors even at similar peer-review ratings, with the difference statistically significant (p = 0.033) (Liang et al., 2023, p. 2). The bias survives into professional academic writing. It scales with the writer’s first language, not their competence. This is the direct line from detection to the penalty carried by non-native English writers: the tool mistakes a second language for a language model.
Why “structural” is the right word
It would be comforting to call this a calibration error — feed the tool more non-native training data and the problem dissolves. It doesn’t, and the reason is baked into how the measurement works. Aounon Kumar and colleagues, studying detector reliability from the ground up, warn that a low average error rate hides exactly this failure: a detector “may have very large errors within a sub-population of samples such as essays written by non-native English writers or essays on a particular topic or with a particular writing style” (Kumar et al., 2024, p. 21). The headline accuracy figure a vendor quotes is an average across populations. Non-native writers are a sub-population the average quietly buries.
The same paper shows how brittle the underlying signal is even for its intended target. A light paraphrase drops a soft-watermarking detector’s accuracy from 97% to 80%, and a heavier one to 57%, with the perplexity of the text barely moving (Kumar et al., 2024, pp. 4, 6). If a five-minute rewrite can collapse the score for machine text, the “signal” the detector reads was never a solid boundary between human and AI — it was a fuzzy statistical gradient. Non-native writers happen to sit on the wrong side of that gradient without doing anything at all. As the authors put it, deploying a tool with a high false positive rate “will cause more harm than good in society” (Kumar et al., 2024, p. 6).
Other researchers converge on the same read. A 2025 literature review of 34 studies concluded that most detectors clear 50% accuracy yet remain unreliable for academic use, and its synthesis of the field states flatly that “AI text detectors are unreliable due to their negative bias towards English written by those of foreign origin” (Gotoman et al., 2025, pp. 1, 5). The review also notes the field’s cruel trajectory: the growth rate of AI content generators outpaces that of detectors, so the gap the tools rely on keeps shrinking (Gotoman et al., 2025, p. 7). Halah Mohammed Abbas frames the consequence in the sharpest terms, warning that biased detectors “may wrongly accuse marginalized groups of plagiarism” and that the fairness problem lives “whether through perplexity scores or classifiers” — that is, in both the old method and the new one (Abbas, 2025, p. 14).
That last point is the whole argument in miniature. Swapping raw perplexity for a machine-learned classifier does not remove the bias, because the classifier learns from the same statistical regularities. The predictability of L2 writing is still there in the features. The model just finds it a different way.
What’s happening right now
This is not confined to a 2023 dataset. In the 2024/25 academic year, 1,177,766 international students were enrolled at US institutions — about 6% of all higher-education students, a 5% rise on the prior year (IIE, 2025). More than a million people are writing graded English in a language most of them learned second, feeding assignments into tools that structurally distrust exactly the writing style that second-language learning produces.
The lawsuits have followed the mechanism. In February 2025, Thierry Rignol, a French national enrolled in Yale’s Executive MBA program, sued the university after his final exam was flagged by GPTZero; his complaint argued the tool discriminates against non-native English speakers and that Yale’s own policies caution against detectors precisely because of their high false-positive rates on such writers (Poets&Quants, 2025). A federal judge denied his injunction in May 2025, so he did not walk with his class — the mechanism reached all the way to a diploma. It is the same story we trace through other cases of students falsely accused of using AI: a statistical artifact treated as proof.
Even the vendors have effectively conceded the point. GPTZero now runs an “ESL debiasing” layer aimed specifically at non-native writing and reports cutting its false-positive rate on the original TOEFL essays to roughly 1% — though on that same dataset about 6.6% of essays still come back as “possible AI,” and the figure is self-reported rather than independently verified (TextSight, 2026). Read that as an admission: you only build a debiasing layer for a problem you know your metric creates. The fix is a patch bolted onto a measurement that punishes predictability, not a redesign of the idea that predictability equals guilt.
What follows from the mechanism
If you grade or adjudicate, the practical consequence is specific. A low perplexity score is not evidence of AI use for a non-native writer — it is the expected reading for careful second-language prose, and treating it as a red flag inverts the burden of proof onto the writers least able to absorb it. That is the distinction the accuracy debate around AI detectors keeps missing: even a tool that never technically malfunctions still routes its errors toward one group, because the error is in what it measures, not in whether it measures correctly. The same theme runs through the broader evidence on whether AI detectors are accurate — the number is a probability about predictability wearing the costume of a verdict about honesty.
For the writer on the receiving end, the argument to make is not “the detector made a mistake.” It is stronger than that: the detector did exactly what it was built to do, and what it was built to do cannot distinguish a second language from a machine. The 61.22% false-positive rate, the 5.19%-to-56.65% jump when a human simply used plainer words, the sub-population errors the averages hide — all of it says the same thing. The flag is a fact about your perplexity. It was never a fact about you.
Sources
Every factual claim above is tied to a source you can open and check, with page numbers wherever the source has them.
- 1. Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, James Zou, GPT Detectors Are Biased Against Non-Native English Writers , Patterns; Stanford Institute for Human-Centered AI (HAI), 2023 , pp. 1, 2. link
- 2. Aounon Kumar, Sriram Balasubramanian, Vinu Sankar Sadasivan, Wenxiao Wang, Soheil Feizi, Can AI-Generated Text be Reliably Detected? , arXiv preprint, 2024 , pp. 4, 6, 21. link
- 3. Jezreel Edriene J. Gotoman, Harenz Lloyd T. Luna, John Carlo S. Sangria, Cereneo S. Santiago Jr, Danel Dave Barbuco, Accuracy and Reliability of AI-Generated Text Detection Tools: A Literature Review , American Journal of Interdisciplinary Research and Bibliometrics, 2025 , pp. 1, 5, 7. 10.54536/ajirb.v4i1.3795
- 4. Halah Mohammed Abbas, A Novel Approach to Automated Detection of AI-Generated Text , Journal of Al-Qadisiyah for Computer Science and Mathematics, 2025 , p. 14. 10.29304/jqcsm.2025.17.11958
- 5. Institute of International Education (IIE), Open Doors 2025: International Student Enrollment Data , Institute of International Education, 2025 . link
- 6. Nikki Waller / Poets&Quants, Judge Denies Injunction In Yale Student's AI-Linked Suspension , Poets&Quants, 2025 . link
- 7. TextSight, GPTZero Review 2026 — Accuracy, Pricing and Verdict , TextSight, 2026 . link