It is a probability estimate from a tool that fails in known, documented ways. If a score has been used against you, this is the pillar that tells you exactly how it fails.
A light paraphrase collapses detector accuracy from 97% to 57% without changing the meaning. That isn't a loophole to exploit — it's proof the score was never evidence.
A detector score is not proof of cheating: a 2025 case study scored human writing 0% and the same text, lightly AI-edited, 100%. Here is the policy that follows.
Text watermarking is the one AI-detection method that actually works — but it needs control of the model, which the person grading your essay never has.
AI detectors clear human text far better than they catch machine text — and the false positives they do produce land on non-native writers up to 98% of the time.
A detector score is not proof. Here is what to say, what to show, and which peer-reviewed studies to cite when you have to defend work you actually wrote.
Two completely different technologies both get called 'AI detection.' One guesses from the finished text; the other reads a signal built in at generation. Knowing which is which explains everything else.
A 2026 experiment gave 214 teachers the same paper. The only thing that changed was the AI-detection number on top — and the grade fell 12 points on a 100-point scale.
AI detectors measure how predictable your writing is — and second-language English is predictable by design. The bias isn't a bug in the tool; it's the metric working exactly as built.
7 cited sources
What the research shows
Four documented failure modes
These are not complaints from angry students. They are findings from published evaluations of detection tools.
Biased against non-native writers
Stanford researchers found GPT detectors systematically misclassify writing by non-native English speakers as machine-generated, because that writing has lower perplexity by nature.
Paraphrasing defeats them
Detection accuracy drops significantly across all tested language models, detectors and tasks once the text is paraphrased — regardless of diversity control codes.
Human text gets flagged
Review studies conclude that most detectors are ineffective at identifying generative-AI text, and that genuinely human writing is regularly detected as AI-generated.
Watermarking is a different approach
Active methods embed a signal at generation time instead of guessing afterwards. They work — but only for platforms that control the model, which is why they do not help your professor.
Common questions
Questions people actually ask
I wrote it myself and the detector says it is AI. How is that possible?
Detectors estimate how predictable your text is. Writing that is clear, plain and grammatically regular scores as predictable — and so does machine output. That is why the same tools disproportionately flag non-native English writers, whose lexical and syntactic variability is lower on average. Low variability is not evidence of cheating.
Should an institution act on a detector score alone?
The literature on detector accuracy does not support it. Published reviews report that tools are inconsistent across models and tasks, that paraphrase defeats them, and that false positives on human text are common. A score is a prompt to have a conversation, not a finding of fact.
How do I prove I wrote something?
Process evidence beats output evidence. Version history, drafts with timestamps, notes and outlines, and the ability to explain your own argument in person are all stronger than anything a classifier says about the finished file.