AI Writing Tips

Should Schools Use AI Detectors at All? What a Score Can and Can't Prove

· 10 cited sources

Stop asking whether AI detectors should be banned. That framing loses the argument before it starts, because it invites a yes-or-no fight over a tool and skips the only question that actually decides policy: what is a detector score logically licensed to support? Answer that honestly and the policy writes itself. A detection percentage is not a measurement of who wrote a text — it is a measurement of how statistically ordinary the prose looks. Those are different things, and the gap between them is where every wrongful accusation lives. Once a committee understands what the number is, the range of defensible uses collapses to a narrow one: a detector output can open a conversation, and it can never close a case.

What the number actually measures

The cleanest demonstration of the gap comes from a 2025 case study by Anna Kamińska at the University of Silesia. She wrote a biography of a fictional person herself, then edited two copies of that human-authored text with ChatGPT and with Claude, giving each model precise instructions that limited it to a purely editorial role — fix punctuation, smooth a clumsy sentence, find a better synonym — with no new facts, arguments, or claims added. She then ran all three versions through the detector ZeroGPT. The original scored a 0% probability of being AI-generated. The ChatGPT-polished version scored 96.7%, and the Claude-polished version scored 100% (Kamińska, 2025, p. 76). Same author. Same ideas. Same intellectual content. The detector’s verdict swung from “certainly human” to “certainly machine” on the basis of nothing but stylistic tidiness.

Kamińska’s conclusion is the sentence every academic-integrity committee should paste at the top of its policy: detectors “identify stylistic features characteristic of well-edited texts, rather than the actual origin of the intellectual contribution” (Kamińska, 2025, p. 76). The whole intellectual-property question in academic honesty is about origin — did this person’s mind produce this argument? The detector cannot see origin. It sees surface regularity, and it mistakes a clean edit for a stolen one. A student who runs a genuinely self-written essay through Grammarly, or through a spell-checker aggressive enough to reshape phrasing, is walking straight into that trap (Kamińska, 2025, p. 63).

This is why the accuracy debate around AI detectors — is a given tool 60% right, 90%, a coin flip? — is necessary but not sufficient. Even a tool that were perfectly calibrated on its own terms would still be reporting the wrong quantity. It would be answering “does this read like machine-typical text?” while the teacher hears “did a machine write this?” No accuracy figure fixes a category error.

The numbers that show the floor is real

If Kamińska’s study were an outlier, you could dismiss it. It is not — the failure to clear genuine human writing shows up across languages and tool sets. A 2026 evaluation of Persian-language academic abstracts found that when scoring texts written by humans, ZeroGPT correctly recognized them as human only 1.43% of the time and GPTZero only 2.14% of the time (Shamsi et al., 2026, p. 131). Overall, across human and machine samples combined, ZeroGPT’s accuracy landed at 53.89% and GPTZero’s at 28.34% — the latter worse than a coin flip (Shamsi et al., 2026, p. 131). The same paper cites Chaka’s 2024 test of thirty detection tools on student essays, in which only two of the thirty accurately identified human-written text (Shamsi et al., 2026, p. 127). Thirty tools, twenty-eight of them unable to reliably clear a real student.

The spread between tools is enormous, which matters for any policy that names a specific product. In one empirical comparison, Originality reached 97.0% accuracy while GPTKit managed only 55.29%, with GPTZero, Sapling, Writer and Zylalab scattered between roughly 63% and 69% (Akram, 2023). A 2025 literature review found the same instability from a different pile of studies, reporting that paid detectors average around 87% accuracy and free ones around 77% — and, more damningly, quoting Weber-Wulff’s team that current detection technologies are “neither accurate nor dependable,” and Sadasivan’s that the tools are simply “unreliable” for the task (Gotoman et al., 2025, pp. 5–6). “The detector said 87%” is therefore not one fact but a family of very different facts, depending on which tool, which text, which language, and which measurement convention produced it.

None of these figures come with the one thing a disciplinary process needs: a stable, known error rate for the specific student in front of you. And the errors are not random. As covered in detail in our piece on why detectors flag non-native English writers, the false positives concentrate on second-language writers, whose narrower vocabulary reads to a perplexity-based detector exactly like a language model’s calm, predictable output. The tool is most confident where it is most likely to be wrong, and it is most likely to be wrong against the students least able to fund an appeal.

Why “just one input” is not a real safeguard

Committees reassure themselves with a phrase: the detector is “just one input” a teacher weighs alongside their own judgment. The evidence says that reassurance is false, because the number does not sit quietly beside the teacher’s judgment — it reshapes it. A 2026 experiment on automation bias gave 214 teachers the same student paper and varied only the detection score printed on top; the group shown a high score marked the identical essay down by roughly twelve points on a hundred-point scale, and rated its originality and logic lower too. The full breakdown is in our write-up on how the detector score itself lowers your grade, but the implication for policy is blunt: you cannot treat a detector as one neutral input when its mere presence bends every other input around it. By the time a human “independently” reviews the work, the number has already told them what to see.

That is the mechanism behind the stream of students falsely accused of using AI on the strength of a percentage and little else. The tool does not merely risk a wrong verdict; it manufactures the suspicion that then goes looking for confirmation.

What you may legitimately conclude — and what follows

Put the pieces together and the permissible inference is small and specific. A high detector score licenses exactly one conclusion: this text has surface features that some tool associates with machine-typical writing. It does not license “the student used AI.” It does not license “the student cheated.” It cannot distinguish generation from editing, a second language from a second author, or a formulaic genre — lab reports, legal briefs, structured abstracts — from a language model. Kamińska’s zero-to-one-hundred swing on unchanged authorship is the proof that the leap from score to accusation is unsupported (Kamińska, 2025, p. 76).

From that single honest premise, a workable policy follows without anyone having to argue about a ban:

  • A detector output may open a conversation; it may never be the evidence in a case. This is precisely where institutions have landed as they retreat from detection. Johns Hopkins moved its detector to “advisory only” — an instructor may run a submission, but the result can only start a discussion, never found a formal charge. That is the correct altitude for the tool.
  • No accusation rests on a percentage alone. The MLA-CCCC Joint Task Force on Writing and AI is explicit that faculty should not rely on an automated report as their only evidence, warning that false accusations “disproportionately affect marginalized groups” and are often driven by culture and unconscious native-speakerism rather than misconduct (MLA-CCCC, 2024). Corroboration means process evidence a detector cannot fake: draft history, version logs, a conversation in which the student can talk through their own argument.
  • Grade before you look. The automation-bias research points to a simple procedural fix — complete the full quality evaluation before opening any detection report, so the number cannot anchor a judgment that has not yet been formed.
  • Design the assignment so detection is beside the point. Process-visible work — staged drafts, in-class writing, oral defenses, reflections tied to the student’s own sources — deters misuse far more reliably than a post-hoc scan, and it does so without accusing anyone.

Notice that none of this required deciding whether detectors are “good” or “bad.” It required only taking the score at its true weight.

The institutions are already voting

This is not a fringe position; it is where the sector is moving in real time. Vanderbilt University disabled Turnitin’s AI detector back in 2023, noting that of the 75,000 papers it had submitted the previous year, roughly 750 could have been wrongly flagged had the tool been running (Vanderbilt University, 2023). Curtin University announced it would switch off Turnitin’s AI writing detection across all campuses from 1 January 2026, framing the move as a way to keep assessment fair and trustworthy while the reliability debate continues (Loosley, 2026). More than fifty universities across the US, UK, Canada and Australia — Yale and Johns Hopkins among them — have now disabled or restricted the same feature.

The pressure is not going down, which is exactly why a defensible policy matters now rather than later. Teacher adoption is climbing: a Center for Democracy & Technology survey found the share of secondary-school teachers using AI-detection tools jumped thirty percentage points in a single year, to 68% in 2023–24, while student discipline tied to suspected AI use rose from 48% to 64% (Center for Democracy & Technology, 2024). On the higher-education side, Turnitin reported that between October 2025 and February 2026 an average of 14.8% of English-language submissions to its detector came back with 80% or more of the text flagged as AI-written (Turnitin, 2026). Whatever fraction of that 14.8% are false positives — and the research above says it is not small — each one now lands in a disciplinary pipeline. A policy that treats the flag as a prompt rather than a proof is the difference between a conversation and a wrongful conviction.

The honest answer

So — should schools use AI detectors at all? The useful answer is not yes or no but a boundary. Use them, if you must, for the one thing they can support: as a low-stakes signal that a piece of writing might be worth a closer, human look. Never use them for the thing they cannot support: as evidence that a specific person did a specific wrong. The distinction is not a compromise between two camps. It is just the difference between what a detector measures and what an accusation claims — and the research from Kamińska, Shamsi, Weber-Wulff, and the 214 teachers in the automation-bias study all converge on the same line. The number describes a texture of prose. It does not know who wrote it. Any policy that forgets that will keep punishing the wrong people, and the courts and the disabled Turnitin dashboards are already the receipts.

Sources

Every factual claim above is tied to a source you can open and check, with page numbers wherever the source has them.

  1. 1. Anna Małgorzata Kamińska, Czy ten artykuł napisała sztuczna inteligencja? Studium przypadku „wykrywania plagiatu” , Zagadnienia Informacji Naukowej (ZIN), 2025 , pp. 61, 63, 76. 10.36702/zin2025.02.04
  2. 2. Amrollah Shamsi, Ting Wang, Maryam Amraei, Narayanaswamy Vasantha Raju, Evaluating AI Text Detection Tools for Distinguishing Human-Written from AI-Generated Abstracts in Persian-Language Journals of Library and Information Science , Acta Informatica Pragensia, 2026 , pp. 127, 131. 10.18267/j.aip.293
  3. 3. Jezreel Edriene J. Gotoman, Harenz Lloyd T. Luna, John Carlo S. Sangria, Cereneo S. Santiago Jr, Danel Dave Barbuco, Accuracy and Reliability of AI-Generated Text Detection Tools: A Literature Review , American Journal of Interdisciplinary Research and Bibliometrics, 2025 , pp. 5, 6. 10.54536/ajirb.v4i1.3795
  4. 4. N. Pireci Sejdiu, S. Grassini, S. Sejdiu, Generative AI in education: Process-aware pedagogy, assessment integrity, and institutional governance , Open Research Europe, 2026 , pp. 3, 9. 10.12688/openreseurope.23343.1
  5. 5. Arslan Akram, An Empirical Study of AI-Generated Text Detection Tools , Advances in Machine Learning & Artificial Intelligence, 2023 . 10.33140/amlai.04.02.03
  6. 6. MLA-CCCC Joint Task Force on Writing and AI, Generative AI and Policy Development: Guidance from the MLA-CCCC Task Force , Modern Language Association & Conference on College Composition and Communication, 2024 . link
  7. 7. Chris Loosley, Curtin University to disable Turnitin AI detection tool in 2026 as debate over reliability continues , EdTech Innovation Hub, 2026 . link
  8. 8. Vanderbilt University (Center for Teaching), Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector , Vanderbilt University, 2023 . link
  9. 9. Turnitin, Turnitin Data Shows Transparency About AI Use Benefits Students and Educators , Turnitin, 2026 . link
  10. 10. Center for Democracy & Technology (reported by K-12 Dive), As teacher use of AI detection grows, discipline guidance a mixed bag , K-12 Dive, 2024 . link
More on AI Detectors
Built by Sitario.com