AI Writing Tips

When Your Teacher Cannot Tell Any More

· 11 cited sources

For twenty years, a teacher’s read on a suspicious essay was worth something. You knew your students’ voices, you noticed when a paragraph outran a kid’s usual range, and you could act on that instinct. That instinct is now worthless, and the research says so with numbers. When Johanna Fleckenstein and colleagues asked 89 pre-service and 200 experienced teachers to pick the ChatGPT essays out of a pile of student writing, the teachers could not do it — and, worse, both groups were overconfident in judgments that were no better than chance (Fleckenstein et al., 2024, p. 1). The eye has failed. The software that was supposed to replace the eye has failed too. This piece is written for the people that failure lands on hardest: the language instructors and course leaders who not only grade the essays but also decide which detection tools their department pays for. If you are signing that purchase order, the honest state of the evidence should change what you sign.

Start with the finding that the whole thing turns on, because it is bleaker than the marketing admits. This is not a story about detectors being 80% accurate instead of 99%. It is a story about two separate systems of judgment — the human teacher and the machine detector — both landing near a coin flip on exactly the writing that matters, and about a shortcut teachers reach for when they cannot tell, which happens to punish the most proficient non-native writers in the room.

The human read is already gone

Fleckenstein’s study is the one to sit with, because it tested the thing every teacher privately believes they can still do. The design was simple: mix AI-generated essays into genuine student work, hand them to teachers, and see who spots what. Neither novices nor veterans could identify the ChatGPT texts among the student-written ones, and text quality did not rescue them — a strong essay was as likely to be misjudged as a weak one (Fleckenstein et al., 2024, p. 1). The detail that should worry anyone who grades is the confidence gap. The teachers were sure. Being experienced did not make them right; it made them certain while being wrong. That combination — low accuracy, high confidence — is precisely the profile that produces a false accusation, because the person making the call never doubts it enough to check.

There is a reason the read collapsed, and it is not that teachers got worse. The machines got better. A 2026 detection study opens by conceding the ground plainly: AI text “has become increasingly fluent and contextually coherent, often approaching the quality of human writing,” which is exactly why identifying its origin has turned into an urgent research problem rather than a solved one (Chen et al., 2026, p. 1). The gap a teacher used to hear — the tin ear, the phrase no student would write — has been engineered out. You are not slipping. The target moved.

The shortcut that fails twice

When the direct read stops working, teachers fall back on a heuristic, and this is where the story bends toward the non-native writer. Katarzyna Alexander, Christine Savvidou and Chris Alexander watched six ESL lecturers try to sort AI essays from human ones and documented the rule the teachers actually used. They leaned on a deficit model of assessment: error was read as the signature of a real second-language student, while “high levels of technical and grammatical accuracy and sophisticated language use” were treated as indicators of AI (Alexander et al., 2023, p. 25). Read that back slowly. The cleaner the English, the more likely the teacher was to call it a machine.

That heuristic misfires in the one direction that does the most damage. A proficient non-native writer — someone who drilled the grammar, checked the article, built the sentence on a pattern they trust — produces exactly the clean, accurate prose the deficit model files under “AI.” The same students who worked hardest to write correct English are the ones the fallback rule flags. It is the mirror image of the machine-side problem we trace in why detectors flag non-native English writers: there, a statistical model reads low-perplexity L2 fluency as machine output; here, a human teacher reads high grammatical accuracy as machine output. Two different mechanisms, one victim. The writer whose English is good is guilty on both counts.

The same study caught a second blind spot with teeth. The lecturers paid little attention to “the veracity of facts and references generated by ChatGPT” (Alexander et al., 2023, p. 25) — they were chasing style and ignoring substance. That is backwards, and it points to the one tell that still works. Style is now forgeable; a fabricated citation is not. When a model invents a plausible-looking source, the invention is checkable in a way that prose rhythm never was — you can look for the paper and find nothing there. A teacher who runs a suspect bibliography through a tool like Cytado’s citation checker to see whether the cited works actually exist is testing a fact, not a vibe, and facts are where AI text is still weakest.

The software cannot bail you out

The obvious move — if I can’t tell, let the detector tell — is the one the department is being sold, and it is the one the evidence dismantles most thoroughly. Debora Weber-Wulff’s team ran one of the field’s most comprehensive tests: 14 tools, including Turnitin, against an original document set. Their conclusion is not hedged. The available detectors “are neither accurate nor reliable” and carry a built-in bias toward classifying text as human-written rather than catching the AI (Weber-Wulff et al., 2023, p. 1). In one telling case they cite, a plagiarism checker recognized nearly all fifty ChatGPT-generated scientific abstracts as fully original (Weber-Wulff et al., 2023, p. 6). The tool marketed as your safety net was, in that test, blind to almost everything it was pointed at.

It gets more specific, and worse for the non-native case. Weber-Wulff’s group warns directly that using machine translation such as Google Translate or DeepL “can lead to a higher number of false positives, leaving L2 students (and researchers) at risk of being falsely accused” (Weber-Wulff et al., 2023, p. 26). So the second-language student who drafts in their first language and translates — a legitimate, common workflow, the subject of publishing in English from a non-dominant language — is precisely the one the detector is most likely to finger. The tool’s errors are not random. They pool on the people least able to absorb them.

The reliability picture across the wider literature is the same shape. A 2025 review of the field concluded that authorities “should only partially trust these tools, for they are imperfect,” and flagged the fairness problem for non-native writers as a live concern (Gotoman et al., 2025, pp. 1, 7). A 2026 evaluation drove the point home with a blunt statistic from Chaka’s testing: of 30 AI detectors run over student essays, only two accurately identified human-written text (Shamsi et al., 2026, p. 127) — and the same body of work found detection accuracy drops sharply on non-English languages and collapses under simple adversarial edits (Shamsi et al., 2026, p. 128). If a department is choosing between “trust the tool” and “trust the teacher,” the data says both readers are unreliable, and the tool is unreliable in a way that is systematically unfair.

The trend is toward less detectable, not more

The temptation is to treat this as a temporary lag — buy the tool now, wait for version four. The theory says otherwise. Aounon Kumar and colleagues proved a formal bound on detection: as a language model’s output distribution approaches the distribution of human text, the accuracy ceiling of any possible detector falls toward a coin flip, and this is a mathematical limit, not an engineering bug (Kumar et al., 2024, p. 5). Their empirical work shows the direction of travel — as models grow more powerful, their outputs become “more indistinguishable from human text, making them harder to detect” (Kumar et al., 2024, p. 17). Detection does not get easier as the technology matures. It gets structurally harder.

Two more reviews confirm the trajectory from different angles. A 2025 survey describes the classification boundary between human and AI text as “blurred and dynamic,” in contrast to the “clear and well-separated” boundaries of ordinary text classification (Xiang et al., 2025, p. 4) — you are trying to draw a line through a cloud that keeps moving. And the earlier reliability review found detectors are more effective on older generative models and lose accuracy as newer ones arrive (Gotoman et al., 2025, p. 7). Every model release quietly degrades the tool the department bought last year. This is the same losing race we lay out in full in are AI detectors accurate: the thing you are detecting improves faster than the detector.

There is even a nastier failure mode the detection literature keeps surfacing: the plainer the writing, the more human it is assumed to be. As Halah Mohammed Abbas puts it, “simple and easily understandable content is less likely to be perceived as artificial” (Abbas, 2025, p. 13) — a rule that inverts reality, since simple, careful, high-frequency vocabulary is exactly what both good AI output and good L2 writing produce. Abbas is explicit that datasets rarely capture the gap between native and non-native fluency, and that the resulting bias “may wrongly accuse marginalized groups of plagiarism” (Abbas, 2025, p. 14). The tool does not just miss the machine text; it can convict the human.

What this means for the person signing the purchase order

Here is where the audience matters. Language instructors are not passive recipients of this technology — in most departments they are the ones who evaluate and recommend the assessment tools the institution buys. That makes the collapse of detection a procurement question, not just a grading one, and the evidence points to a specific set of decisions.

First, stop buying certainty you cannot get. A detector that returns a confident percentage is selling the same overconfidence Fleckenstein measured in human teachers (2024, p. 1), just dressed as arithmetic. The moment a tool’s output is treated as proof rather than a hint, it will produce false accusations, and those accusations will land disproportionately on translated and high-proficiency L2 work (Weber-Wulff et al., 2023, p. 26). Weber-Wulff’s team is unambiguous about the correct status of these reports: they can give faculty a hint that something may have happened, but they “cannot be used as the only basis for reporting students for cheating” (Weber-Wulff et al., 2023, p. 26). Any procurement policy that lets a score trigger a misconduct case on its own is buying a lawsuit.

Institutions have already started acting on this. Vanderbilt disabled Turnitin’s AI detector in August 2023, reasoning that at the vendor’s claimed 1% false-positive rate, the 75,000 papers the university processed in a year would have produced roughly 750 wrongful flags — and noting that detectors are more likely to mislabel non-native English writers (Coley / Vanderbilt, 2023). That is not a fringe position anymore; it is a documented institutional judgment that the tool’s costs exceed its value.

Second, spend the budget where the signal still lives. Since 2005, the reliable defense against machine-generated academic text has not been a magic classifier but provenance — traditional plagiarism detectors were already unable to catch GPT-3 output, and the field’s own answer was to look at process and history rather than the finished string (Santra & Majhi, 2023, p. 177). For a language classroom that means assessment redesign over surveillance: drafts, revision history, in-class writing, oral defense of a submitted text, and rubrics that reward the reasoning a model cannot fake. It also means teaching students what disclosure looks like, the ground we cover in what AI feedback actually improves in L2 writing — the tool is legitimate for grammar and flow, and pretending otherwise just drives honest use underground.

Third, protect the writers your current heuristics endanger. The deficit model Alexander documented (2023, p. 25) is not a personal failing; it is what any of us reach for when the direct read fails. Naming it is the fix. If your department’s unwritten rule is “too clean to be a real student,” that rule is a bias against the non-native writers who worked hardest, and it belongs in professional-development sessions the way any other assessment bias does. The broader cost of these misfires — to the students on the receiving end — is the subject of being falsely accused of using AI, and the whole non-native cluster of writing in a second language sits under that shadow.

The uncomfortable truth is that “can you tell it’s AI?” is the wrong question to build a policy on, because the answer is no — not by eye, not by software, not next year. The question that still has a good answer is “can this student show me how they got here, and does the substance hold up?” That is a question about drafts, dialogue, and checkable facts, and it does not care whether a teacher’s instinct has quietly stopped working. It has. The sooner a department stops buying tools that promise otherwise, the sooner it can spend on the things that still tell the truth.

Sources

Every factual claim above is tied to a source you can open and check, with page numbers wherever the source has them.

  1. 1. Johanna Fleckenstein, Jennifer Meyer, Thorben Jansen, Stefan D. Keller, Olaf Köller, Jens Möller, Do teachers spot AI? Evaluating the detectability of AI-generated texts among student essays , Computers and Education: Artificial Intelligence, 2024 , p. 1. 10.1016/j.caeai.2024.100209
  2. 2. Katarzyna Alexander, Christine Savvidou, Chris Alexander, Who Wrote This Essay? Detecting AI-Generated Writing in Second Language Education in Higher Education , Teaching English with Technology, 2023 , p. 25. 10.56297/BUKA4060/XHLD5365
  3. 3. Debora Weber-Wulff, Alla Anohina-Naumeca, Sonja Bjelobaba, Tomáš Foltýnek, Jean Guerrero-Dib, Olumide Popoola, Petr Šigut, Lorna Waddington, Testing of detection tools for AI-generated text , International Journal for Educational Integrity, 2023 , pp. 1, 6, 26. 10.1007/s40979-023-00146-z
  4. 4. Jezreel Edriene J. Gotoman, Harenz Lloyd T. Luna, John Carlo S. Sangria, Cereneo S. Santiago Jr, Danel Dave Barbuco, Accuracy and Reliability of AI-Generated Text Detection Tools: A Literature Review , American Journal of Interdisciplinary Research and Bibliometrics, 2025 , pp. 1, 7. 10.54536/ajirb.v4i1.3795
  5. 5. Aounon Kumar, Sriram Balasubramanian, Vinu Sankar Sadasivan, Wenxiao Wang, Soheil Feizi, Can AI-Generated Text be Reliably Detected? , arXiv preprint, 2024 , pp. 5, 17. link
  6. 6. Lingyun Xiang, Nian Li, Yuling Liu, Jiayong Hu, AI-Generated Text Detection: A Comprehensive Review of Active and Passive Approaches , Computers, Materials & Continua, 2025 , p. 4. 10.32604/cmc.2025.073347
  7. 7. Hongyi Chen, Han Chen, Bo Hu, Jie Chai, Hui Zhang, Xin Wang, Jun Wang, Research on ChatGPT generated text detection model based on phonetic feature extraction and semantic features , Scientific Reports, 2026 , p. 1. 10.1038/s41598-026-49952-8
  8. 8. Amrollah Shamsi, Ting Wang, Maryam Amraei, Narayanaswamy Vasantha Raju, Evaluating AI Text Detection Tools for Distinguishing Human-Written from AI-Generated Abstracts in Persian-Language Journals of Library and Information Science , Acta Informatica Pragensia, 2026 , pp. 127, 128. 10.18267/j.aip.293
  9. 9. Halah Mohammed Abbas, A Novel Approach to Automated Detection of AI-Generated Text , Journal of Al-Qadisiyah for Computer Science and Mathematics, 2025 , pp. 13, 14. 10.29304/jqcsm.2025.17.11958
  10. 10. Patit Paban Santra, Debasis Majhi, Scholarly Communication and Machine-Generated Text: Is it Finally AI vs AI in Plagiarism Detection? , Journal of Information and Knowledge (SRELS), 2023 , p. 177. 10.17821/srels/2023/v60i3/171028
  11. 11. Michael Coley / Vanderbilt University (Brightspace), Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector , Vanderbilt University, 2023 . link
More on Writing English as a Non-Native
Built by Sitario.com