Does Paraphrasing Beat AI Detectors? Yes — And That's the Problem
Yes, paraphrasing beats AI detectors — and the reason that answer matters is the opposite of the reason most people ask the question. The interesting finding is not that machine text can slip past a checker. It is how little it takes, and what that reveals about the checker itself. When a rewrite that leaves the meaning untouched can swing a detector’s verdict from “definitely AI” to “definitely human,” the number the detector produced was never measuring what anyone thought it was. A score you can flip without changing a single idea is not evidence of authorship. It is a reading of surface texture, dressed up as a judgment about who wrote the words. That is the case this article makes, and the research behind it is unusually clean.
This is a companion argument to the one running through our AI detectors coverage. Elsewhere we’ve shown these tools clear human writing far more reliably than they catch machine writing, and that the false positives they do produce land disproportionately on specific groups. The paraphrasing result is the other half of the same problem, approached from the machine side: not “they wrongly accuse humans,” but “they can be silently talked out of a correct accusation by a change that a reader would barely notice.” Both failures trace to the same root, and the root is not a bug to be patched. It is what passive detection is.
The number that ends the argument
The cleanest experiment comes from Aounon Kumar and colleagues at the University of Maryland, in a 2024 paper with a title that gives away its conclusion: Can AI-Generated Text be Reliably Detected? They took watermarked model output — text a detector caught with 97% accuracy — and ran it through a lightweight neural paraphraser. Detector accuracy fell to 80%. A slightly heavier paraphrase pushed it to 64%, then 57%, and the crucial detail is that the text’s quality barely moved: the perplexity score, a rough proxy for fluency, degraded by only 3.5 points on the first pass (Kumar et al., 2024, p. 6). The meaning survived. The verdict did not.
The collapse is more dramatic for the “zero-shot” detectors that need no training data. Running a T5-based paraphraser over model text dropped the area under the ROC curve of those detectors from 96.5% to 25.2% (Kumar et al., 2024, p. 4). To read that number correctly you need one fact: an AUROC of 50% is a coin flip. The paraphrased detector scored below 50% — meaning it was now worse than guessing, systematically betting the wrong way. Even OpenAI’s own trained RoBERTa-Large detector, which caught the raw model text 100% of the time at a strict error threshold, fell to a 60% catch rate after the same treatment (Kumar et al., 2024, p. 10).
The authors do not frame this as a temporary weakness that better engineering will fix. They frame it as structural. Their results, they write, “indicate the impossibility of developing reliable detectors in practical scenarios — to maintain reliable detection performance, LLMs would have to trade off their performance.” A paraphrase attack works because the space of ways to express a given idea is enormous, and a detector trying to fence off “the AI ways of saying it” is fencing off a region that overlaps almost completely with “the human ways of saying it.” You cannot draw that border without either missing machine text or catching people. That is the same trade-off, viewed from the machine side, that our piece on how AI detectors actually work traces from the human side.
Why a rewrite is enough
To see why the collapse is so easy, it helps to know what a passive detector actually reads. It does not understand your argument. It measures statistical texture — chiefly how predictable your word choices are to a language model, a quantity called perplexity. Machine text tends to sit in a low, smooth band of predictability, because that is what the models are trained to produce. Human text is usually bumpier. The detector draws a line and calls everything below it “AI.”
A paraphrase moves the text across that line without moving its content. Swap the model’s most-likely phrasing for a second- or third-choice phrasing and the perplexity rises; the statistical fingerprint the detector was keyed to smudges, and the classifier loses its grip. Nothing about the ideas changed. This is why even research teams building new detectors concede the point in their own methods sections. A 2026 paper in Scientific Reports proposing a fusion detector notes plainly that “semantic-only detectors tend to experience performance degradation under paraphrasing conditions,” which is why they had to bolt on surface-level features as a workaround (Chen et al., 2026, p. 11). The same authors put their finger on the deeper reason the whole enterprise is losing ground: there is an “increasing convergence of stylistic characteristics between advanced language model outputs and carefully edited human writing” (Chen et al., 2026, p. 12). As the machines get more human and the humans use the machines, the two distributions merge — and a border between merging distributions is a border that catches the wrong things.
The pattern is not one lab’s result. A 2025 literature review synthesizing dozens of studies records the same finding across independent groups: it reports that paraphrasing “is found to drastically reduce detection accuracy while keeping input semantics,” and that detector accuracy “drops significantly after using recursive paraphrasing and spoofing” (Gotoman et al., 2025, p. 4). The reviewers also surface a companion number that should unsettle anyone who trusts a human backstop instead: in one study, people could tell a genuine news story from a machine-written one only 52% of the time (Gotoman et al., 2025, p. 5) — again, a coin flip. Neither the tool nor the unaided reader has a reliable grip on the distinction.
The detectors’ own makers say the same
You do not have to take this from academics. The vendors’ own disclosures point the same way once you read past the marketing number. Copyleaks advertises a 99% accuracy rate (Elkhatat et al., 2023, p. 3) — the kind of figure that ends up in a disciplinary email. But when OpenAI shipped its own classifier, it told users to treat the result “as supplementary information rather than relying on them exclusively for determining AI-generated content” (Elkhatat et al., 2023, p. 3), then discontinued the tool entirely for low accuracy. The company that built the model could not reliably detect the model.
Turnitin, the tool most students actually meet, now carries a similar caution in its own documentation: it warns that its detector “may misidentify” human-written, AI-generated, and AI-paraphrased text, and that scores “should not be used as the sole basis” for penalizing a student (Turnitin, 2025). Institutions have started acting on that admission rather than the accuracy claim. Vanderbilt disabled Turnitin’s AI detector in 2023; Curtin University announced it would switch the feature off from January 2026, keeping ordinary plagiarism checks but dropping AI detection over exactly the reliability concerns described here (EdTech Innovation Hub, 2026). The vendors and the universities are converging on the position the research reached first: the score is not the thing to hang a decision on.
There is an arms-race response, of course. Turnitin rolled out a “bypasser” model in late 2025 aimed at text that has been run through rewriting tools, and the fusion detectors in the literature add stylistic features precisely to survive a paraphrase. But note what the escalation concedes. Each new layer is an admission that the previous layer was defeated by a rewrite — and each new layer defines a new statistical border that a different rewrite will, in turn, cross. The Maryland team’s argument is that this loop has no fixed point. You are not converging on a reliable detector; you are chasing a distribution that keeps merging with human writing.
What the score actually proves — in both directions
Here is the symmetry that makes the paraphrasing result damning rather than merely inconvenient. A detector that can be talked out of a correct verdict by a harmless rewrite is, by the same mechanism, a detector that can be talked into a wrong one by nothing at all. If second- or third-choice phrasing reads as “human,” then any writer whose natural style already avoids the model’s favorite phrasings — a careful editor, a formulaic technical writer, a second-language author writing deliberately plain prose — reads as suspicious for free. That is not a hypothetical. It is the documented mechanism behind the false-positive rates that fall on non-native English writers, whose prose is statistically calm by design.
So the two failure modes are one failure mode. The tool is measuring predictability and reporting it as authorship. Predictability is cheap to raise (paraphrase the machine text) and naturally low for whole populations of real writers (flag the humans). A measurement with those properties cannot function as evidence in a proceeding where a person’s grade or reputation is at stake — because the quantity it measures is only loosely, and manipulably, correlated with the thing it claims to detect.
The older literature said this before the paraphrase experiments sharpened it. Weber-Wulff and a large European team, testing the major detectors across 750-plus cases, concluded the tools were in one evaluation “no better than random classifiers,” and specifically that they “found it challenging to detect a piece of human-written text that was rewritten by ChatGPT” (Weber-Wulff et al., 2023, p. 7). In another test they cite, adding a “beat detection” instruction pushed the number of tools fooled from 10 of 16 up to 12 of 16 (Weber-Wulff et al., 2023, p. 7). And in a wrinkle that dissolves the naive fix, they note that ChatGPT “does not produce original texts after paraphrasing” — meaning a paraphrase good enough to fool an AI detector can still trip an ordinary plagiarism checker, because the paraphrased text stays too close to its source (Weber-Wulff et al., 2023, p. 5). The tools do not even fail cleanly.
The practical reading
None of this is a how-to, and it is worth saying why. The point of the paraphrasing evidence is not that a rewrite is a clever way to cheat. It is that the ease of the rewrite retroactively voids the number. If the verdict was that fragile all along, then a “97% AI” reading on an unedited essay never carried the weight an accuser wants to put on it — because an identical essay, trivially reworded, would have read as human, and the tool would have been just as confident either way. Confidence that survives being wrong is not confidence worth acting on.
For anyone who grades, adjudicates, or gets accused, the operational consequence is narrow and firm. A detector percentage is a probability about statistical texture, not a finding about who wrote the text, and it can move in either direction for reasons that have nothing to do with authorship. It is a starting point for a conversation with the writer, never the end of one — a position the accused-student literature spells out in what happens when you’re falsely accused of using AI, and one the detector score itself quietly undermines even before a human weighs in. The vendors now say the same in their fine print, and the universities are switching the feature off. The paraphrase experiment is simply the cleanest proof of what all of them have concluded: the score was never the evidence.
Sources
Every factual claim above is tied to a source you can open and check, with page numbers wherever the source has them.
- 1. Aounon Kumar, Sriram Balasubramanian, Vinu Sankar Sadasivan, Wenxiao Wang, Soheil Feizi, Can AI-Generated Text be Reliably Detected? , arXiv preprint, 2024 , pp. 4, 6, 10. link
- 2. Debora Weber-Wulff, Alla Anohina-Naumeca, Sonja Bjelobaba, Tomáš Foltýnek, Jean Guerrero-Dib, Olumide Popoola, Petr Šigut, Lorna Waddington, Testing of detection tools for AI-generated text , International Journal for Educational Integrity, 2023 , pp. 5, 7. 10.1007/s40979-023-00146-z
- 3. Jezreel Edriene J. Gotoman, Harenz Lloyd T. Luna, John Carlo S. Sangria, Cereneo S. Santiago Jr, Danel Dave Barbuco, Accuracy and Reliability of AI-Generated Text Detection Tools: A Literature Review , American Journal of Interdisciplinary Research and Bibliometrics, 2025 , pp. 4, 5. 10.54536/ajirb.v4i1.3795
- 4. Hao Chen, Huancheng Chen, Boyu Hu, Jihua Chai, Hui Zhang, Xitong Wang, Jupeng Wang, Research on ChatGPT generated text detection model based on phonetic feature extraction and semantic features , Scientific Reports, 2026 , pp. 11, 12. 10.1038/s41598-026-49952-8
- 5. Ahmed M. Elkhatat, Khaled Elsaid, Saeed Al-Meer, Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text , International Journal for Educational Integrity, 2023 , p. 3. 10.1007/s40979-023-00140-5
- 6. Turnitin, Understanding the false positive rate for sentences of our AI writing detection capability , Turnitin, 2025 . link
- 7. EdTech Innovation Hub, Curtin University to disable Turnitin AI detection tool in 2026 amid debate over reliability , EdTech Innovation Hub, 2026 . link