Falsely Accused of Using AI: The Evidence That Wins an Appeal
If a professor emails to say your essay came back “98% AI” and you know you wrote every word, the worst thing you can do is answer from panic. A detector percentage feels like proof. It is not proof. It is a probability estimate from a tool whose own field cannot agree that it works — and in a growing number of hearings and courtrooms, that distinction is exactly what clears the accused. This is not a page about calming down. It is a page about what to say, what to show, and which studies to put in front of a committee so that a number stops being treated as a verdict.
Start with the one fact that reframes the whole conversation: the burden is not on you to prove innocence. In academic-integrity procedures the institution has to demonstrate misconduct, and a single automated score is not that demonstration. Your job is not to prove a negative. Your job is to make the committee see that their evidence is one unreliable signal, and that you have real evidence they don’t. Both halves of that are winnable.
First moves — before you write a single reply
Do these in the first hour, in order:
- Do not apologize “just in case.” A hedged “I’m sorry if it looked that way” reads as a confession in a transcript. You did not do the thing; do not perform contrition for it.
- Do not touch the document. Do not tidy it, re-save it, or “clean up” the file. Editing the artifact now can destroy the timestamped history that will later exonerate you.
- Ask, in writing, for the specific evidence. Under the Family Educational Rights and Privacy Act (FERPA), a US student generally has the right to see the records used against them. Request the exact detector report, the score, and which passages were flagged. Getting the accusation reduced to “one tool said a number” is half the battle.
- Preserve your trail immediately. Export your draft version history, notes, outlines, and any messages where you discussed the assignment. That record is the thing detectors cannot rebut.
Everything after this depends on those four steps. The evidence you gather in hour one is worth more than anything you argue in week three.
The evidence that actually wins: your process, not your prose
Detectors judge the finished text. They cannot see how it was made — and how it was made is where a real writer’s defense lives. Cloud editors keep a timestamped, attributed record of every edit, and that record has already cleared students. At UC Davis, William Quarterman was exonerated after producing several hours of time-stamped Google Docs revision history for a flagged exam essay; a second UC Davis student, Louise Stivers, was cleared the same way. A verdict built on “the output looks like AI” collapses against two hundred incremental saves showing a human building the output sentence by sentence.
Assemble a process file and hand it over as one package:
- Version history from Google Docs, Word (with AutoSave/OneDrive history), or Overleaf — the single strongest exhibit.
- Earlier drafts and outlines, ideally with their own timestamps.
- Prior graded work in your own voice, to show the flagged piece matches how you always write.
- Your sources, and proof they’re real. Fabricated references are one of the most reliable fingerprints of machine-generated text, so a bibliography where every entry resolves to a genuine, correctly cited work is affirmative evidence that a person did the reading. It is worth running your reference list through a free checker such as cytado.com’s bibliography check and printing the result — a clean pass says the citations aren’t hallucinated, which is precisely what an AI-authored draft cannot show.
If your institution offers an oral defense or “viva,” take it. Being able to explain your argument, your choices, and your sources in real time is something no detector output can survive.
The studies to cite — and the exact numbers to quote
This is the part most students skip and shouldn’t. A committee that treats a detector as authoritative is usually unaware of the peer-reviewed record. Put the numbers in front of them. Every figure below comes from a published study you can name in an appeal.
The tools are formally unreliable — this is the field’s own conclusion, not a complaint. A 2025 literature review synthesizing dozens of evaluations states plainly that AI-generated-text detectors are “generally proven” to be unreliable and have “not reached a level in which faith can be put in the results” (Gotoman et al., 2025, p. 7). The same review records how wide the accuracy spread is across tools — from a low of 55.29% to a high of 97.09% depending on which detector you use (Gotoman et al., 2025, p. 4). Two tools can look at the same paragraph and disagree completely. A signal that inconsistent is not a basis for punishment.
The most rigorous head-to-head test quantified the false-accusation risk directly. Debora Weber-Wulff and a European research team ran 54 test cases through fourteen detectors, 756 tests in all (Weber-Wulff et al., 2023, p. 11). They then computed, for each tool, the likelihood it would falsely accuse a student. Six of the fourteen tools produced false positives on genuine human text, and the risk “increased dramatically for machine-translated texts” — the average false-accusation ratio jumped from 2.4% on original human writing to 11.1% once the human text had been translated, with GPTZero alone posting a 50% false-positive rate in that condition (Weber-Wulff et al., 2023, p. 19). If your work involved translating your own thoughts from another language, that finding is about you specifically.
A five-minute rewrite breaks the tools — which means their scores mean nothing consistent. Vinu Sadasivan and colleagues showed that running text through an ordinary paraphraser dropped a leading detector’s accuracy from 97% to 80%, and pushed a zero-shot detector’s discrimination score (AUROC) from 96.5% down to 59.8% — where 50% is the score of pure random guessing (Kumar et al., 2024, pp. 4, 10). They also proved the failure runs toward the innocent: by feeding paraphrased human passages into a “retrieval” detector’s database, they made it falsely flag all 100 human passages as AI-generated, and warned that an adversary could use exactly this to “falsely accuse these humans of plagiarism” (Kumar et al., 2024, p. 19). Their conclusion is a sentence worth quoting to any administrator: “a detector with a high false positive rate will cause more harm than good in society” (Kumar et al., 2024, p. 6).
The errors are not random — they concentrate on specific writers. Detectors lean on “perplexity,” a measure of how statistically predictable your word choices are, and second-language writers use narrower, more predictable vocabulary by nature. A Stanford team found seven major detectors falsely flagged non-native English essays at rates above 60%, while barely erring on native-speaker essays (Liang et al., 2023, p. 2). If you are a non-native English writer, the tool may have penalized you for your second language, not your honesty — a bias covered in depth in our piece on why AI detectors punish non-native writers. And detectors let AI slip through in the other direction: Ahmed Elkhatat’s team watched tools flag genuine human samples as false positives while repeatedly passing GPT-4 text off as human (Elkhatat et al., 2023, p. 9).
Finally, the deepest point — an accusation of AI use is structurally unprovable from text alone. A 2026 analysis frames it precisely: unlike plagiarism, AI-generated content “lacks verifiable textual overlap, rendering accusations inherently ambiguous,” and false-positive labels are “difficult to conclusively refute” (Doğan & Doğan, 2026). The authors note that both human experts and current detectors perform only marginally better than chance at telling the two apart, and argue that any detector output must be disclosed and paired with “a transparent process that allows authors to respond to or contest such findings” (Doğan & Doğan, 2026). That is the standard you are entitled to. For the fuller breakdown of what the comparison tests measured, see our companion piece on whether AI detectors are accurate and how detectors work at all.
The precedent is now on your side
You are not the first person to fight this, and the institutions themselves are retreating. Vanderbilt University disabled Turnitin’s AI detector back in 2023, noting that of the 75,000 papers it had run the prior year, roughly 750 could have been wrongly flagged had the tool been active (Vanderbilt University, 2023). More than fifty universities across the US, UK, Canada and Australia have since disabled or restricted the same feature.
The courts have started to agree. In Matter of Newby v. Adelphi University, decided by the Supreme Court of Nassau County, New York on January 29, 2026, a judge annulled a university’s AI-cheating finding as “without valid basis and devoid of reason” and ordered the school to expunge the student’s record (Newby v. Adelphi University, 2026; Palmer, 2026). The student — a freshman with autism whose paper Turnitin had flagged as 100% AI — had submitted independent checks from other tools rating his essay human-written, and the court held the university’s process failed basic fairness (Palmer, 2026). A detector score, the ruling made clear, is not due process.
What to actually say
Keep your written response short, factual, and unemotional. Three moves, in this order:
- State it plainly: “I wrote this work myself and I’m requesting the opportunity to demonstrate that.”
- Offer the process file: “I can provide my complete draft version history, outlines, and prior work, and I’m available to discuss the paper in person.”
- Put the tool in its place — with citations, not indignation: “I’d ask that the detector result be weighed against the peer-reviewed evidence that these tools are unreliable, including Weber-Wulff et al. (2023), which measured a 50% false-positive rate for one widely used detector, and Kumar et al. (2024), which showed detectors can be made to flag 100% of genuine human text.”
That is the whole playbook: refuse to confess, preserve your trail, produce your process, and force the score to compete with the science instead of standing in for it. The number on that report describes a probability, and the research is unanimous that the probability is worst exactly where the cost of being wrong is highest. You are allowed to say so — and now you have the page and paragraph to prove it.
Sources
Every factual claim above is tied to a source you can open and check, with page numbers wherever the source has them.
- 1. Debora Weber-Wulff, Alla Anohina-Naumeca, Sonja Bjelobaba, Tomáš Foltýnek, Jean Guerrero-Dib, Olumide Popoola, Petr Šigut, Lorna Waddington, Testing of detection tools for AI-generated text , International Journal for Educational Integrity, 2023 , pp. 11, 19. 10.1007/s40979-023-00146-z
- 2. Aounon Kumar, Sriram Balasubramanian, Vinu Sankar Sadasivan, Wenxiao Wang, Soheil Feizi, Can AI-Generated Text be Reliably Detected? , arXiv preprint, 2024 , pp. 4, 6, 10, 19. link
- 3. Ahmed M. Elkhatat, Khaled Elsaid, Saeed Al-Meer, Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text , International Journal for Educational Integrity, 2023 , p. 9. 10.1007/s40979-023-00140-5
- 4. Jezreel Edriene J. Gotoman, Harenz Lloyd T. Luna, John Carlo S. Sangria, Cereneo S. Santiago Jr, Danel Dave Barbuco, Accuracy and Reliability of AI-Generated Text Detection Tools: A Literature Review , American Journal of Interdisciplinary Research and Bibliometrics, 2025 , pp. 4, 7. 10.54536/ajirb.v4i1.3795
- 5. Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, James Zou, GPT Detectors Are Biased Against Non-Native English Writers , Patterns; Stanford Institute for Human-Centered AI (HAI), 2023 , p. 2. link
- 6. Ahmet Rıdvan Doğan, Ali İrfan Doğan, From the Turing Test to AI detectors: an epistemological mismatch in scholarly publishing , Brazilian Journal of Anesthesiology, 2026 . 10.1016/j.bjane.2026.844740
- 7. Matter of Newby v. Adelphi University, 2026 NY Slip Op 26021 (Supreme Court, Nassau County) , New York Other Courts / Justia, 2026 . link
- 8. Kathryn Palmer, Adelphi Student Wins AI Plagiarism Lawsuit , Inside Higher Ed, 2026 . link
- 9. Vanderbilt University (Center for Teaching), Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector , Vanderbilt University, 2023 . link