AI Writing Tips

The Detector Score Itself Lowers Your Grade

· 7 cited sources

Here is the finding that should end the argument about whether AI detectors are “just one input” a teacher weighs alongside their own judgment: the number on the detector report changes the judgment. Not the writing — the writing stays identical. Show a grader a paper with a low AI-detection score and they read it one way. Show a different grader the same paper with a high score printed on top, and they mark it down across every dimension they were asked to assess — originality, language, logical structure, the overall grade. The content did nothing. The percentage did all the work. That is not a hypothetical risk. In 2026 it was measured, on 214 teachers, and the effect was enormous.

The experiment that isolates the number

The study comes from Peitao Du and colleagues, published in Frontiers in Psychology. Its design is the thing to understand, because it strips away every confound. In a 2×2 between-subjects experiment, 214 university teachers each evaluated the same medium-quality student course paper, accompanied by a fictitious AI-detection report. Two things varied between groups: the detection rate shown (a calm 7% versus an alarming 87%), and whether the report used red highlighting on the flagged passages (Du et al., 2026, p. 1). Nobody changed a word of the essay. The only variable was what the report said about it.

The headline result is in the abstract, in plain language: a high detection rate “increased perceived AI authorship likelihood and risk and lowered overall quality evaluations, percentage-based scores, originality, language expression, and logical structure” (Du et al., 2026, p. 1). Read that list again. Originality. Logical structure. Those are judgments about the argument on the page. The argument did not move. The teachers’ reading of it moved, because a machine told them to be suspicious.

What the drop actually looks like

The magnitudes are the part that stops you. On the 0–100 percentage score, the group shown a 7% detection rate gave the paper an average of 73.30. The group shown 87% — same paper — gave it 60.64 (Du et al., 2026, Table 1). That is a 12.7-point penalty applied to writing that was, by construction, exactly as good in both conditions. On a 1–10 overall rating, the paper fell from 6.83 to 5.02. Originality dropped from 4.39 to 2.89. Language expression from 5.73 to 4.05. Logical structure from 4.49 to 3.53. Every dimension, the same direction, all of it downstream of a two-digit number.

These are not noisy trends fished out of a small sample. The main effect of the warning on the percentage score was F(1,210) = 615.06 with a partial eta-squared of 0.745 (Du et al., 2026, Table 2) — meaning the detection number alone accounted for roughly three-quarters of the variance in how teachers scored an unchanging paper. In social-science terms, an effect that size is not a whisper. It is the loudest thing in the room.

The red ink makes it worse

The visual design of the report had its own, separate effect. Add red highlighting to the high-detection condition and the average score falls further still — down to 51.61 in the group that saw both an 87% rate and red-marked passages (Du et al., 2026, Table 1). The interaction was statistically significant: warning strength and visual alarm reinforce each other rather than simply adding up. Under the high-detection condition, the effect of red highlighting on the percentage score reached a Cohen’s d of −1.94, which is the kind of effect size you rarely see outside a manipulation designed to produce one.

And it did not stop at the grade. Teachers’ self-reported intervention tendency — how inclined they felt to escalate, question, or accuse — climbed from an average of 3.40 in the calmest condition to 5.66 in the loudest, a large effect (Du et al., 2026, p. 1). The report does not just lower a mark. It nudges the grader toward treating the student as a suspect. This is the mechanism behind so many stories of students falsely accused of using AI: by the time a human looks at the essay, the number has already told them what to see.

This has a name, and it is not new

What Du and colleagues documented is automation bias — the well-studied human tendency to over-trust an automated recommendation and to stop weighing evidence that contradicts it. Pair it with anchoring, the pull of the first number you see, and confirmatory reading, the habit of hunting for evidence that fits a conclusion you have already reached, and you get the exact chain the study describes: a warning forms suspicion, the visual cues trigger a confirmatory search through the text, and the evaluation slides downward while the impulse to intervene rises. The authors are careful to frame the detection report as a “decision-support display,” not a technical readout — a thing that shapes perception through numbers, labels, and color before the reader has consciously decided anything (Du et al., 2026, p. 1).

The uncomfortable implication is that the accuracy debate around AI detectors — is it 80% right, 60%, a coin flip? — misses half the damage. Even a perfectly calibrated detector would still be feeding a number into a human process that demonstrably bends around numbers. The report doesn’t have to be wrong to distort the grade. It only has to be present.

And the number it feeds you is usually shaky

Now stack the two problems. The grader anchors hard on the detection score — and that score is, on its own, one of the least reliable inputs in the room. The landmark evidence here belongs to Weixin Liang’s team at Stanford, who ran seven widely used detectors over essays by non-native English speakers and recorded an average false positive rate of 61.22%, versus near-perfect accuracy on native-speaker essays (Liang et al., 2023, p. 2). When they prompted a model to simplify an American student’s vocabulary “as if written by a non-native speaker,” the misclassification rate for that genuine human writing jumped from 5.19% to 56.65% (Liang et al., 2023, p. 2). So the very number that drags a grade down by twelve points is the one most likely to be inflated for a non-native English writer whose only offense is a smaller vocabulary.

The fragility runs deeper than one population. Vinu Sadasivan and colleagues showed that a light paraphrase drops a detector’s accuracy from 97% to 80%, and a slightly heavier one to 57%, without meaningfully degrading the text (Kumar et al., 2024, p. 6) — so the score a teacher anchors on can be moved at will by anyone who knows the trick. Ahmed Elkhatat’s team found detectors throwing false positives and “uncertain” verdicts on human-written control samples while letting GPT-4 text pass (Elkhatat et al., 2023, p. 11). A 2025 review of the whole field summarized it bluntly: detectors “produce more accurate findings for human-made input than ChatGPT-created input, yet they are still unreliable for use in the academic setting” (Gotoman et al., 2025, p. 5). This is the same failure that makes detector outputs so dangerous elsewhere — a confident, precise-looking figure with very little truth underneath, the theme running through everything we cover on whether AI detectors are accurate.

Why this matters more every semester

If detectors were a fading fad, the automation-bias finding would be a curiosity. They are not fading. According to a Center for Democracy & Technology survey of secondary-school teachers, the share using AI-detection tools jumped 30 percentage points in a single year, to 68% in 2023–24, and the share of student discipline tied to suspected generative-AI use rose from 48% to 64% over the same period (Center for Democracy & Technology, 2024). On the higher-education side, Turnitin reported that between October 2025 and February 2026, an average of 14.8% of English-language submissions to its detector came back with 80% or more flagged as AI-written (Turnitin, 2026). Whatever fraction of those flags are false positives — and the research above says it is not small — each one now lands on a grader whose scoring we know bends around the number.

That is the whole problem in one sentence. The tool that most reliably lowers a grade is also one of the least reliable pieces of evidence in the file, and it does its damage before anyone has finished reading the essay.

The fix is procedural, not technological

Du and colleagues do not argue for better detectors. They argue for a different order of operations. Their central recommendation is that AI-detection reports “should not be treated as direct evidence of writing independence” but as auxiliary cues, and — crucially — that institutions adopt a two-stage procedure: teachers complete their full quality evaluation before they open the detection report, using it only as supplementary review material afterward (Du et al., 2026, p. 13). Grade the paper blind. Then, and only then, look at the number. The sequence is the safeguard, because anchoring can only pull you toward a conclusion you have not yet reached.

Their second recommendation targets the red ink directly: report interfaces should drop the strong colors, warning icons, and high-risk labels that the experiment showed inflate the bias (Du et al., 2026, p. 13). A detection tool that must exist should at least be designed to inform a judgment rather than stampede it.

For a student on the receiving end, the practical lesson is smaller but sharper. If your grade or an accusation rests on a detection report, the number was not a neutral observation of your work — the research says it actively reshaped how your work was read. That is a legitimate thing to raise on appeal: not only “the detector is inaccurate,” but “the detector’s score measurably biases the grader, and the paper was never evaluated on its own merits.” The evidence for that sentence is now peer-reviewed, quantified, and sitting in a 2026 journal with a sample of 214 teachers behind it.

The detector was sold as a second opinion. The measurement says it is closer to a thumb on the scale.

Sources

Every factual claim above is tied to a source you can open and check, with page numbers wherever the source has them.

  1. 1. Peitao Du, Tingting Liu, Xujin Xian, Automation bias in teachers' evaluation of student writing: effects of algorithmic warnings and visual risk cues in AI detection reports , Frontiers in Psychology, 2026 , pp. 1, 13. 10.3389/fpsyg.2026.1889402
  2. 2. Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, James Zou, GPT Detectors Are Biased Against Non-Native English Writers , Patterns; Stanford Institute for Human-Centered AI (HAI), 2023 , p. 2. link
  3. 3. Aounon Kumar, Sriram Balasubramanian, Vinu Sankar Sadasivan, Wenxiao Wang, Soheil Feizi, Can AI-Generated Text be Reliably Detected? , arXiv preprint, 2024 , p. 6. link
  4. 4. Ahmed M. Elkhatat, Khaled Elsaid, Saeed Al-Meer, Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text , International Journal for Educational Integrity, 2023 , p. 11. 10.1007/s40979-023-00140-5
  5. 5. Jezreel Edriene J. Gotoman, Harenz Lloyd T. Luna, John Carlo S. Sangria, Cereneo S. Santiago Jr, Danel Dave Barbuco, Accuracy and Reliability of AI-Generated Text Detection Tools: A Literature Review , American Journal of Interdisciplinary Research and Bibliometrics, 2025 , p. 5. 10.54536/ajirb.v4i1.3795
  6. 6. Center for Democracy & Technology (reported by K-12 Dive), As teacher use of AI detection grows, discipline guidance a mixed bag , K-12 Dive, 2024 . link
  7. 7. Turnitin, Turnitin Data Shows Transparency About AI Use Benefits Students and Educators , Turnitin, 2026 . link
More on AI Detectors
Built by Sitario.com