AI Writing Tips

What AI Feedback Actually Improves in L2 Writing

· 6 cited sources

Ask what AI feedback does for a second-language writer and you get a sales pitch: it makes you a better writer. That claim is too big and too vague to be true. Strip it down to what the studies actually measure and two specific things survive. AI feedback reliably improves grammatical accuracy — fewer article, preposition, and agreement errors in the finished text — and it improves lexical fluency, meaning the writing flows with less repetition and fewer word-choice stumbles. Everything past that gets shaky fast. It does not reliably improve how you organize an argument, it can quietly flatten your vocabulary, and — the distinction that matters most — improving the text is not the same as teaching the writer. This piece is about where that line falls, with the numbers attached.

The reason to be precise is practical. If you know AI feedback is a grammar-and-flow tool, you use it for grammar and flow and keep doing the harder work yourself. If you believe the pitch, you outsource the parts it was never good at and wonder why your unaided writing does not move.

What “accuracy” and “fluency” actually mean here

Before the results, two definitions, because the whole argument rests on keeping them apart. Interlanguage error research does not treat L2 writing as one blurry quality score — it sorts errors into a taxonomy. The Cambridge Learner Corpus alone annotates over 70 error types spanning the lexical, syntactic, and morphosyntactic properties of English, and the European SLATE project built its CEFR descriptors around distinct axes: grammatical accuracy, vocabulary control, vocabulary range, orthographic control, and cohesion and coherence (MacDonald et al., 2013, p. 39). Those are not the same skill. A tool can lift one and leave the others untouched, which is exactly what happens.

And the errors AI feedback is best at are the small, local, rule-bound ones. In a Spanish-learner corpus, MacDonald and colleagues found lexical errors clustering on prepositions — in versus on, of versus from — where the learner’s first language offered one form and English demanded a choice, a textbook case of L1 interference (MacDonald et al., 2013, p. 47). Their corpus also confirmed lexical errors were the heavy category, with a single study they cite concluding that vocabulary is “probably grossly under-taught” to these learners (MacDonald et al., 2013, p. 47). This is the terrain: articles, prepositions, agreement, word choice. It is also precisely the terrain a pattern-matching tool handles well — and, as we will see, the terrain where its help ends.

Grammatical accuracy: the clear win

This is the finding that holds up across tools and studies. When Lee (2021) compared machine-translation output to EFL university students’ own writing, the machine beat the students on spelling, vocabulary, and grammatical accuracy (Chon & Shin, 2023, p. 2). Cancino and Panes (2021) randomly assigned high-schoolers to translation-tool access or none and found higher accuracy scores in the groups with the tool (Chon & Shin, 2023, p. 2). Chung and Ahn (2021) analyzed L2 texts and reported “major improvements in accuracy” from the same kind of assistance (Chon & Shin, 2023, p. 3). Different learners, different designs, same direction.

Current commercial tools land in the same place. A 2025 study of Grammarly on 99 students’ argumentative writing measured its feedback precision at 84.93% across error categories, with determiner errors — the a/an/the problem that torments most non-native writers — detected especially accurately (Hoang & Van, 2025). A systematic review of 24 Grammarly studies reached the sober version of this: the tool is genuinely effective on local, surface-level errors — articles, prepositions, verb-noun agreement — while awkward wording and cohesion still need a human (Dizon & Gayed, 2024). That boundary is the honest one. Grammar accuracy: yes. The rest: not really.

None of this is surprising once you see the mechanism. Grammatical accuracy at the sentence level is close to a rule-checking problem, and rule-checking is what these systems do best. The same predictability that makes second-language prose read as “safe” — a trait we unpack in the context of why detectors flag non-native English writers — is what makes its errors legible and fixable to a model.

Lexical fluency improves — but read the fine print

The vocabulary story is where the honest answer diverges from the pitch. Yes, AI feedback improves lexical fluency. Chung and Ahn found that translation assistance “helped learners increase lexical variation” — more different words, less clunky repetition (Chon & Shin, 2023, p. 3). If fluency means the text stops circling the same three verbs, the tool delivers.

But the same study, in the same sentence, records the catch: that gain in variation came with lower lexical sophistication, because the tool was “recommending a wide range of common and frequently-used vocabulary” (Chon & Shin, 2023, p. 3). Read that carefully. The writing gets smoother and more varied while simultaneously getting plainer — the model reaches for high-frequency words because those are the safe, probable choices. You trade up on flow and down on precision at the same time. This is why “AI improved my vocabulary” is a half-truth: it improved the fluency of your word choices and dampened their sophistication, and which of those you care about depends entirely on what you are writing.

There is a sharper warning buried in the same body of work. Chon and Shin found that when less skilled learners post-edited machine output, their heavy use of paraphrase and deletion strategies actually lowered sentence adequacy scores — the edits made meaning worse, not better (Chon & Shin, 2023, p. 20). The tool hands a weak writer suggestions they are not yet equipped to judge, and the result can degrade. AI feedback is not uniformly additive. Its value scales with the reader’s ability to tell a good suggestion from a bad one — which is a skill the tool does not supply.

Where the improvement stops

Push past grammar and vocabulary and the gains thin out. Zhihui Zhang and Thomas Chiu ran 60 Chinese EFL students through 12 weeks of feedback, one group on generative-AI feedback alone and one on hybrid feedback that added a human tutor. The AI-only group improved on grammar and sentence variety — exactly the two things this article says AI is good at — but its impact on higher-order skills like organization and critical thinking was limited (Zhang & Chiu, 2025). The hybrid group, with a human in the loop, pulled ahead precisely on organization and critical thinking. The AI moved the sentence-level metrics and stalled on the structural ones.

That pattern is consistent with everything above. Coherence and argument are the parts of writing that a next-word predictor is worst at, because they are global properties of a whole text, not local rule violations inside a sentence. The systematic review put the same limit in institutional terms: Grammarly “lacks teaching presence,” which is the reason it cannot replace instruction even where it improves the draft (Dizon & Gayed, 2024). The tool corrects. It does not teach the shape of an argument, and the studies stop pretending it does. If you want the higher-leverage prompting moves that do touch structure, they are a different discipline — one we cover in prompt techniques that measurably improve AI writing.

The line the pitch erases: the text versus the writer

Here is the distinction the marketing collapses on purpose. Improving a document is not the same as improving the person who wrote it. Vocabulary research draws this line cleanly. Bialystok and Sharwood Smith separated two things: lexical knowledge, the words stored in a learner’s mental lexicon, and lexical control, the processing system that retrieves and deploys those words in real time under the pressure of actual writing (Liu, 1998, p. 8). AI feedback operates on the page. It can hand you a better word for this sentence. It does nothing directly to the mental lexicon or the control system that would let you produce that word yourself next Tuesday with no tool open.

This is why the honest claim is narrow. When AI feedback cleans up your prepositions, your finished essay has correct prepositions — that is real and worth having. But whether you now know the preposition rule depends on what you did with the correction, not on the correction itself. The learning happens in the reaction, not the receipt. Zhang and Chiu’s hybrid advantage points the same way: the group that had to negotiate feedback with a human internalized more than the group that just received machine corrections and moved on (Zhang & Chiu, 2025). The tool that makes revision effortless is, for exactly that reason, the one that teaches least — because the friction you remove is the friction where learning lives.

How to use it, given all that

None of this is an argument against AI feedback. It is an argument for aiming it correctly. The evidence says use it for what it demonstrably improves and keep your hands on what it does not.

Lean on it for the local layer. Articles, prepositions, subject-verb agreement, spelling, and obvious word-repetition — the high-frequency L1-interference errors MacDonald’s corpus catalogs (2013, p. 47) — are what these tools clear with real precision (Hoang & Van, 2025). Run the draft through, accept the mechanical fixes, and reclaim the time you would have spent hunting for a missing the.

Distrust it on sophistication. When it swaps a precise term for a common one, that is the plainness bias operating (Chon & Shin, 2023, p. 3), not an improvement. If the exact word mattered, keep yours. Fluency is the tool’s gift; precision is still your job.

Do the structural work before the sentence work. Organization, argument, and coherence are where AI feedback stalls (Zhang & Chiu, 2025), so build them yourself and use the tool only for the surface cleanup afterward. Cleaning sentences inside a broken structure just produces well-punctuated incoherence.

Read every correction as a lesson, or you are renting the skill. The gap between lexical knowledge and lexical control (Liu, 1998, p. 8) is closed by paying attention to why a fix was made, not by collecting fixes. A writer who reacts to feedback keeps the skill; a writer who accepts it keeps only the document. That, in the end, is what AI feedback actually improves in L2 writing: your text, reliably — and your writing, only if you make it.

Sources

Every factual claim above is tied to a source you can open and check, with page numbers wherever the source has them.

  1. 1. Yuah V. Chon, Dongkwang Shin, Second language learners' post-editing strategies for machine translation errors , Language Learning & Technology, 2023 , pp. 2, 3, 20. 10.64152/10125/73523
  2. 2. Penny MacDonald, Amparo García-Carbonell, José Miguel Carot-Sierra, Computer Learner Corpora: Analysing Interlanguage Errors in Synchronous and Asynchronous Communication , Language Learning & Technology, 2013 , pp. 39, 45, 47. 10.64152/10125/44323
  3. 3. Jiawei Liu, A corpus-based investigation of the lexis of the postgraduate engineering textbooks with reference to the needs of Southeast Asian students , University of Liverpool (doctoral thesis), 1998 , p. 8. link
  4. 4. Giang Thi Linh Hoang, Khanh Tran Gia Van, Grammarly Feedback on EFL Learners' Writing: Feedback Precision and Student Perceptions , EIKI Journal of Effective Teaching Methods, 2025 . link
  5. 5. Gilbert Dizon, John M. Gayed, A systematic review of Grammarly in L2 English writing contexts , Cogent Education, 2024 . 10.1080/2331186X.2024.2397882
  6. 6. Zhihui Zhang, Thomas K. F. Chiu, The role of generative AI and hybrid feedback in improving L2 writing skills: a comparative study , Innovation in Language Learning and Teaching, 2025 . 10.1080/17501229.2025.2503890
More on Writing English as a Non-Native
Built by Sitario.com