AI Writing Tips

Chain-of-Thought for Writers, Not Mathematicians

· 5 cited sources

Chain-of-thought prompting was born on a math benchmark, tested on math and logic, and then repackaged as a universal spell you cast on any request. “Let’s think step by step” now appears in prompt guides for cover letters, marketing copy, and short stories — genres that have no steps to think through. The trick is real, and the evidence for it is strong, but the evidence is about arithmetic and symbolic manipulation, not prose. This article separates the part of chain-of-thought that genuinely helps a writer from the much larger part that was measured on problems a writer never has to solve.

Start with what the technique actually is, because the name is doing a lot of persuasive work. Chain-of-thought means adding a step-by-step breakdown — “a coherent set of intermediate reasoning steps towards the final answer” — to the prompt, and it was, in the words of the team that folded it into medical AI, “designed to mimic the human thought process when solving problems that require multi-step computation and reasoning” (Singhal et al., 2023, p. 183). Read that description closely. It is a method for problems that decompose into a sequence of computations. That is the shape of a math word problem. It is not the shape of an essay.

The benchmark that started it was arithmetic

When Jason Wei and colleagues introduced chain-of-thought in 2022, they demonstrated it on GSM8K — grade-school math word problems — and on arithmetic, commonsense, and symbolic reasoning tasks, showing that a handful of worked exemplars pushed a large model to state-of-the-art accuracy (Wei et al., 2022). The Nature team that adopted it a year later was blunt about the domain: chain-of-thought “can elicit reasoning abilities in sufficiently [large] LLMs and dramatically improve performance on tasks such as mathematical problems,” and its very appearance “appears to be an emergent ability” that has “been used to achieve breakthrough LLM performance on several STEM benchmarks” (Singhal et al., 2023, pp. 183–184). Every foundational result points the same way: math, logic, STEM. Nobody in the origin literature was claiming it made writing better, because that is not what they measured.

The leap from “dramatically improves math” to “improves everything you ask a model to do” happened in blog posts, not in papers. And it is exactly the leap our field guide to the prompt-engineering techniques that actually improve writing warns against: a move that lifts arithmetic accuracy is not automatically a move that improves an argument. The burden of proof was quietly skipped.

The meta-analysis that measured the gap

Someone finally checked. In 2024, Zayne Sprague and colleagues ran the largest audit of chain-of-thought to date: a quantitative meta-analysis of over 100 papers, plus their own evaluations across 20 datasets and 14 models. The headline finding is the one the folklore ignores — “CoT gives strong performance benefits primarily on tasks involving math or logic, with much smaller gains on other types of tasks” (Sprague et al., 2024). The gains are not spread evenly across everything a language model can do. They pool almost entirely in the symbolic corner.

The sharpest detail is what they found on MMLU, a broad knowledge benchmark: “directly generating the answer without CoT leads to almost identical accuracy as CoT unless the question or model’s response contains an equals sign” (Sprague et al., 2024). An equals sign. The presence of a literal symbolic operation was the switch that turned chain-of-thought from useless to useful. When they dug into the mechanism, they concluded that “much of CoT’s gain comes from improving symbolic execution” — and that even there, it underperforms a real symbolic solver (Sprague et al., 2024). This is the empirical spine of the whole argument: chain-of-thought is, at bottom, a way of getting a language model to carry out step-by-step computation it would otherwise botch. Writing is not computation.

What does transfer to writing

So the technique is not worthless for writers — it is just narrow, and knowing the narrow band is the point. Chain-of-thought helps writing precisely where writing borrows the shape of a reasoning task: when the piece has a logical spine that has to hold. Building a structured argument, working through a comparison, reaching a conclusion that has to be defensible rather than merely fluent — these have intermediate steps, and forcing the model to lay them out before it drafts stops it from skating to a smooth, unsupported ending.

The usable move is not “think step by step” sprinkled over a creative request. It is: separate the planning pass from the drafting pass. Ask the model to outline the argument — claim, evidence, objection, response — as its own step, inspect that skeleton, fix it, and only then ask for prose built on it. That is chain-of-thought applied honestly: you are using it for the part of writing that actually decomposes into steps, and you are doing the decomposition where you can see it. It pairs naturally with treating the outline as the real deliverable, which is the discipline behind the brief being the prompt — the thinking happens before the sentences, not inside them.

There is a second honest use, closer to the technique’s mechanical origin. A 2026 comparison of prompting styles on document-based extraction argued that forcing sequential reasoning gives a model room to “review” the order of retrieved information before answering, redirecting its attention toward the supplied document and away from irrelevant internal knowledge (Yahya et al., 2026, pp. 2399–2400). For a writer, the analogue is fact-heavy work grounded in a source you paste in: making the model walk through the document step by step before it summarizes does reduce the odds it wanders off into what it half-remembers. That is a real, if modest, benefit — and it is about grounding, not eloquence.

What does not transfer

Now the larger territory, where the trick is cargo-culted. Two failures matter for writers.

First, chain-of-thought does not make prose truer or the voice better. It targets the reasoning trace, and a model’s central failure in writing is not faulty logic — it is bland averaging and confident invention, neither of which a “let’s reason this out” preamble fixes. Sprague’s own conclusion is that the benefit lives in symbolic execution; there is no symbolic execution in “make this paragraph warmer.” Asking a model to reason step by step about tone produces a reasoning-shaped monologue followed by the same generic paragraph you would have gotten anyway.

Second, and more damning: even where chain-of-thought works, it is brittle. The 2026 survey of large language models notes that its performance “degrades sharply under distribution shifts in task, length, or format,” and that surface-similar prompts diverge — “different CoT prompts might impact model performance on certain tasks, despite their apparent surface similarity” (Zhao et al., 2026, p. 2009). Brittleness under format and length shifts is fatal for writing, because writing is a distribution shift away from the benchmarks: long, open-ended, formatless, with no single correct answer to converge on. A method that wobbles when you change the length of the problem is a method you should not lean on for a 2,000-word draft.

This is also why the famous incantations belong in a museum. The survey catalogues them plainly: zero-shot chain-of-thought prompts models with phrases like “Let’s think step by step,” and researchers have even explored “Take a deep breath and work on this problem step-by-step” (Zhao et al., 2026, p. 2009). Those phrases were tuned and tested against reasoning benchmarks. Pasting them onto a request to write a LinkedIn post is imitating the ritual while discarding the only condition under which it was ever shown to work.

The 2026 twist: the models already do it

Here is the part that retires the debate for most writers. The step-by-step behavior has been trained into current models. The same survey observes that “CoT-style reasoning data has been widely incorporated into model training,” so “mainstream LLMs can often generate step-by-step solutions even without explicit CoT prompting” — and for recent long-reasoning models like OpenAI’s o1, it “even suggests avoiding CoT prompts” (Zhao et al., 2026, p. 2009). By 2026 the frontier has split: reasoning models with extended-thinking modes run the chain internally, spend real time on it, and hide most of it from you. Telling one of them to “think step by step” is at best redundant and at worst a contradiction of the process it already runs.

For the writer, the takeaway inverts the folklore. On a modern reasoning model, you do not need to script the reasoning; you need to give it something worth reasoning about — a real brief, a real structure, real constraints. On a smaller or cost-optimized model, explicit chain-of-thought still earns its keep, but only on the sliver of writing tasks with a genuine logical spine, and only if you accept the brittleness. Either way, the instruction that changes your draft is not “reason first.” It is the specificity you feed in and the self-critique pass you run afterward — the two moves that survive measurement on prose, rather than the one imported wholesale from a math benchmark.

Chain-of-thought is a good tool aimed at the wrong job when a writer reaches for it by reflex. Use it as a planning and grounding step, on the tasks that decompose. Drop it everywhere else. And when you catch yourself typing “let’s think step by step” over a request that has no steps, remember what it was measured on — an equals sign — and ask whether your sentence has one. The rest of the pillar collects the moves that do hold up for writing at the prompting hub.

Sources

Every factual claim above is tied to a source you can open and check, with page numbers wherever the source has them.

  1. 1. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, Denny Zhou, Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , Advances in Neural Information Processing Systems (NeurIPS), 2022 . link
  2. 2. Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Lee, Hyung Won Chung, Nathan Scales, Ajay Kumar Tanwani, et al., Large language models encode clinical knowledge , Nature, 2023 , pp. 183, 184. 10.1038/s41586-023-06291-2
  3. 3. Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, Greg Durrett, To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning , International Conference on Learning Representations (ICLR 2025); arXiv, 2024 . link
  4. 4. Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Zican Dong, Yupeng Hou, Beichen Zhang, Yingqian Min, et al., A Survey of Large Language Models , Frontiers of Computer Science, 2026 , pp. 2008, 2009. 10.1007/s11704-026-60308-3
  5. 5. Kurnia Yahya, Dikwan Moeis, Musdalifa Thamrin, Sry Yunarti, Alvina Felicia Watratan, Satriawaty Mallu, Analisis Perbandingan Efektivitas Zero-Shot vs Chain-of-Thought Prompting dalam Meningkatkan Presisi Informasi pada LLM Berbasis Dokumen PDF , Journal of Artificial Intelligence and Digital Business (RIGGS), 2026 , pp. 2399, 2400. 10.31004/riggs.v5i1.6524
More on Prompting for Usable Drafts
Built by Sitario.com