Show, Don't Tell: Examples Beat Instructions
The most common question a working copywriter asks about AI is some version of how do I get it to write like me? The instinct is to answer it with adjectives — tell the model your voice is “warm but authoritative,” “punchy,” “conversational, not corporate.” That instinct is wrong, and the research on style transfer has been clear about why for four years. You do not describe a style to a language model. You show it one. Paste two or three passages of your actual writing and the model has the target in front of it; write a paragraph of adjectives and you have handed it a riddle it will solve with the average of everything anyone has ever called “warm but authoritative.” Show, don’t tell is not a writing-class cliché here. It is the single best-supported move in the whole field of prompting for prose.
Why adjectives fail and examples don’t
The problem with describing a style is that the words you reach for are the same words everyone reaches for. “Professional,” “engaging,” “clear” — these are requests for the mean, because every brand asks for them. An example carries information that an adjective cannot: rhythm, sentence length, where you break a paragraph, which words you would never use, how you open. It is the difference between telling a session musician “play it bluesy” and handing them a recording.
The style-transfer literature formalized this early. When Mirac Suzgun, Luke Melas-Kyriazi, and Dan Jurafsky broke the task down mathematically, they decomposed “write it in this style” into three measurable components — textual similarity to the source, target style strength, and fluency — and showed that prompting a model with the right framing lets even small language models perform arbitrary style transfer on par with far larger ones, at two orders of magnitude less compute (Suzgun et al., 2022, pp. 1, 3). The lever was never model size. It was the quality of the demonstration you put in front of it. A well-shaped example does more than a bigger model with a vague instruction.
The clean experiment: examples versus a plain instruction
The cleanest head-to-head comes from Emily Reif and colleagues at Google Research, in the paper that gave the field its working recipe. They compared three ways of asking a model to rewrite a sentence in a target style: zero-shot (a plain instruction, “rewrite this to be more positive”), few-shot (a handful of worked input–output pairs in the exact style you want), and their own hybrid. Their finding on the plain instruction is the one to tape above your desk: zero-shot prompting “can be prone to failure modes such as not returning well-formatted or logical outputs,” while few-shot prompting — showing examples — “has been shown to achieve higher performance” (Reif et al., 2022, p. 837). Instructions alone are the unreliable path. Examples are the reliable one.
They then put it to human raters. Six professional evaluators scored outputs across six non-standard styles on transfer strength, semantic preservation, and fluency — 3,600 ratings in total — and the example-driven method was rated comparably to human-written ground truth (Reif et al., 2022, p. 839). On the standard sentiment and formality benchmarks, the example-based approach “substantially outperforms vanilla zero-shot learning and almost reaches the accuracy of five-shot learning” (Reif et al., 2022, p. 840). Read that last clause slowly: their clever instruction-based method got close to plain five-shot — that is, close to just showing five examples. Showing examples was the bar everything else was trying to reach.
This is the same principle that runs under the whole prompting pillar: the input is the lever, and specificity delivered as a sample beats specificity delivered as a description.
Examples are how the model learns a task at all
None of this is special to style. It is how in-context learning works. The medical-AI literature describes few-shot prompting plainly as encoding a task “through text-based demonstrations,” typically as input–output pairs, so the model infers the pattern from the cases rather than from an abstraction (Singhal et al., 2023, p. 183). The systematic survey of the field, The Prompt Report, treats few-shot prompting as one of six top-level families of technique — not a tip but a category — and notes that well-chosen exemplars can carry instruction-following on their own, sometimes outperforming task-specific written instructions (Schulhoff et al., 2024). When practitioners describe the same thing, they land in the identical place: Anthropic’s own guidance for Claude calls examples “an effective shortcut,” recommends three to five diverse ones, and reports that more examples generally mean better performance, especially on anything nuanced (Anthropic, 2026). Every corner of the field — theory, benchmark, vendor — points the same way.
What the newest research says you can and can’t get from examples
Here is where honesty matters, because “show it examples” is not a magic wand, and the most recent study on the exact question a copywriter cares about says so out loud. In late 2025, Zhengxiang Wang and colleagues ran the largest test to date of whether models can imitate an ordinary person’s implicit writing style from a few samples — over 40,000 generations per model, more than 400 real authors, across news, email, forums, and blogs (Wang et al., 2025, p. 10040). Their verdict is split. Models can approximate your style in structured genres like news and email, but “they struggle with nuanced, informal writing in blogs and forums,” where outputs “often default to an average, generic tone and remain readily detectable as AI-written” (Wang et al., 2025, pp. 10040–10041).
Two findings from that paper should change how you work. First, they confirm why you need examples at all: left to itself, an LLM “often default[s] to a generic style learned from vast web data, stripping away the personal touch that makes writing feel authentic” (Wang et al., 2025, p. 10040). The generic draft is the baseline you are fighting, and examples are the only cheap tool that moves it. Second — and this is the counterintuitive part — they found that simply “increasing the number of demonstrations offers limited gains in stylistic alignment” (Wang et al., 2025, p. 10041). Piling in twenty samples of your writing does not linearly buy you twenty samples’ worth of fidelity. Past a small number of well-chosen examples, more is not better; better is better.
The same team notes the alternative most tools push you toward and why it fails. Configurable sliders — tone, voice, formality — “fall short as users’ personal styles are nuanced and rarely reducible to a few sliders” (Wang et al., 2025, p. 10040). That is the adjective problem again, dressed up as a UI. Your voice is not three dials. It is a pattern, and a pattern is transmitted by example.
The practical recipe
So the advice is not “show more” but “show right.” Choose two to four passages that are unambiguously you and unambiguously the kind of thing you are asking for — if you want a product email, the examples are your best product emails, not your best anything. Reif’s team found that keeping the format of the exemplars constant, so the model can see the template clearly, was part of what made the method work (Reif et al., 2022, p. 837); Anthropic’s guidance says the same thing operationally, recommending you wrap each example in tags so the model can tell your sample from your instruction (Anthropic, 2026). Concretely: label each example, keep the shape identical across them, and put them before the task.
There is one more design choice the research surfaces. Suzgun and colleagues found that a prompt specifying only the target style — the naive default — left performance on the table; giving the model the source style too, and drawing an explicit contrast between where the text is and where it should go, helped the model grasp the actual transformation (Suzgun et al., 2022, p. 5). The lesson for a copywriter is direct. Do not only show the model good examples of the destination. Show it the before and the after — a flat draft and your rewrite of it — so it learns the move, not just the endpoint. That is a stronger brief than any pile of finished samples, and it is the same discipline that makes the brief the real prompt: you are doing the thinking the model cannot do, then handing it the pattern instead of the abstraction.
Where examples stop and their limits begin
Know what you are buying. Example-driven style transfer trades control for range. Reif’s team was candid that their example-based method “offers less fine-grained controllability in the properties of the style-transferred text than methods which see task-specific training data” (Reif et al., 2022, p. 841). You get breadth — any style you can demonstrate — but not the surgical precision of a model fine-tuned on your entire corpus. For most writing that is a fine trade. When it is not, the fix is more or better examples and a sharper before/after, not a paragraph of new adjectives.
And remember what examples cannot rescue. They shape voice; they do not verify facts, and a model imitating your confident house style will imitate the confidence even when it is inventing. That is a separate failure with a separate fix — the kind of grounding and structure covered in the techniques that measurably improve AI writing. Style is a pattern problem, solved by demonstration. Truth is a retrieval problem, and no example teaches it.
The whole answer to how do I teach AI my style fits in one line. Stop describing your voice and start showing it. The flattening toward a common register that everyone complains about is measurable and real — frontier models normalize the markers of individual voice in the same direction the moment you leave them to their defaults (van Nuenen, 2026, pp. 8–9) — and the cheapest thing standing between you and that mean is a few paragraphs of your own writing, pasted in, before you ask for anything at all.
Sources
Every factual claim above is tied to a source you can open and check, with page numbers wherever the source has them.
- 1. Mirac Suzgun, Luke Melas-Kyriazi, Dan Jurafsky, Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language Models , arXiv, 2022 , pp. 1, 3, 5. 10.48550/arXiv.2205.11503
- 2. Emily Reif, Daphne Ippolito, Ann Yuan, Andy Coenen, Chris Callison-Burch, Jason Wei, A Recipe for Arbitrary Text Style Transfer with Large Language Models , Proceedings of ACL 2022 (Short Papers), 2022 , pp. 837, 839, 840, 841. 10.18653/v1/2022.acl-short.94
- 3. Zhengxiang Wang, Nafis Irtiza Tripto, Solha Park, Zhenzhen Li, Jiawei Zhou, Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authors , Findings of the Association for Computational Linguistics: EMNLP 2025, 2025 , pp. 10040, 10041. link
- 4. Tom van Nuenen, Voice Under Revision: Large Language Models and the Normalization of Personal Narrative , arXiv, 2026 , pp. 8, 9. 10.48550/arXiv.2604.22142
- 5. Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Lee, Hyung Won Chung, et al., Large language models encode clinical knowledge , Nature, 2023 , p. 183. 10.1038/s41586-023-06291-2
- 6. Sander Schulhoff, Michael Ilie, Nishant Balepur, et al., The Prompt Report: A Systematic Survey of Prompt Engineering Techniques , arXiv, 2024 . link
- 7. Anthropic, Use examples (multishot prompting) to guide Claude's behavior , Anthropic Documentation, 2026 . link