AI Writing Tips

What a Fabricated Citation Actually Costs

· 10 cited sources

A fabricated citation used to feel like a victimless error — an embarrassing typo the model made, caught or not, no harm done. That is no longer a defensible view, and the reason is that the cost has become measurable. It has a docket number, a dollar figure, and a growing public ledger. The first thing this article does is name the price. The second, and more important, thing it does is name the price nobody puts on an invoice: the fabricated citation that slips through, that no judge ever sees, that quietly becomes part of the record everyone else builds on.

Start with the one that got a number attached.

The case with a docket number

Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023), began as an ordinary personal-injury suit — a passenger said a metal serving cart injured his knee on an Avianca flight. It ended as the reference point for what AI fabrication costs, because the plaintiff’s lawyers filed a brief built on six court decisions that do not exist. Varghese v. China Southern Airlines, Martinez v. Delta Airlines, Shaboon v. Egyptair, Petersen v. Iran Air, Estate of Durden v. KLM, Miller v. United Airlines — all invented, all supplied by ChatGPT, all cited as binding authority.

The mechanics of how it fell apart are worth keeping in view, because they are the same mechanics that recur in every case since. Opposing counsel could not find the cases. The court could not find the cases. When Judge P. Kevin Castel ordered the attorney to produce them, ChatGPT obligingly generated full text for the fake opinions, complete with fabricated internal citations to further fake cases. Steven Schwartz, who had done the research, later said he had asked ChatGPT whether the cases were real; it said yes. That is the whole failure in miniature — the tool that invents the citation is also the tool you ask to verify it, and it will confirm its own fabrication without hesitation.

On June 22, 2023, Castel imposed a $5,000 sanction jointly on Schwartz, the attorney of record Peter LoDuca, and their firm Levidow, Levidow & Oberman (Mata v. Avianca, 2023). By later standards the fine was small. Its significance was categorical: a court had ruled that submitting an AI-fabricated citation is not a clerical slip but a false statement to the tribunal, sanctionable as such. The number was almost beside the point. The precedent was that there would now be numbers.

Why it was not a one-off

The instinct in 2023 was to treat Mata as a freak event — one careless lawyer, one viral news cycle. Three years of data have retired that reading. The clearest accounting is a public database maintained by the legal researcher Damien Charlotin, which logs court decisions worldwide where a party relied on AI-hallucinated material and a court responded. As of mid-2026 it records on the order of 1,600 such cases, more than a thousand of them in the United States (Charlotin, 2026). What was a headline is now a category of docket.

And the price has climbed steeply. The $5,000 in Mata has been eclipsed many times over. By 2026, single-matter sanctions have reached the tens of thousands; a federal appeals court has fined individual attorneys $15,000 apiece; and the largest aggregate penalty tied to one matter — Couvrette v. Wisnovsky, in the District of Oregon — ran to roughly $109,700 in sanctions, fines, and opposing fees after counsel filed fifteen fake citations and eight fabricated quotes across three briefs. Courts have moved past money, too: in a 2026 Mississippi case both sides filed hallucinated citations, and the judge suspended two lead attorneys from practice in the district for two years and canceled the trial outright.

This is not confined to the American courtroom, and it is not confined to law. Legal scholarship now treats it as a structural liability: where a representative asks an AI tool to summarize cases or draft parts of a brief, a failure to review the output can produce material “far from the factual and legal perspective,” and both the party and its counsel bear responsibility for the uncorrected result (Łągiewska, 2025, p. 101). The escalation from a $5,000 slap to five- and six-figure penalties, suspensions, and canceled trials is the market — such as it is — pricing the error in real time.

Why the model builds something this expensive

None of this happens because the model is trying to deceive. It happens because a language model predicts the next plausible token and has no table of real cases to consult. Fabricated citations are dangerous precisely because they are engineered — by the training objective — to look legitimate at first glance (Walters & Wilder, 2023, p. 2). Plausibility is the only property the model optimizes for, which is exactly why a fake Varghese v. China Southern reads like a real airline case and why the model will vouch for it when asked.

The most vivid demonstration comes from medicine. When Hussam Alkaissi and Samy McFarlane asked ChatGPT for references on bone metabolism, it returned citations carrying real PubMed IDs — but the IDs pointed to unrelated papers. One fabricated reference about homocysteine and bone metabolism wore a PMID that actually belonged to a urology paper on titanium surgical staples (Alkaissi & McFarlane, 2023, p. 2). Every component looked verifiable. The connection was invented. Oxford researchers have a sharper word for this specific failure — a confabulation: a case where the model “fluently” produces an answer that is both wrong and arbitrary, sensitive to nothing more than the random seed, so that asking twice can yield two different fake citations (Farquhar et al., 2024, p. 625). That same Nature paper reaches, for its opening example of the stakes, the lawyer sanctioned for fabricated legal precedents — the Mata case, cited in the scientific literature as the canonical instance of the problem.

The quieter cost — the one nobody catches

Here is the part that never reaches a docket. Every sanctioned case is, by definition, a fabrication that was caught. Opposing counsel checked, the judge checked, the citation collapsed. But the catch rate is not 100%, and the base rates are high enough that plenty must be getting through. A legal-hallucination study found rates running from 58% with GPT-4 to 88% with an open model when asked verifiable questions about federal cases (reported in Misra & Udandarao, 2026). Not every brief with an invented citation draws an adversary diligent enough to run it down.

In fields without an opposing counsel, the uncaught fabrication is the normal case. Robin Emsley, editor of the journal Schizophrenia, described being fed references for a study he was planning and, on checking, finding them fictitious; he cites one audit of AI-generated medical articles in which, of 115 references, 47% were fabricated, 46% were authentic but inaccurate, and only 7% were both real and accurate (Emsley, 2023). When he confronted the model, it “doubled down.” A citation like that, dropped into a manuscript by a rushed author, has no judge downstream. It has a peer reviewer who may not check every reference, and then it has the printed record.

That is the true cost of the fabrication nobody catches: it does not stay contained. Walters and Wilder note that at least two of the fabricated citations in their own study pointed to journals whose publishers have been identified as predatory, and they warn plainly that editors and publishers should ensure fabricated citations “do not find their way into the scholarly literature” (Walters & Wilder, 2023, p. 6). Once a fake reference is printed, it can be cited by the next paper, which lends it the appearance of provenance, which makes it harder to dislodge. The sanctioned lawyer pays $5,000 and the ledger closes. The uncaught citation in a review article keeps charging interest, to every reader who trusts it, for as long as the paper is in circulation. No one ever gets the invoice, which is exactly why it is the more expensive kind.

What the loud cost and the quiet cost share

They are the same failure priced two different ways, and they have the same defense. In both, a fluent, confident string of author-title-year-DOI is being treated as evidence when it is only a probability. The counter is not to trust the model more or prompt it more politely; it is to verify existence before anything else. Confirm the paper is before you argue about what it says — and for a reference list, that means checking each entry against a real index rather than re-asking the model that produced it. A free tool such as cytado.com’s bibliography checker — cytado is owned by this site’s operator, so treat that as a disclosed interest, not a neutral recommendation — resolves each entry against bibliographic databases and flags the ones with no credible match, which is the same first move a diligent opposing counsel makes by hand. This is the workflow the rest of our fact-checking guides are built around.

Two things sharpen the point. First, self-verification helps but does not close the gap: when researchers prompted ChatGPT to check and revise its own citations, the fully-accurate share rose from 7.5% to 77.5% (Safran & Çalı, 2025, p. 695) — a real gain that still leaves roughly a fifth of the list wrong, now wearing fresh confidence. Second, the durable fix is architectural. A model asked to recall a citation confabulates; a model connected to a real database and asked to retrieve one does far better. One such system, fetching entries directly from an authoritative index rather than through the model’s own text, hit an 82.7% perfect-match rate against 28.2% for plain web search, with metadata corruption eliminated entirely (Szeider, 2025, pp. 2–3). The lesson is the same one Mata taught the courts: the fabrication is cheap for the model to produce and expensive for everyone downstream to absorb, so the only economical move is to catch it before it ships.

If you want the underlying rates that make this risk concrete — which models fabricate how often, and why a single percentage is always misleading — our reference table of AI hallucination rates keeps each figure attached to its model and task, and the companion piece on whether ChatGPT makes up sources walks through the measurements in detail.

The bill, itemized

So what does a fabricated citation actually cost? When it is caught, a documented range: $5,000 in Mata v. Avianca, up to roughly $109,700 in Couvrette, plus fees, suspensions, canceled trials, and a name in a public database now well past a thousand entries. When it is not caught, something harder to bill but larger in aggregate — a false reference propagating through a literature, cited onward by people who assume someone already checked.

Set that against the cost of prevention, which is a few minutes running a reference list past an index that says yes or no. That asymmetry is the entire argument. The expensive part was never the checking. It was skipping it.

Sources

Every factual claim above is tied to a source you can open and check, with page numbers wherever the source has them.

  1. 1. P. Kevin Castel (opinion & order), Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023) , U.S. District Court, Southern District of New York, 2023 .
  2. 2. Damien Charlotin, AI Hallucination Cases (database) , damiencharlotin.com, 2026 . link
  3. 3. William H. Walters, Esther Isabelle Wilder, Fabrication and errors in the bibliographic citations generated by ChatGPT , Scientific Reports, 2023 , pp. 1, 2, 6. 10.1038/s41598-023-41032-5
  4. 4. Hussam Alkaissi, Samy I. McFarlane, Artificial Hallucinations in ChatGPT: Implications in Scientific Writing , Cureus, 2023 , p. 2. 10.7759/cureus.35179
  5. 5. Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, Yarin Gal, Detecting hallucinations in large language models using semantic entropy , Nature, 2024 , p. 625. 10.1038/s41586-024-07421-0
  6. 6. Robin A. Emsley, ChatGPT: these are not hallucinations – they're fabrications and falsifications , Schizophrenia (npj), 2023 . 10.1038/s41537-023-00379-4
  7. 7. Nipun Misra, Vikranth Udandarao, Detecting Citation Hallucinations in Large Language Model Outputs , Proceedings of the AAAI Conference on Artificial Intelligence, 2026 . 10.1609/aaai.v40i48.42257
  8. 8. Ertuğrul Safran, Adem Çalı, Fabricated or accurate? Ethical concerns and citation hallucination in AI-generated scientific writing on musculoskeletal topics , Anatolian Current Medical Journal, 2025 , p. 695. 10.38053/acmj.1746227
  9. 9. Magdalena Łągiewska, Artificial Intelligence and International Arbitration Law: Revolution or Evolution , Routledge, 2025 , p. 101. 10.4324/9781003667834
  10. 10. Stefan Szeider, Unmediated AI-Assisted Scholarly Citations , Open Conference Proceedings (AAAI-26), 2025 , pp. 2, 3. 10.52825/ocp.v8i.3161
More on Fact-Checking AI Output
Built by Sitario.com