AI Writing Tips

Does ChatGPT Make Up Sources? What the Studies Actually Measured

· 9 cited sources

Yes, ChatGPT makes up sources. That part has never been in serious dispute. The useful question is how often, and the answer is not a vibe — it has been measured, repeatedly, by researchers who took thousands of AI-generated references and checked whether the papers behind them exist. This article puts five of those measurements in one place, because a single number pulled out of context (“18%!” or “90%!”) tells you almost nothing without the model, the field, and the year attached to it.

The short version: fabrication rates for ChatGPT citations run from about 18% for GPT-4 to 55% for GPT-3.5, and climb as high as 88% once you ask about case law. Other assistants do worse still — Google’s Bard failed on 91.3% of references in one head-to-head test, which is worth keeping straight, because that number gets quoted as if it were ChatGPT’s. Newer models fabricate less than older ones. Medicine and general topics are safer than law. And “fabricated” is only half the problem — even the real citations are frequently wrong in ways that matter.

The headline study: 55% versus 18%

The most careful public accounting comes from William Walters and Esther Wilder, published in Scientific Reports. They had GPT-3.5 and GPT-4 write short literature reviews on 42 topics, then hand-checked all 636 references in the 84 resulting papers against real databases.

The result is the number worth memorizing: 55% of the GPT-3.5 citations were fabricated, but just 18% of the GPT-4 citations were (Walters & Wilder, 2023, p. 1). That single generational jump — from more than half made up to fewer than one in five — is the whole story of how these models have improved. It also tells you that “does ChatGPT hallucinate citations” has no fixed answer; it depends entirely on which ChatGPT.

Their second finding is the one people forget. Fabrication is not the only failure. Among the citations that pointed to real papers, 43% of the GPT-3.5 references and 24% of the GPT-4 references still contained substantive errors (Walters & Wilder, 2023, p. 1) — wrong volume numbers, wrong pages, wrong years, wrong authors. A citation can be “real” and still send you to the wrong place. Interestingly, every single citation was formatted in clean APA style, yet more than 40% had formatting errors, most often improper capitalization (Walters & Wilder, 2023, p. 5). The model is very good at looking correct and much worse at being correct.

The range is enormous, and the field is why

Walters and Wilder also pooled six early studies and found that, across 732 citations, 51% were fabricated (Walters & Wilder, 2023, p. 2). But that average hides a huge spread, and other measurements show why.

Mikaël Chelli and colleagues, writing in the Journal of Medical Internet Research, tested ChatGPT and Google Bard on systematic-review references for rotator cuff disease. GPT-3.5 produced 39.6% (55 of 139) nonexistent references, while Bard failed on 91.3% (95 of 104) (Chelli et al., 2024, p. 7). GPT-4 came out best of the three, but hallucination rates across the models still ranged from 28.6% to that ceiling of 91.3% (Chelli et al., 2024, p. 9). Same task, same week — a nearly threefold difference depending on which chatbot you opened.

Robin Emsley, editor of the journal Schizophrenia, described the failure from the inside after ChatGPT fed him fictitious references for a study he was planning. He cites one investigation of AI-generated medical articles in which, of 115 references, 47% were fabricated, 46% were authentic but inaccurate, and only 7% were both authentic and accurate (Emsley, 2023). Another study he reports found that of 35 generated citations, only two were real. His conclusion doubles as a warning about the word we all use: these, he argues, “are not hallucinations — they’re fabrications and falsifications.”

A 2025 study by Ertuğrul Safran and Adem Çalı sharpened the picture with a three-point score. Across 40 references on musculoskeletal topics, only 7.5% were fully accurate on the first try, 42.5% were completely fabricated, and 50% were partially correct (Safran & Çalı, 2025, p. 695) — usually real papers wearing a wrong DOI or a mangled journal name.

Put those side by side and the “how often” question resolves into a table, not a number:

StudyModel(s)Fabrication / hallucination rate
Walters & Wilder, 2023GPT-3.5 / GPT-455% / 18%
Chelli et al., 2024GPT-3.5 / Bard39.6% / 91.3%
Emsley, 2023 (cited study)ChatGPT (medical)47% fabricated, 46% inaccurate
Safran & Çalı, 2025ChatGPT42.5% fabricated, 7.5% fully accurate
Misra & Udandarao, 2026 (review)GPT-4 → legal LLMs18% → 88%

Law is the worst case

If you want the scariest numbers, leave medicine and go to law. A 2026 review compiling recent measurements reports hallucination rates spanning 18% for GPT-4 up to over 70% for other frontier models, and as high as 88% in legal contexts (Misra & Udandarao, 2026). The same review notes a controlled comparison where ChatGPT-4o hallucinated on 20% of prompts while another assistant reached 76.7%. Stefan Szeider’s survey of the problem echoes the trajectory: across GPT-3, ChatGPT, and GPT-4, error rates fell from over 70% to under 50%, yet for some publication categories fabrication still exceeded 80% (Szeider, 2025, p. 2).

Those legal percentages are not academic. Since 2023, roughly 900 AI-fabricated citations have been documented in real US court filings, and the sanctions have escalated fast — from the original $5,000 penalty against the lawyer in Mata v. Avianca to a combined $110,000 fine against two attorneys in Oregon in 2026 for submitting 23 fabricated citations. Courts now treat a made-up case the way they treat any other false statement to a judge.

Why does a model invent a citation at all?

The mechanism is not mysterious. A language model predicts the next token from statistical patterns; it has no lookup table of real papers. As one group quoted by Walters and Wilder puts it, the model “cannot distinguish between accurate and false information” and simply produces the most plausible-looking string of author, title, journal, and year (Walters & Wilder, 2023, p. 6). That is why a fabricated citation looks exactly like a real one — plausibility is the only thing the model is optimizing for.

The clinical demonstration of this is memorable. When Hussam Alkaissi and Samy McFarlane asked ChatGPT for references on bone metabolism, it produced citations with real-looking PubMed IDs — but every ID pointed to a completely unrelated paper. One “reference” about homocysteine and bone metabolism carried a PMID that actually belonged to a urology paper about titanium surgical staples (Alkaissi & McFarlane, 2023, p. 2). The number was real. The paper was real. The connection was invented.

Researchers at Oxford have started calling this specific failure a confabulation rather than a hallucination: a case where the model is “fluently” wrong and arbitrary, meaning the answer changes if you rerun it with a different random seed (Farquhar et al., 2024, p. 625). That arbitrariness is the tell. Ask twice and a confabulated citation often comes back with a slightly different title or year — because there was never a fact underneath it, only a probability distribution. This is closely related to how AI-detection tools misfire: both problems come from treating fluent, confident text as evidence of anything.

Does GPT-5 fix it? Partly.

The honest update is that newer models fabricate meaningfully less — but the headline numbers come with a condition almost nobody quotes.

OpenAI’s own GPT-5 system card reports that gpt-5-main produces a claim-level hallucination rate 26% lower than GPT-4o, and gpt-5-thinking 65% lower than o3, with 44% fewer responses containing at least one major factual error (OpenAI, 2025, p. 1). Those are real improvements, measured by grading individual factual claims rather than whole answers.

Here is the condition: that evaluation is titled Factuality on ChatGPT Production Traffic (Browsing Enabled). The model was allowed to look things up. Strip the retrieval away and the picture changes completely — on SimpleQA, a benchmark of deliberately hard factual questions answered from the model’s own parameters, later flagship models still miss roughly half the time.

That gap is the whole lesson of this article compressed into one benchmark. The improvement is not the model getting a better memory for bibliographies; it is the model being handed a way to check. Which is also why the fix for fabricated citations is architectural rather than a matter of asking more nicely.

There is also a trap worth naming. Asking the model to check its own references helps, but not as much as it looks. When Safran and Çalı prompted ChatGPT to verify and revise its citations, the fully-accurate share jumped from 7.5% to 77.5% (Safran & Çalı, 2025, p. 695) — a real gain, but one that still leaves roughly a fifth of the list wrong, now wearing a fresh coat of confidence. Emsley described the darker version of this: when he confronted ChatGPT about its false references, it “doubled down.” Self-verification narrows the problem; it does not close it.

What to actually do

The practical takeaway is not “never use it.” It is that every citation a model gives you is a claim to be checked, not a fact to be trusted — closer to a lead than a source. Three habits cover most of the risk:

  • Verify existence before anything else. Before you worry about whether a paper supports your point, confirm the paper exists. Paste the reference list into a tool like cytado.com’s bibliography checker, which tells you for free whether each entry resolves to a real work, then chase the survivors by DOI. This is exactly the workflow our fact-checking guides are built around.
  • Distrust the DOI most. The Walters and Wilder data is blunt on this: numeric fields — DOIs, page numbers, volumes, years — are where even real citations go wrong. A clickable DOI that lands on an unrelated paper is the single most common trap.
  • Prompt for retrieval, not recall. A model asked to remember a citation confabulates; a model connected to a real database and asked to retrieve one does far better. If your tool can search the web or a bibliographic API, that grounding is worth more than any wording trick, though smart prompting still helps you get verifiable output rather than fluent guesses.

So — does ChatGPT make up sources? Yes. Between roughly 18% and 55% of the time depending on which version you asked, and up to 88% once the question turns to case law. The number keeps falling, and it has never once been zero. Treat every reference as unverified until a database says otherwise, and the model becomes a genuinely useful drafting partner instead of a very confident liar.

Sources

Every factual claim above is tied to a source you can open and check, with page numbers wherever the source has them.

  1. 1. William H. Walters, Esther Isabelle Wilder, Fabrication and errors in the bibliographic citations generated by ChatGPT , Scientific Reports, 2023 , pp. 1, 2, 5, 6. 10.1038/s41598-023-41032-5
  2. 2. OpenAI, GPT-5 System Card , OpenAI, 2025 , p. 1. link
  3. 3. Mikaël Chelli, Jules Descamps, Vincent Lavoué, et al., Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis , Journal of Medical Internet Research, 2024 , pp. 7, 9. 10.2196/53164
  4. 4. Ertuğrul Safran, Adem Çalı, Fabricated or accurate? Ethical concerns and citation hallucination in AI-generated scientific writing on musculoskeletal topics , Anatolian Current Medical Journal, 2025 , p. 695. 10.38053/acmj.1746227
  5. 5. Hussam Alkaissi, Samy I. McFarlane, Artificial Hallucinations in ChatGPT: Implications in Scientific Writing , Cureus, 2023 , p. 2. 10.7759/cureus.35179
  6. 6. Robin A. Emsley, ChatGPT: these are not hallucinations – they're fabrications and falsifications , Schizophrenia (npj), 2023 . 10.1038/s41537-023-00379-4
  7. 7. Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, Yarin Gal, Detecting hallucinations in large language models using semantic entropy , Nature, 2024 , p. 625. 10.1038/s41586-024-07421-0
  8. 8. Stefan Szeider, Unmediated AI-Assisted Scholarly Citations , Open Conference Proceedings (AAAI-26), 2025 , p. 2. 10.52825/ocp.v8i.3161
  9. 9. Nipun Misra, Vikranth Udandarao, Detecting Citation Hallucinations in Large Language Model Outputs , Proceedings of the AAAI Conference on Artificial Intelligence, 2026 . 10.1609/aaai.v40i48.42257
More on Fact-Checking AI Output
Built by Sitario.com