The model is confident and wrong more often than you think.
Fabricated references are not an edge case — they are a measured, reproducible failure mode. This is the pillar where we put numbers on it and show you how to check.
Yes — a follow-up verify prompt measurably improves citation accuracy (7.5% to 77.5% in one study). No — self-checking still leaves a fifth wrong, because a model can't reliably audit itself.
A fake citation isn't a botched copy of a real record — the model never had one. It assembles the most probable author, title, and journal, so plausibility is baked in.
A sourced table of measured AI hallucination rates by model, task, and domain — with the methodology behind each figure kept intact, not averaged away.
A minute is enough to catch a fabricated citation — if you check the fields in the order the errors actually sit. Start with the DOI and PMID, not the title.
A made-up citation now has a price you can read off a docket: $5,000 in Mata v. Avianca, up to $109,700 by 2026 — plus a quieter cost when no one catches it.
Yes, ChatGPT fabricates citations. Five peer-reviewed studies measured how often: 18% for GPT-4, 55% for GPT-3.5, and up to 88% for legal questions.
9 cited sources
What the research shows
Four findings worth knowing before you trust a draft
Each of these comes from a peer-reviewed study, cited in full on the post that covers it.
Most citations were wrong
In a controlled evaluation of AI-generated scientific writing, the majority of references were either partially incorrect or entirely fabricated — plausible authors, real-looking journals, invalid DOIs.
Rates vary wildly by model
Citation hallucination runs from roughly 18% in GPT-4 to over 70% in other frontier models. In legal questions about federal cases, published rates reached 58–88%.
Follow-up prompts genuinely help
Asking the model to verify its own bibliography converts partially correct references into accurate ones at a measurable rate. It does not fix everything — but it is not folklore either.
Uncertainty can be measured
Entropy-based methods published in Nature detect a subset of hallucinations — confabulations — by measuring uncertainty at the level of meaning rather than wording.
Common questions
Questions people actually ask
Does ChatGPT make up sources?
Yes, measurably and repeatedly. The phenomenon has a name in the literature — citation hallucination — and published rates range from about 18% to over 70% depending on the model. The references usually look right: real author names, plausible journals, well-formed DOIs that resolve to nothing or to a different paper.
Can I just ask the model whether its sources are real?
Partly. Structured follow-up prompts that explicitly ask the model to verify bibliographic accuracy do improve reference validity — that effect is documented. But the model cannot reliably check what it never had access to, so a follow-up prompt narrows the problem rather than removing it. Resolve the DOI yourself.
What is the fastest reliable check?
Resolve every DOI and confirm the title, first author and year match. Studies that verified AI-generated references against Crossref, PubMed and Google Scholar found that mismatches cluster in exactly those three fields, which makes them a cheap, high-yield check.