AI Writing Tips
Fact-checking

The model is confident and wrong more often than you think.

Fabricated references are not an edge case — they are a measured, reproducible failure mode. This is the pillar where we put numbers on it and show you how to check.

Guides

Everything in this pillar

Verification workflows, model-by-model error rates, and what to do when a source does not resolve.

What the research shows

Four findings worth knowing before you trust a draft

Each of these comes from a peer-reviewed study, cited in full on the post that covers it.

Most citations were wrong

In a controlled evaluation of AI-generated scientific writing, the majority of references were either partially incorrect or entirely fabricated — plausible authors, real-looking journals, invalid DOIs.

Rates vary wildly by model

Citation hallucination runs from roughly 18% in GPT-4 to over 70% in other frontier models. In legal questions about federal cases, published rates reached 58–88%.

Follow-up prompts genuinely help

Asking the model to verify its own bibliography converts partially correct references into accurate ones at a measurable rate. It does not fix everything — but it is not folklore either.

Uncertainty can be measured

Entropy-based methods published in Nature detect a subset of hallucinations — confabulations — by measuring uncertainty at the level of meaning rather than wording.

Common questions

Questions people actually ask

Does ChatGPT make up sources?

Yes, measurably and repeatedly. The phenomenon has a name in the literature — citation hallucination — and published rates range from about 18% to over 70% depending on the model. The references usually look right: real author names, plausible journals, well-formed DOIs that resolve to nothing or to a different paper.

Can I just ask the model whether its sources are real?

Partly. Structured follow-up prompts that explicitly ask the model to verify bibliographic accuracy do improve reference validity — that effect is documented. But the model cannot reliably check what it never had access to, so a follow-up prompt narrows the problem rather than removing it. Resolve the DOI yourself.

What is the fastest reliable check?

Resolve every DOI and confirm the title, first author and year match. Studies that verified AI-generated references against Crossref, PubMed and Google Scholar found that mismatches cluster in exactly those three fields, which makes them a cheap, high-yield check.

Built by Sitario.com