AI Writing Tips

Check an AI Citation in 60 Seconds

· 8 cited sources

You can decide whether an AI-generated citation is real in about a minute. The trick is not speed-reading — it’s checking the fields in the order the errors actually sit. Most people read a citation the way it’s written: author, title, journal, year, and the DOI last, if they look at it at all. That order is exactly backwards. The failures do not cluster where your eye lands first. They cluster in the numeric fields — the DOI, the PubMed ID, the volume, the pages, the year — so a 60-second check should start there and treat the title and author names as the last thing to trust, not the first.

This is not a hunch about how to read carefully. It’s what the measurement studies found when they took thousands of AI-generated references apart field by field. Walters and Wilder, checking 636 citations produced by GPT-3.5 and GPT-4, reported that among the citations pointing to real papers, the incorrect numeric values — volume, issue, page numbers, and years — were the single most common problem, more frequent than errors in authors or titles (Walters & Wilder, 2023, p. 5). An earlier audit they cite is starker: Bhattacharyya and colleagues found that 87% of ChatGPT citations to genuine works carried at least one of seven errors, and errors in the numeric components were “especially common” (Walters & Wilder, 2023, p. 2). If you only have a minute, spend it where the mistakes are.

Step 1 (0–20 seconds): resolve the DOI

Copy the DOI and put https://doi.org/ in front of it. Paste it into a browser. This is the fastest, highest-yield check you can run, because the DOI system gives you an unambiguous binary answer: a registered DOI resolves to the publisher’s page for that exact work, and an unregistered one returns a “DOI NOT FOUND” error. There is no middle state. The identifier either exists in the global registry or it does not.

The registry behind that check is enormous and current. Crossref’s 2026 public data file holds close to 180 million DOI records, and its metadata API fields on the order of a billion requests a month (Crossref, 2026) — which means that if a real, recently published paper has a DOI at all, doi.org knows about it within days of publication. A DOI that returns “not found” is therefore a strong signal, though not yet proof, that the citation was invented. Note the common false alarm: a single transposed digit or letter will make a genuine DOI fail to resolve, so before you condemn a reference, retype the identifier once by hand.

Step 2 (20–35 seconds): check where the DOI lands

Resolving is necessary but not sufficient. The nastier failure mode is a DOI that resolves perfectly — to the wrong paper. This is the confabulation that fools careful readers, because every surface signal says “verified.” The clinical demonstration is now a small classic. When Alkaissi and McFarlane asked ChatGPT for references on bone metabolism, it returned citations with real, resolvable PubMed IDs — but each ID pointed to an unrelated paper. Their example: a fabricated reference titled “Homocysteine and bone metabolism,” carrying PMID 12352394, which when searched actually returns a urology paper about titanium surgical staples in laparoscopic surgery (Alkaissi & McFarlane, 2023, p. 2). The number was real. The paper it named was real. The connection between them was invented.

So the second check is: does the page the DOI opened actually match the title and authors the citation claimed? Robin Emsley, editor of Schizophrenia, described the same trap from the receiving end — entering the DOIs from ChatGPT-generated references “took me to totally unrelated publications” (Emsley, 2023). Fifteen seconds of glancing at the resolved page — is this the right title, the right authors, the right journal — catches the entire class of “real identifier, fake mapping” errors that Step 1 waves through.

When there is no DOI: absence is not proof

Here is where fast checkers get it wrong in the other direction. A missing DOI is not evidence of fabrication, and treating it as such will make you reject real sources. Digital Object Identifiers only launched in 1997 and were adopted gradually, so a large share of legitimate older literature never got one. Boudry and Chartron, studying PubMed-archived articles, found that across 1966–2015 only 40.48% of articles had a DOI, even though the share for articles published in 2015 alone had climbed to 86.42% (Boudry & Chartron, 2017, p. 2). A 1998 paper with no DOI is completely normal. A 2024 paper with no DOI is a little suspicious.

When the DOI is absent, don’t stall — pivot to the title. Paste the exact title into Crossref’s search (search.crossref.org) or, for a biomedical claim, PubMed. If the paper exists, one of them will surface it in seconds, usually with the DOI the model failed to supply. If neither the title nor the author-plus-year combination turns up anything, you’ve learned something: the reference has now failed two independent indexes, which is a much stronger verdict than a missing identifier alone.

Step 3 (35–50 seconds): the PMID, for anything biomedical

If the citation is medical or life-sciences and carries a PubMed ID, check it the same way you checked the DOI: paste the bare number into PubMed. A PMID is narrower than a DOI — it only covers journals PubMed indexes — but within that scope it’s just as binary. Either the ID returns the paper the citation names, or it returns something else, or it returns nothing. The Alkaissi example is the pattern to watch for: the ID resolves to a paper, just not the paper. Many biomedical articles carry both a PMID and a DOI, so when you have both, checking them against each other is a free cross-validation — they must point to the same work.

Step 4 (50–60 seconds): title and authors, checked last for a reason

Only now do you look hard at the title and author list — and the reason they come last is that they are the least reliable indicator of a fake. Fabricated citations are engineered by the training objective to look legitimate at a glance; plausibility is the only property the model optimizes for (Walters & Wilder, 2023, p. 2). A confabulated title reads exactly like a real one, and the authors are frequently real people who genuinely work in the field — just not on that paper. Safran and Çalı, scoring 40 ChatGPT-generated references on musculoskeletal topics, found the inaccurate ones often paired real authors with invented or subtly altered titles, and in their worst prompt 4 of 10 DOIs simply failed to resolve (Safran & Çalı, 2025, p. 698). The title is where the model’s fluency is most convincing and least informative. That’s precisely why it’s the wrong place to start and the right place to finish.

When you have a whole reference list, not one entry

Sixty seconds per citation is fine for the three references you’re about to quote. It does not scale to a 60-item bibliography, and that’s where fabrications hide, because nobody checks all sixty. For a full list, run it through an automated checker that resolves each entry against real indexes rather than re-asking the model that produced it — for example cytado.com’s bibliography checker, which flags entries with no credible match (cytado is owned by this site’s operator, so treat that as a disclosed interest rather than a neutral tip). The one move that never works is asking the chatbot to verify its own output: Emsley found the model “doubled down” when confronted, and the mechanism guarantees it — the same system that invents a citation will vouch for it.

Automation also outperforms manual search on the metadata itself. Szeider built a system that fetches bibliographic entries directly from an authoritative database instead of through the model’s text, and measured an 82.7% perfect-match rate against 28.2% for standard web search, with zero cases of metadata corruption (Szeider, 2025, p. 3). The gap is the whole argument for checking against an index rather than eyeballing: the index doesn’t hallucinate volume numbers.

What 60 seconds cannot catch

Be honest about the ceiling. This method verifies that a citation exists and that its metadata is internally consistent. It does not verify that the paper supports the claim it’s attached to — a real, correctly-cited source can still be pointed at a sentence it never makes. That deeper check takes reading, not a minute. But existence-checking is the right first filter precisely because it’s cheap and it clears the majority of the risk: if the paper doesn’t exist, nothing downstream matters.

It’s also worth knowing why the numbers stay bad even as models improve. Across GPT-3, ChatGPT, and GPT-4, fabrication rates fell from over 70% to under 50%, yet for some publication categories they still exceed 80% (Szeider, 2025, p. 2), and in legal contexts one benchmark measured hallucination rising from 58% with GPT-4 to 88% with an open model (Misra & Udandarao, 2026). Self-correction helps but doesn’t close it: when Safran and Çalı asked ChatGPT to verify and revise its own citations, the fully-accurate share jumped from 7.5% to 77.5% (Safran & Çalı, 2025, p. 695) — a real gain that still leaves roughly a fifth of the list wrong, now wearing fresh confidence. The base rate is high enough that a 60-second check is not paranoia. It’s arithmetic.

For the underlying rates broken down by model and task, our reference table of AI hallucination rates keeps each figure attached to its conditions, and if you want the evidence that fabrication happens at all, the companion piece on whether ChatGPT makes up sources walks through the studies. This check is the front line of the broader workflow our fact-checking guides are built around — and skipping it is exactly what a fabricated citation ends up costing the people who trusted it.

The whole method fits on a sticky note: resolve the DOI, confirm it lands on the right paper, forgive a missing DOI on an old source but not a new one, cross-check the PMID, and read the title last. Start where the errors are, not where the sentence begins.

Sources

Every factual claim above is tied to a source you can open and check, with page numbers wherever the source has them.

  1. 1. William H. Walters, Esther Isabelle Wilder, Fabrication and errors in the bibliographic citations generated by ChatGPT , Scientific Reports, 2023 , pp. 1, 2, 5. 10.1038/s41598-023-41032-5
  2. 2. Hussam Alkaissi, Samy I. McFarlane, Artificial Hallucinations in ChatGPT: Implications in Scientific Writing , Cureus, 2023 , p. 2. 10.7759/cureus.35179
  3. 3. Ertuğrul Safran, Adem Çalı, Fabricated or accurate? Ethical concerns and citation hallucination in AI-generated scientific writing on musculoskeletal topics , Anatolian Current Medical Journal, 2025 , pp. 695, 698. 10.38053/acmj.1746227
  4. 4. Christophe Boudry, Ghislaine Chartron, Availability of digital object identifiers in publications archived by PubMed , Scientometrics, 2017 , p. 2. 10.1007/s11192-016-2225-6
  5. 5. Robin A. Emsley, ChatGPT: these are not hallucinations – they're fabrications and falsifications , Schizophrenia (npj), 2023 . 10.1038/s41537-023-00379-4
  6. 6. Stefan Szeider, Unmediated AI-Assisted Scholarly Citations , Open Conference Proceedings (AAAI-26), 2025 , pp. 2, 3. 10.52825/ocp.v8i.3161
  7. 7. Nipun Misra, Vikranth Udandarao, Detecting Citation Hallucinations in Large Language Model Outputs , Proceedings of the AAAI Conference on Artificial Intelligence, 2026 . 10.1609/aaai.v40i48.42257
  8. 8. Crossref, 2026 public data file now available , Crossref, 2026 . link
More on Fact-Checking AI Output
Built by Sitario.com