Short answer
Break every citation into three tests: does the source exist; do the bibliographic details and link identify that source; and does it actually support the claim beside it? Open the link, compare title, author and date, check the record through a DOI or the issuing institution, and read the relevant passage. Proving that a paper exists is not enough because a real paper can be misrepresented. A formal-looking URL is even weaker: its path, page number and statistic may all have been generated.
Why fabricated references look unusually real
Academic references and web addresses follow strong patterns. Authors, years, titles, journals, volumes, pages and DOI strings have learnable forms. A model can create a perfectly formatted record by combining a real researcher, a plausible title and a realistic identifier into a paper that never existed. It can also identify a real paper while giving the wrong year, author, population or finding.
A study manually checked 636 bibliographic citations generated in literature reviews by GPT-3.5 and GPT-4. In that particular setting, 55 per cent of GPT-3.5 citations and 18 per cent of GPT-4 citations were fabricated; substantial errors also remained in some references to real works. Those rates belong to the tested models, date and task, not to every current system. They nevertheless demonstrate that correct form is not evidence of existence. Walters and Wilder: Fabrication and errors in bibliographic citations generated by ChatGPT
The problem is persuasive because each field can look individually reasonable. A journal operates in the relevant discipline. An author publishes on the topic. A DOI has the expected punctuation. Human recognition of these parts can suppress the final question: does this exact combination resolve to a record?
Gate one: does the source exist?
For scholarly work, start with the DOI. Put the complete identifier after https://doi.org/ and see whether it resolves to a publisher. Search Crossref by title, author or DOI. Crossref's public REST API exposes bibliographic metadata deposited by publishing members and trusted sources, allowing comparison of title, author, journal, publication date and update relationships. Crossref REST API documentation
Absence of a DOI does not prove nonexistence. Books, government reports, legislation, standards, preprints and institutional pages may use ISBNs, report numbers, legislative identifiers or stable URLs. Search the appropriate authority: publisher, library catalogue, government agency, legislation database, standards body or research repository. Do not stop because an ordinary web search repeats the title. Pages copying generated text can create circular confirmation.
Open web links and inspect the final domain after redirects. A link can resemble a real institutional URL but contain a nonexistent path. It may land on a login page, homepage, search results or unrelated document. Record a 404, redirect and page title; do not let blue link styling testify on its behalf.
Gate two: does the metadata match?
Finding a similar paper is not enough. Compare title, principal or complete authors, publication year, journal or institution, version and identifier. One study may exist as a preprint, conference paper, final journal article and corrected version. Authors can share names. A webpage's “updated” date may not be the date the work was first issued.
For a direct quotation, search the original for a distinctive phrase. If it is absent, do not assume only the page number is wrong. A paraphrase may have been placed inside quotation marks, or the wording may be invented. When a similar sentence exists, read its surroundings for negation, scope and speaker.
For data, match the table or dataset version. Statistical agencies revise historical values. A chart may use seasonally adjusted data while the AI describes it as raw. A percentage may represent survey respondents rather than the population. Preserve dataset name, variable, period, geography, unit and access date.
Check corrections and retractions. A DOI resolving successfully proves registration, not that a finding remains accepted. Crossref metadata can include update relationships, and publishers may place correction or retraction notices on the record. The status is part of the evidence.
Gate three: does the source support the adjacent claim?
This gate is most often omitted. AI may cite a genuine privacy study for the proposition that a specific product is illegal, use one occupational experiment to claim productivity rises for every worker, or apply a general legal principle to an individual case. Relevance is not a completed inference.
Rewrite the AI sentence as its smallest claim and find the passage that supports it. Compare strength. Does the source report association while AI says cause? Does it say “in this sample” while AI says everyone? Does it say “may” while AI says “will”? Does the study examine an old model while the answer describes a current product?
If only an abstract is available, verify only what the abstract explicitly covers. Do not pretend to have inspected methods, limitations or appendices. A paywall does not invalidate research, but it limits what you can honestly say you checked.
A useful evidence table has columns for the claim, exact source passage, source type, scope, limitations and your final wording. If the passage column is empty, the citation is decorative rather than evidential.
Trace a statistic to its denominator
“Forty per cent improvement”, “95 per cent accuracy” and “most users” require denominators. Ask: improvement over which baseline; what are the absolute values; how large was the sample; were only completers counted; over what period; what uncertainty was reported; who funded the study; and does the metric represent the outcome you care about?
Faster processing is not necessarily lower total cost if review and rework increase. Higher average accuracy does not show that every group improved. AI often selects the most striking figure and omits the measurement design that gives it meaning.
Recalculate one load-bearing percentage from the source table. Failure to reproduce may reflect rounding, a different subgroup, a unit error or invention. Whatever the cause, it should not be published unresolved.
When a number is derived across several sources, preserve the calculation. Mixing a population estimate from one year with a rate from another can create a precise answer that no source actually reports.
A link is not evidence, and citation volume is not quality
Ten pages copying one another are one information path. A vendor blog may accurately document its own feature but is not independent evaluation of its social effect. News can help discover an event, but consequential facts should lead to the judgment, announcement, dataset or paper. Wikipedia can provide a map; important claims should follow its references to source material.
Match source type to claim. Use official text for law and policy; official documentation plus reproducible tests for product behaviour; original studies and data for scientific findings; methodologically transparent statistical sources for market figures; and personal accounts only for that person's experience.
Beware a “citation halo” in which supported background makes an unsupported conclusion look sourced. Put citations immediately beside the claim they support and split sentences containing multiple propositions.
Handle AI-generated reading lists as candidates
Do not request fifty references and then rescue them one by one. Define the search scope and inclusion criteria, ask for a small candidate set, and request a DOI or stable identifier, original abstract and reason for relevance. Independently search Crossref, PubMed, the publisher or a field-specific database.
After confirming existence, build an evidence table: question, sample, method, primary result, limitations and direct relevance. Mark “full text not obtained”, “not yet checked” and “background only”. AI can assist discovery without disguising a candidate list as a completed review.
If no stable identifier is available, exact title and author may still work. When several consecutive items cannot be found, stop using the batch rather than assuming the remainder are reliable. Batch generation can produce batch-correlated errors.
For a systematic or professional review, preserve the actual search query, database, date and screening rule. A model's unexplained selection cannot establish that relevant contrary evidence was considered.
Check versions of web and official documents
Government and product pages change. The same URL may display different text next month. For important evidence, record access date, document version and issue date, and prefer a permanent PDF, notice number or archived version. For law, distinguish passage, commencement, amendment and repeal.
Product documentation needs account tier, region, platform and version. “The service supports this function” may apply only to enterprise accounts or a preview. A product homepage cannot support those boundaries; cite the specific documentation.
If a source is living documentation, quote minimally and summarise with a date. Long copied passages become stale and raise copyright concerns, while a precise description plus link is easier to maintain.
A five-step citation audit
Step one, open: does the link resolve, and what are the final domain and title? Step two, match: do author, date, version and identifier agree? Step three, locate: where is the supporting paragraph, table or page? Step four, narrow: reduce the AI wording to the strength and scope the source supports. Step five, record: preserve verification status, access date and unresolved questions.
Consequential material may warrant a second reviewer for load-bearing sources. Automated checks can test URL status and DOI metadata. They cannot reliably decide whether a source supports a complex inference; that semantic gate still requires understanding both the source and the intended use.
My assessment: AI is a lead generator, not a bibliographic authority
AI is useful for proposing search language, relevant institutions, research directions and candidate works. It accelerates where to begin, not what ultimately deserves citation. Every source entering a formal article, report or decision should pass an independent existence check and a support check.
Be most cautious with a reference that fits too perfectly: a title that repeats the exact question, a figure that supports the desired conclusion, and a long professional-looking link. Real research is commonly narrower and conditional. Perfect fit can indicate excellent retrieval, but it can also indicate evidence generated to order.
Pre-publication checklist
- Does the link open the right document rather than a homepage or search page?
- Do DOI, title, author, year, journal and institution agree?
- Have preprint, final version, correction and retraction been distinguished?
- Which exact passage or table supports the adjacent claim?
- Has association become causation, or a limited sample become everyone?
- What are the figure's denominator, unit, sample, period and geography?
- Can at least one load-bearing number be recalculated independently?
- Is the source primary, current and appropriate for the claim type?
- Are inaccessible full text and unverified items clearly marked?
Conclusion
Checking an AI citation requires three levels: existence, metadata and support. A real link may not support the claim, a correct title may be paired with the wrong result, and an impressive figure may lack the relevant denominator. Let AI discover leads, but let DOI records, official documents, source passages and independent calculations determine what can be cited. The value of a reference is not its presence. It is whether the reader can follow it back to evidence.
Related questions
- Why Can AI Sound Confident and Still Be Completely Wrong?
- What Does AI Most Often Miss When Summarising Long Documents?
Continue reading: All articles in How Far Should You Trust AI?
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.