AllCitations logo

AllCitations

How to Check If a Citation Is Real Before You Submit

Nora Ellison··16 min read
APAMLAChicagocitation guideresearch integrity

Try it now - paste a URL, DOI, or ISBN

To check if a citation is real, resolve its identifier first: paste the DOI into https://doi.org/ or look it up at https://api.crossref.org/works/10.xxxx/yyyy. A DOI that returns nothing is the fastest possible answer. But resolution only settles the easiest question, and a reference can pass it while still being wrong in three further ways. This guide works through all four checks, with a real DOI resolved as a worked example, the shapes a fabricated reference takes, and a method for auditing forty references in one sitting.

One framing before the steps. Libor Ansorge's 2026 perspective piece in Frontiers in Research Metrics and Analytics describes the shift the field is making, from existence checking to semantic auditing: confirming that a work exists is the shallow end, and confirming that it says what you cited it for is the deep end. Most checking advice stops at the shallow end. The expensive errors live past it.

The four levels of a citation check

Every reference in your list can fail at a different depth, and each depth takes a different amount of your time.

Level 1, existence. Does the identifier resolve to anything at all? Seconds per reference, and it catches the crudest fabrications.

Level 2, identity. Does the record it resolves to match the reference you wrote? A DOI can resolve perfectly and point at a different paper.

Level 3, findability without an identifier. Books, older articles, theses, and reports often have no DOI, so existence has to be established some other way.

Level 4, semantics. Does the source support the claim you attached to it? This is the one no lookup tool can do for you.

Levels 1 through 3 are mechanical and you can batch them. Level 4 requires reading. The practical consequence is that you should run the mechanical levels across your whole list first, because they are cheap, and reserve your reading time for the references that survive.

Level 1: does the identifier resolve

A DOI is a persistent identifier that resolves through the Handle System to wherever the publisher currently hosts the work, which is why the International DOI Foundation describes DOI names as more than handles: the resolution layer is the point. Two free ways to test one:

Browser: https://doi.org/10.1037/0003-066X.59.1.29 Metadata: https://api.crossref.org/works/10.1037/0003-066X.59.1.29

The first lands you on the publisher's page. The second returns the registered metadata as JSON. Crossref's REST API needs no account and no key, which makes it the faster of the two when you are checking several references and do not want to load a publisher page each time.

A researcher at a sunlit desk sorting printed journal articles into a tabbed pile and an unchecked pile

Here is that DOI resolved. The Crossref record returns the title "How the Mind Hurts and Heals the Body," a single author, Oakley Ray, the container American Psychologist, volume 59, issue 1, pages 29 to 40, published 2004. Formatted in APA 7:

Ray, O. (2004). How the mind hurts and heals the body. American Psychologist, 59(1), 29–40. https://doi.org/10.1037/0003-066X.59.1.29

A fabricated DOI behaves differently. Request one that was never registered and the Crossref API answers with a 404 and no record. There is no partial credit and no ambiguity, which is what makes this check worth running first.

While you have the DOI in front of you, fix its formatting. APA Style's guidance on DOIs and URLs is specific: present the DOI as a hyperlink beginning https://doi.org/, standardise older forms such as doi: or dx.doi.org into that current format, drop "Retrieved from," and add no period after it, because a trailing period can break the link. APA also tells you to copy and paste the DOI rather than retype it, to avoid transcription errors. That instruction does double duty here. A hand-typed DOI is the most common way a real reference acquires a broken identifier and starts looking fabricated.

Level 2: a resolving DOI is not a correct DOI

This is the check almost everyone skips. Existence-checking asks whether the DOI resolves. Correctness-checking asks whether the record it resolves to is the work described in your reference. A DOI pointing at the wrong paper passes the first test and fails the second, and it fails silently, because nothing about a successful resolution tells you it went somewhere unintended.

Compare the resolved record against your entry field by field:

  • Author surnames and initials, and how many authors there are
  • The article title, word for word
  • The journal or book title
  • Year, volume, issue, and page range
  • Publisher, for books and reports

The University of Phoenix library guide on verifying citations from generative AI names the failure mode this produces. A generated citation may name an article that exists in a journal without that journal ever having published it, because the tool combined elements from several sources into one plausible entry. The guide also flags a subtler version, where the body of one article gets attached to an entirely different citation. Both produce references that survive a title search and collapse under a field-by-field comparison.

Two mismatches deserve particular suspicion. A DOI prefix identifies the registrant, so a prefix belonging to one publisher attached to a journal published by another is worth resolving before you trust it. And a year that predates the journal's founding, or a volume number outside the range the journal has reached, means the entry was assembled from parts.

When the resolved record and your reference disagree, the record wins. It is the registered metadata; your entry is a transcription of it.

Level 3: checking a source that has no DOI

Most books, a great many older articles, conference papers, theses, and government reports have no DOI. Existence still has to be established, through a catalogue.

A researcher in a quiet library aisle pulling a bound journal volume from a high shelf

Work down this order, stopping at the first thing that confirms or refutes the entry:

  1. Your library's discovery layer. Northwestern's library guide on evaluating AI-generated content puts this first for a reason: search the title of the article or book in the catalogue, then check your entry against the complete record, including author and dates. Library databases are indexed from publisher deposits, so a match there is strong evidence.
  2. A subject database. PubMed for biomedical work, arXiv for preprints in physics and computing, ERIC for education. These carry records for material that predates DOI registration.
  3. The journal's own archive. If a journal is named, go to its site and open the volume and issue your reference claims. An article either appears in that issue's table of contents or it does not. This is the check that catches the plausible article in the real journal.
  4. The publisher's catalogue, for books. Search the ISBN. An ISBN that returns nothing at the publisher and nothing in a union catalogue is not a real edition.
  5. Google Scholar, last. It indexes widely and it indexes loosely, including preprint mirrors, reading lists, and citation stubs for works nobody has verified. A Scholar hit is a lead to follow, and a Scholar hit alone is weak confirmation.

For a web source that has gone dead, the Internet Archive's Wayback Machine settles whether the page existed at the URL and date you cited. A page that was never archived under a URL you claim to have read is a problem worth resolving before submission.

One honest caveat on the bound volume in the photograph above: physically checking the issue is the most conclusive test on this list and the slowest. Save it for the references your argument actually leans on.

Level 4: does the source say what you cited it for

A reference can clear every check above and still be wrong, because the work exists, the metadata matches, and it simply does not support your sentence. Ansorge's paper is direct about how common this is. Alongside outright fabrication, it describes the mis-citation of real works as the more widespread problem, and identifies two causes: an author who never read the cited work and copied the reference out of secondary literature, and an author padding a reference list to look better read. Both turn citations into decoration.

The method that catches this comes from researchers who review manuscripts. In the r/AskAcademia thread that prompted this guide, a reviewer described going through every reference in a paper they were co-authoring and writing a short note on whether each one was consistent with the claim it was attached to. Their framing was blunt: if your name is on it, you should make sure it makes sense. Another commenter suggested writing that note when you first read the article, keeping all of them in one file with the title and authors, so the work is already done by the time you are drafting. A third pointed out that reference manager note fields do the same job and stay searchable, which is the version worth adopting if you already keep a library in Zotero.

If you have written an annotated bibliography before, you have done this exercise under a different name. The annotation is the note, and it doubles as a semantic check.

Where you genuinely read a claim in one work quoting another, cite it as a secondary source. Every major style has a format for it, and our guide to citing a source you cannot find or access covers the APA, MLA, and Chicago wording.

The four shapes a fabricated reference takes

The Phoenix guide's taxonomy is the most useful one published, because it separates cases that look identical in a reference list:

Fully fabricated. No such work. Title, authors, journal, and DOI are all generated. Dies at Level 1.

A real list with fabricated entries mixed in. Some references were in the training data and are genuine, so spot-checking three of twenty proves nothing about the other seventeen. This is the case that makes people over-trust a list.

Real journal, invented article. The journal exists and publishes in the field. The specific article was never published there. Survives Level 1 if a DOI was borrowed, and dies at Level 2 or at the journal's own archive.

Real citation, wrong content. The reference is entirely genuine and the claim attached to it is not something the work says. Only Level 4 catches this.

Two tells worth learning. Page ranges that are suspiciously round, such as 100 to 110, occur less often in real journals than generated ones. And an author who is real but works in a different field, credited with an article in your field, usually means a name was borrowed to make an entry look solid.

Auditing a whole reference list in one sitting

For a list of thirty to forty references, the mechanical levels take well under an hour if you stop doing them one reference at a time.

Overhead view of blank index cards laid out in three sorted columns on a desk, with one card set apart

Sort first. Split the list into entries that carry a DOI, entries with an ISBN, and everything else. The first group is a batch job, the second is a catalogue search, and the third needs judgment.

Run the DOI group through resolution in one pass, marking each entry verified, unresolved, or mismatched. Keep those three states separate, because they mean different things: unresolved may be a typo, while mismatched is either a wrong DOI or a fabricated entry, and the repair differs.

Work the ISBN group next, then the remainder by hand in the Level 3 order.

Finally, read the marked-up list and ask which references carry real argumentative weight. Those get Level 4. A reference supporting your central claim deserves the fifteen minutes; one supporting a passing observation in your introduction probably does not.

Two cautions on automated checkers, both of which came up in the same reviewer discussion. Uploading an unpublished manuscript to a third-party citation checker sends confidential work to an outside service, which is a genuine problem for anything under review and a matter your institution may have rules about. And AI-text detectors are a poor substitute for checking references, because their accuracy is contested and, as one commenter put it, the provenance of the text is beside the point: what matters is whether the references are real, which you can determine directly.

Once the list is clean, our five-pass audit for citation errors handles the separate problem of formatting, matching every in-text citation to its entry and repairing element order.

What reviewers do when they find one

Worth knowing, because it calibrates how much checking is enough. In the r/AskAcademia thread, the response to a manuscript with fabricated references was near-unanimous: reject and tell the editor why. One reviewer offered the sentence they would write, that half the references not existing casts major doubt on the accuracy of any statement in the manuscript. Several said they now check references before reading anything else, precisely because it is fast and decisive.

Two things follow for anyone submitting work. Checking is no longer a courtesy, since the people receiving your manuscript are doing it themselves and doing it first. And the damage from one fabricated reference is out of proportion to its size, because it transfers doubt onto everything else you wrote.

None of that requires treating your reference list as a hazard. It is a mechanical task with a known method, and the method above is most of an hour.

Common mistakes and how to avoid them

  • Checking that the title exists and stopping there. A title search confirms a work exists somewhere; it says nothing about whether the journal, year, volume, and pages in your entry belong to it. Compare every field against the resolved record.
  • Trusting a DOI because it resolved. Resolution proves registration. Open what it resolved to and read the title.
  • Spot-checking a few entries and generalising. Mixed lists are the common case, so a clean sample tells you nothing about the rest. Run the cheap levels across everything.
  • Retyping a DOI by hand. APA asks you to copy and paste for exactly this reason. One transposed character turns a real reference into an unresolvable one and costs you an hour of confused searching.
  • Treating Google Scholar as confirmation. It indexes reading lists and citation stubs, so a hit can reflect someone else's unverified reference rather than a published work.
  • Leaving Level 4 until after submission. The semantic check is the only one that requires reading, so it is the one that gets deferred, and it is the one that produces the errors a reader in your field will notice.

Quick-reference table

LevelQuestionHow to checkTime per reference
1. ExistenceDoes the identifier resolve?doi.org or api.crossref.org/works/{doi}Seconds
2. IdentityDoes the record match the entry?Compare authors, title, journal, year, volume, pagesUnder a minute
3. No identifierDoes the work exist at all?Library discovery, subject database, journal archive, ISBN1 to 5 minutes
4. SemanticsDoes it support the claim?Read the relevant section and write a one-line note10 to 20 minutes

Frequently Asked Questions

Try AllCitations for Free

No account required. Generate your first citation in seconds.

Start Citing for Free