
paper-preflight
io.github.amos689v0.2.0Updated Oct 4, 2026
Checks every LaTeX reference against Crossref, dblp, arXiv, DataCite, PubMed and OpenAlex. No LLM.
Overview
Checks a LaTeX paper's references against scholarly registries to flag nonexistent, mismatched, retracted, or superseded citations.
- What it does
- Reads .tex and .bib files (or .bbl, plain-text reference lists, PDFs, or an arXiv ID) and queries Crossref, dblp, arXiv, DataCite, PubMed and OpenAlex about every cited work. It reports whether a work exists, whether authors, title, year or venue match, whether it was retracted or corrected, and whether a cited preprint has since been published. It can also fetch verified BibTeX by DOI, arXiv ID or title, and propose bibliography fixes as a diff. Verdicts are rule-based, with no LLM involved.
- When to use it
- Useful before submitting or revising a paper, when you want citations verified against real records rather than trusting memory or copy-pasted BibTeX. Also suited to coding agents that should check references before declaring a paper finished, and to pre-commit or CI checks on citation keys.
- Requirements
- Runs locally as a Python package via uvx or pip; no account is required. Optional environment variables PAPER_PREFLIGHT_EMAIL, OPENALEX_API_KEY and S2_API_KEY improve speed and coverage. Network access to the scholarly APIs is needed unless running with --offline against the local cache.
Installation
In SourceWeft
- Open paper-preflight in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
paper-preflight
English · 简体中文
[CI] [License: MIT] [Python 3.11–3.14]
Check every reference of a LaTeX paper against real scholarly records before you submit. No LLM guessing, no false accusations.
Language models invent references, and copy-pasted BibTeX carries wrong years, wrong authors
and dead DOIs. paper-preflight reads your .tex and .bib files and asks Crossref, dblp,
arXiv, DataCite, PubMed and OpenAlex (and Semantic Scholar, if you have a key) about every cited
work:
- Does it exist?
- Does it match what you wrote?
- Has it been retracted?
- Has the preprint you cite been published since?
When it cannot tell, it says so instead of guessing.
Status: v0.2, an early release. False positives are the bugs we most want to hear about: please open an issue.
The repository's demo paper cites eleven works, several of them wrong on purpose. A real run, against the live sources:
Each finding is backed by a record (or by every source answering "no"). The correct NeurIPS paper is verified through dblp even though Crossref only holds fake copies of it, and the two books without identifiers are reported as "cannot determine" instead of "not found".
What it catches
paper-preflight explain REF003 describes any rule.
How accurate is it?
Three measurements, all against the live sources: the bibliographies of real papers, a head-to-head with published tools, and a public benchmark.
On real papers
The bibliographies of 20 arXiv papers from the turn of July and August 2026 (cs, stat, q-bio, quant-ph and astro-ph), chosen mechanically and collected only after every fix in this release, with every warning and error reviewed by hand:
- Fewer than one false alarm per paper (50 references on average), against 66 real problems: 31 errors in the entries (invented co-authors and given names, wrong or malformed DOIs, wrong titles and years) and 35 cited preprints that have since been published.
- The false alarms are mostly names written another way (initials without dots, a generational suffix, a nickname or an English name) and records the registries got wrong (a registry listing 3 of 10 authors, a garbled title). The previous release has 2.1 on these papers.
- Four earlier batches of 20 papers were used to find false positives, each first measured
as it came out (0.1.0: 4.5 per 100 references; 0.1.1: 2.3; 0.1.2 before its last fixes: 3.0;
0.1.2: 1.7). On all four, this release has 0.1 to 0.7. Details in
evals/README.md.
Next to other tools
Badalova & Mayr (2026) checked 104 references by hand and published what five tools flagged. On the same references, with their labels:
The sample is small, so the intervals are wide. Some flags count as false here because the study
labels a reference correct when the work exists: five of paper-preflight's flags on such
references point at real errors (a wrong author, a broken DOI). Two causes of false flags found
in this data were fixed, and four names the study's CSV garbled were restored, before the run
above; the first run measured 62.8%. See
evals/results/badalova-mayr.md.
On a benchmark: HALLMARK
HALLMARK is a public benchmark of real and hallucinated BibTeX entries.
HALLMARK v1.2.3, every entry of both public splits, run on 2026-10-04. Fabrication counts a wrong identifier, a work not found and no author in common; any issue also counts wrong authors, title, year or venue.
- The held-out split confirms the development numbers: the same precision and two points less recall on entries no rule was ever tuned on.
- Every flag on a
dev_publicentry labelled VALID was checked by hand. The 11 that remain are not correct citations: DOIs that belong to other papers, author lists naming people who did not write the paper, a shifted year and a truncated title. - Without them, both modes reach 100% precision and 0% false positives. The list, each item
with a reason one lookup confirms, is in
evals/hallmark_disputed.toml. - What is still missed: invented venues on papers known only as preprints (an arXiv record
cannot contradict a venue) and author lists that merely leave people out. See
evals/results/for every hallucination type.
Precision comes first: a reference is called fabricated only on positive evidence, and an
unanswered or ambiguous lookup is reported as "cannot determine", never as "not found". The
evaluation harness and every run's summary are in evals/.
Quick start
With uv nothing needs installing (or pip install paper-preflight):
path/to/paper is the project directory, its main .tex file, or a single .bib file.
A project that ships no .bib, as many arXiv sources do, is read from its compiled .bbl
(checked, but never edited).
No LaTeX at all? A reference list as plain text works too, in the common styles (APA, IEEE, ACM, Nature, Vancouver, Springer, Elsevier, Chicago, MLA), one reference per line, per paragraph or numbered:
Any arXiv paper, by its ID: the source is downloaded to a temporary folder, checked, and deleted.
Only the PDF? Its reference list is read too, with the pdf extra:
Exit codes:
Fetch verified BibTeX
Instead of writing an entry from memory, ask for it by DOI, arXiv ID or title. Every field comes from the registry record, which a comment above the entry names:
- Preprints: an arXiv preprint that has been published comes back as the published version,
with its
eprintkept (--prefer preprintfor the preprint itself). - Titles:
--title(with--author/--yearif needed) lists the candidates instead of choosing when several works match. - Retractions: a retracted work comes with a warning.
- Agents:
--format jsonis for scripts and agents.
Fix the bibliography
bib fix turns findings into edits of your .bib files, taken from the verified records. It
prints a diff and changes nothing until you add --apply:
--level safe(the default) only fixes what cannot change which work is cited: identifiers written so that links break, and DOIs the registry has but the entry lacks.--level unsafealso rewrites authors, title, year and venue from the record, and removes identifiers that point to another work. Review the diff first.- Only the affected fields change; comments, formatting, line endings and encoding are kept. A reference nobody could find is never "fixed": only you can say what was meant.
Silence a finding you have checked
A comment directly above an entry silences rules for that entry, with an optional reason:
The verdict stays in the JSON report; only the finding is dropped. A suppression that silenced nothing is reported as CFG001 (info), so stale comments do not pile up. Reference rules are only judged after a complete online run, since offline answers and outages may leave them unrun.
Use it from your coding agent
Claude Code — install the plugin. It bundles an MCP server and a skill that makes Claude check the references before calling a paper finished, fix only what is proven wrong, and never invent a reference.
Codex, Cursor, VS Code and other MCP clients — run paper-preflight mcp. The tools are
read-only and confined to your workspace; see docs/mcp.md.
pre-commit — check citation keys and cached verdicts on every commit in seconds; see docs/pre-commit.md.
GitHub Actions — uses: amos689/paper-preflight@main checks the paper on every push, with
the report in the job summary and optional code-scanning alerts; see
docs/github-action.md.
Better results with free credentials
paper-preflight works without any account. These optional environment variables make it faster and more complete; their values are never printed or logged.
paper-preflight doctor shows which are set and whether each source answers right now.
How it works
- Source-first. It reads the LaTeX project as LaTeX sees it: comments,
\iffalseblocks and\includeonlyare respected,.auxfiles are used when they are fresh, and the first definition of a duplicated key wins, as in BibTeX. - Identifier-first routing. DOIs go to their registration agency (doi.org tells which: Crossref, DataCite, …). arXiv IDs go to arXiv, with DataCite as a fallback, and PMIDs and PMCIDs to PubMed (which also marks retracted articles). Entries without identifiers are searched by title in dblp and Crossref.
- Field-by-field matching with guards. It compares titles (including earlier arXiv version titles), authors (tolerating transcriptions such as Reiß/Reis), year and venue. A search result is used only when enough of these agree and no other work fits as well; known fake DOI copies are skipped.
- One verdict per reference: verified, metadata mismatch, identifier conflict, not found, or cannot determine with a reason. "Not found" needs every required source to answer "no".
- No LLM anywhere in the verdict. Answers are cached locally (SQLite), so re-runs are fast
and
--offlineworks.
Design principles
- Positive confirmation or abstain. Rate limits, outages and unindexed works lead to "cannot determine", never to "not found".
- Neutral wording. Findings state observations ("not found in Crossref, dblp and Semantic Scholar, and every source responded"), never accusations.
- Local-first, no telemetry. Only the metadata of the cited works (DOIs, titles, authors) is sent to the public scholarly APIs above. Your manuscript never leaves your machine.
What it will never do
Help evade plagiarism or AI-text detection, scrape paywalled or bot-protected sites, recommend or "complete" references from memory, or name and shame authors.
Roadmap
- Done: releases on PyPI (v0.1); references from a
.bbl, plain text, a PDF or an arXiv ID (v0.2) - Next: fewer false alarms, each round measured on a new week of real papers
- Later: Chinese-language references
Progress is tracked in docs/PROGRESS.md (in Chinese) and the changelog.
Contributing
Bug reports with a reproducible .bib entry are the most valuable contribution, especially
false positives. See CONTRIBUTING.md.
License
MIT. See THIRD_PARTY_NOTICES.md for adapted code.
Source: README.md at commit 2bd03f9
Tools
0Version history
1- v0.2.0LatestOct 4, 2026

