The first panic about AI in manuscript submissions involved images. Then it was invented references. Publishers are still catching up to both, but the newest integrity problem is quieter: large language models are generating plausible, strategically padded reference lists, and every unresolved DOI quietly reshapes a researcher's citation record.
STM, the international trade association for scientific, technical and medical publishers, has spent the past several years building infrastructure for exactly this kind of manipulation. The association's Integrity Hub shares signals across participating publishers, and its United2Act initiative coordinates action against paper mills. The latest integrity discussions among STM members have moved past catching individual bad actors and toward treating reference lists as a data stream that can be gamed like any other.
Citation padding is old; industrialising it is new
Citation stacking and citation rings predate generative AI by decades. Editors have been caught requesting self-citations. Groups of authors have traded citations to inflate h-indexes. What changed is the unit cost. A chatbot can produce twenty references in seconds, mixing real authors, real journals, plausible but incorrect DOIs, and citation counts that look organic but are not. Some outputs are pure fabrications. A more dangerous subset includes real articles attributed to the wrong author or wrong year, which defeats a simple DOI lookup because the DOI exists and points somewhere.
The same failure mode reached the courts in 2023, when two New York attorneys were ordered to pay $5,000 in Mata v Avianca after filing briefs containing cases invented by ChatGPT. Scholarly manuscripts rarely produce that instant a sanction. The underlying error is the same: the machine supplies a plausible authority and the human files it without checking. Peer review is not a courtroom, but it is the main gate between an AI-generated bibliography and the permanent record.
Here's the catch
The catch is not that AI occasionally hallucinates a reference. The catch is that a large portion of research evaluation has no reason to notice. A citation is treated as a unit of influence. If a fabricated or coerced citation persists in an indexed paper, a Google Scholar profile, an institutional repository, or a preprint server, downstream aggregators count it. The cited researcher did nothing wrong, yet their metrics improve. The journal's citation profile improves slightly. If enough such citations accumulate, a journal impact factor can move. Everyone with a stake in that number has an incentive not to interrogate the bibliography too closely.
An early-career researcher who manually checks every reference is therefore competing against an automated inflation machine in hiring metrics, grant scoring, and research assessment systems. The integrity problem is a publisher problem and a labour problem: the careful scholar pays in time, while the careless or cynical author pays nothing.
What journals are actually shipping
Publishers have had text-matching tools for years. Crossref's Similarity Check helps editorial offices compare submissions against a large database, but matching text is not the same as validating references. The newer layer is metadata checking: does the DOI exist, does it resolve to the paper described, has the source been retracted, and does the citation fit the claim it supports. Many editorial systems now run those checks at submission, though adoption remains uneven across small society journals and large commercial publishers alike.
The Retraction Watch Database has become a common secondary screen. Some journals now flag citations to retracted papers, though policies vary. The STM Integrity Hub is intended to make those signals portable across publishers, so a journal that identifies a citation ring can share the pattern without waiting for a retraction to be published months later. That cross-publisher coordination is the part of the update that could actually change behaviour, because paper-mill actors spread references across many journals on purpose.
Why this lands on hiring committees
Manuscript integrity looks like an editorial problem until a faculty search or research fellowship committee opens a CV. A candidate with a heavily cited publication record may appear stronger than a colleague with similar output and fewer citations. The automated citation manipulator does not need to touch the candidate's own papers. It only needs to cite them. In high-volume evaluation systems, inflated citation counts become a proxy for visibility, and visibility can outrun the underlying work. The practice is older than AI, but the scale has changed.
The problem is not confined to research-intensive universities. Teaching-focused departments that use citation metrics as one signal in a broader portfolio review are just as exposed, because inflated citations tend to cluster around a small number of heavily authored papers. An administrator relying on a dashboard rather than a close read may never see the discrepancy.
A prior AcademicJobs report on journal indexation made a related point: where a paper appears is not a quality guarantee. Neither is the number of times it is cited. The two problems merge when AI-generated manuscripts pad their bibliographies with citations to real, perhaps irrelevant papers, then those citations feed profiles that hiring committees read in minutes. For anyone applying for permanent academic work, an unverifiable citation record is now a liability even when it looks flattering. The Wiley-Hindawi retractions, covered in an earlier AcademicJobs investigation, showed what happens when paper-mill output spreads before detection catches up.
What to do with a generated reference list
The practical response is unglamorous. It involves checking, not trusting. Authors, reviewers, editors, and research integrity officers should treat any AI-assisted reference list as a draft of claims about the literature, not a finished record.
- Check every DOI against the publisher's page or Crossref, not just the resolution page. A DOI can resolve correctly while pointing to an entirely different article.
- Open the source itself before citing it. A one-sentence summary from a chatbot is not the same as reading the methods, and it is the most common place for a wrong citation to hide.
- Search the Retraction Watch Database for any reference with a retraction history, especially for papers in review articles or systematic reviews.
- Record the use of AI in the cover letter or methods section. Many journals follow COPE guidance that requires disclosure when AI tools suggested, organised, or formatted references.
Using a reference manager creates a second record. When a bibliographic entry lives in Zotero, EndNote, or a similar tool, the source metadata is separate from the prose, which makes it harder for a hallucinated reference to slip through as a pasted block. The extra minute per reference is the real cost of using AI responsibly in manuscript preparation.
The researcher's answer
STM's update matters because it treats citation behaviour as an integrity signal. But tools rarely fix incentive structures on their own. A reference checker catches a fake DOI. It does not catch a real but irrelevant citation, and it does not stop a hiring committee from treating raw citation counts as a proxy for achievement.
The researcher's answer is to stop assuming any reference list is innocent. Read the bibliography the way you read the methods. Check the existence of the source, the accuracy of the claim, and the retraction status. If a machine generates the list, the human obligation to verify it grows, not shrinks. A reference list is a claim about what exists. The publisher can flag the bad DOI. The researcher has to live with a citation record that means something.
Photo by Markus Winkler on Unsplash
