The most consequential number in academic publishing right now is not an impact factor. It is a retraction queue that one publisher still cannot close out. Wiley acquired the open-access publisher Hindawi for $298 million in 2021. By the time Wiley retired the Hindawi name and said it would close 19 Hindawi journals in 2024, thousands of papers had been pulled after investigations found paper mill activity, fake peer review rings, fabricated data, and manipulated images. Across the industry, 2023 set a record with more than 10,000 retracted research papers.
That number did not happen because researchers suddenly lost their ethics. It happened because the industrial production of fake science got cheaper and faster. Generative AI produced passable abstracts and introductions at scale. Paper mills packaged those words with manufactured data and sold authorship slots. Journals that charge article processing charges (APCs) on acceptance made volume profitable. Peer review, the filter meant to catch the fraud, still depends on unpaid academic labour and the assumption that a submission is genuine until someone proves otherwise.
The scale matters because peer review was never designed to authenticate authorship. It was designed to judge methods and arguments. A fabricated manuscript written by a coordinated team can look cleaner than a genuine paper by a tired postdoc. Reviewers cannot be expected to spot that difference in an afternoon, and editors cannot investigate every suspicious flag without budgets they do not have.
Paper mills have worked the same way for decades. A broker recruits a researcher who needs publications for promotion or professional advancement. In a typical operation, the broker arranges a ghostwritten manuscript, supplies or fabricates a dataset, submits the paper, and sometimes suggests reviewers whose email addresses route back to the mill. The reviews come back favourable. The journal collects its APC. The client gets a clean-looking publication.
Generative AI did not invent this business, but it removed its biggest production cost. Cheap language generation means mills can flood journals with text that no longer trips grammar or plagiarism checks. Submission volume rises. Editors and reviewers, usually faculty members working without pay, see more manuscripts that look plausible but are hollow. The system does not catch the bad papers; it burns out the people who might.
Guest-edited special issues are a common entry point. A mill can recruit or impersonate a guest editor, fill an issue with coordinated submissions, exploit a journal's need to publish on schedule, and hide behind fake credentials. Several publishers have suspended special issue programmes after exactly this pattern appeared across their portfolios.
Here's the catch
The catch is not that detection software does not exist. The catch is that every detection tool creates a new cost, and nobody has agreed who bears it. Publishers have built shared screening through the STM Integrity Hub, run by the International Association of Scientific, Technical and Medical Publishers, which flags duplicate submissions and paper mill markers across participating journals. Springer Nature runs its own software, Geppetto, to spot AI-generated text and image manipulation.
Those tools are triage, not proof. A flag still needs a human investigation, and human investigators are exactly the resource in short supply. False positives pile up too. Formal academic prose written by a non-native English speaker can be flagged as machine-generated. Every false flag consumes editor time and risks alienating a genuine author. A paper mill, meanwhile, can route its text through a paraphraser and try again. The machines get better, but the mills get better in response. That is the arms race.
Photo by Brett Jordan on Unsplash
The detection arms race, graded
Hype: AI detectors and integrity software can clean up peer review. Reality: they can identify patterns, but not intent. The distinction matters. COPE, the Committee on Publication Ethics, has been clear that software signals alone are not proof of paper mill involvement; publishers must investigate before acting. What the tools actually flag tends to be narrower than the marketing suggests.
- Duplicate images across supposedly unrelated papers
- Author networks that recur across different journals with no plausible collaboration
- Datasets that do not match the experiments described in the manuscript
- Reviewer email addresses that trace back to the same free providers or IP addresses as the authors
Each of those signals can appear in legitimate work. A real author can reuse a figure by mistake, share a network with a co-author, choose a generic email address, or submit from the same campus network as a colleague. The software's value comes from how quickly it narrows the pool for editors. It does not replace editorial judgement, and it cannot stop a determined mill that varies its submission patterns.
Who pays for the cleanup
The money is the part universities notice. A paper mill submission does not just waste reviewer time; it can end up in a researcher's tenure file, a grant report, or a national research assessment. Retractions often arrive years later, after the promotion is approved and the institution has already advertised the author's output. Research integrity offices then spend months reconstructing what happened. The author may claim ignorance, the middleman has vanished, the journal has moved on, and the institution is left holding the file.
Across campuses, the policy response has splintered. Some universities now screen publication lists during faculty hiring; others rely on search committees with no training in forensic publication checks. This site has tracked the same fragmentation in AI academic integrity rules. Without a shared standard, the same problematic paper can sink one candidate and slip past another committee. The inconsistency is its own kind of crisis.
The publisher problem nobody wants to name
Publishers have an awkward incentive. Journals that charge APCs on acceptance earn money when papers pass review. Journals that run special issues, particularly guest-edited collections, have been hit hard by mills because guest editors can be recruited or impersonated. The pressure to fill pages collides with the duty to reject. Some publishers have closed special issue programmes or suspended journals after fraud. Others have not.
Tracking the scale is not hard. Retraction Watch maintains a public database that lists tens of thousands of retracted papers, many tied to paper mills. The database has become a working tool for journalists, research integrity officers, hiring committees, and tenure review panels. The fact that such a database is necessary at all tells you how far the verification gap has grown.
Photo by Brett Jordan on Unsplash
The personnel problem behind the software fix
The fix is not a smarter detector. It is a change in who checks, and when. If publication lists were verified at hiring, promotion, grant application, and annual review time, the market for fabricated papers would shrink because the credential would lose value. Data and code could be required before review, not after. Editors could be given the time, status, budget, and training to investigate suspicious submissions instead of hoping a volunteer reviewer notices. None of that requires a breakthrough in AI. It requires universities and funders to treat authorship verification as part of the job.
Anna Abalkina, a sociologist at the Free University of Berlin who has documented paper mill networks across multiple countries, has shown repeatedly that these networks respond to publication pressure, not to software. Her work points to the same conclusion: the retraction queue will keep growing until the incentives that feed it are dismantled. That is a personnel problem, not an algorithm problem.
