Academic Jobs - Home of Higher Ed Logo

Generative AI in Referee Reports: Journals Tighten Peer Review Rules

Post a Story
48views
white and black typewriter with white printer paper
Photo by Markus Winkler on Unsplash

The race to put generative AI into scientific publishing has produced a policy scramble no one expected. Not over authors. Over referees.

Through 2023 and 2024, major journals and research funders quietly changed their rules about what a peer reviewer may do with a large language model. The U.S. National Institutes of Health, Elsevier, Springer Nature, Wiley, the JAMA Network and the Committee on Publication Ethics all published or updated guidance. The shared instruction: do not paste an unpublished manuscript into a public AI tool, and do not let one write the referee report.

It sounds like a narrow administrative fix. It isn't. Peer review remains the one part of journal publishing where confidentiality and expert judgment are the entire product. A weak paper can be corrected after publication. A leaked manuscript or a hallucinated critique cannot be recalled.

Generative AI systems such as ChatGPT, Claude and Gemini now sit inside everyday research workflows for drafting, coding, summarising grant text and preparing course materials. Publishers spent years encouraging those uses. A referee report is different. It requires a specific assessment of unpublished data, methods and interpretation under a confidentiality agreement. Once a manuscript enters a model with unclear retention settings, that bargain breaks.

Here's the catch

The rules are simple to publish and difficult to police. Editors see the final report, not the drafting process. A reviewer can run a private model on a local machine, paste a manuscript in sections, lean on an institutional tool with its own data terms, or ask a colleague to summarise the submission, and no editor will notice.

Automated detection doesn't close the gap. Text classifiers remain unreliable on short, technical, person-specific prose, and a false accusation against a referee is institutionally costly in a system that already struggles to find reviewers.

The policies also split on the most basic question: disclosure or prohibition. Some publishers allow generative AI if its use is declared. Others ban it outright. A few funders prohibit any AI assistance in grant review. A single reviewer may face three different standards in the same week.

That is the quiet failure inside the policy wave. The announcements demonstrate intent, but enforcement still depends on the honour system.

What the major players actually changed

In June 2023, the U.S. National Institutes of Health barred reviewers from using generative AI to analyse or critique grant applications, warning that such tools could compromise confidentiality and the integrity of the process. The NIH position is prohibition, not disclosure.

Elsevier's reviewer guidance tells referees not to use generative AI or AI-assisted tools to assist in the scientific review of a manuscript. The publisher's policy centres on confidentiality, personal accountability, originality and disclosure, and asks reviewers to keep submissions out of unapproved systems.

Springer Nature updated its editorial policies to say reviewers should not upload unpublished work to generative AI services and should not use the tools to draft their reports. If AI is used in an acceptable way, it must be disclosed to the editor.

The JAMA Network and journals published by the American Association for the Advancement of Science have issued comparable restrictions on entering manuscript text into AI systems. COPE, the Committee on Publication Ethics, keeps the burden on humans: the listed reviewer owns the report, and any AI involvement must be transparent without removing accountability.

black and white typewriter on white table

Photo by Markus Winkler on Unsplash

What the prohibitions cover

Rules vary, but most new policies cluster around a small set of actions.

  • Uploading an unpublished manuscript, abstract, figure or reviewer comment to a public generative AI tool.
  • Using a language model's output as the basis of a referee report without reading and verifying the underlying paper.
  • Delegating judgments about novelty, rigour, significance or reproducibility to software.
  • Relying on AI-generated literature summaries that may fabricate or misattribute references.

The last point isn't hypothetical. Generative AI has a documented habit of inventing citations, and publishers worry those fabrications will migrate into referee reports. The earlier STM update on AI citation-stuffing in manuscripts described a similar pressure on the author side.

What usually remains allowed is the unglamorous work: grammar correction in a tool with no training access, basic reference formatting, summarising public information before the manuscript arrives and manual proofreading. The dividing line is less the technology than what the technology sees.

Hype versus reality

Reality grade: the bans are meaningful as a norm-setting exercise and weak as an enforcement mechanism. They will stop casual, unintentional leaks. They won't stop a determined referee with a local model or a university AI platform covered by a data-processing agreement.

Advocates for generative AI in peer review point to several legitimate uses. It can help a non-native English speaker polish a report. It can flag statistical inconsistencies at speed. It can summarise a reviewer's own notes. None of those require exposing the manuscript to a consumer chatbot.

The publishers know this. Their policies are not anti-technology so much as anti-convenience. The problem is that convenience is precisely what drives adoption, and journals have built no clear path for safe AI-assisted review at scale.

Indexation has never been a guarantee of review quality, as we've noted in reporting on journal indexation. The new AI rules are another reminder that policies on paper are only as strong as the editorial labour behind them.

Verdict: the policy wave is real and the enforcement gap is large. Clarity varies from journal to journal. Assume the strictest standard applies until an editor says otherwise.

What early-career referees should do

Peer review requests often arrive at the worst possible time for early-career researchers: during thesis writing, between experiments, in a job search and often without pay. The new rules add another layer. A careless upload can compromise a manuscript and your standing with an editor.

Before accepting a review, read the journal's generative AI policy. If you plan to use any AI tool, even for grammar, disclose it to the editor in a short note. If the policy is silent, ask.

Keep the manuscript offline. Use reference managers and statistical software that don't require uploading the submission. If you need language help, use a colleague or a paid human editor under confidentiality terms.

Editors are not your adversaries here. The worst outcome isn't being told no; it's being asked later whether a reviewer report that carries your name was written by software you didn't disclose.

For researchers applying to positions where peer review service counts, a documented record of careful, confidential review remains a credential. The new rules don't devalue that. They make the credential harder to fake.

The most pointed warning comes from the reviewers themselves. Every new policy shifts the burden sideways, not upward. A confidentiality rule only works if the person holding the manuscript refuses the shortcut. That is a professional norm senior researchers can teach, journals can reinforce, but no terms-of-use can install.

Portrait of Dr. Oliver Fenton
About the author

Dr. Oliver FentonView author

Academic Jobs In House Author

Discussion

Sort by:

Be the first to comment on this article!

You

Youโ€™ll be asked to sign in before your comment is posted.

New0 comments

Join the conversation!

Add your comments now!

Have your say

Engagement level

Browse by Faculty

Browse by Subject

Frequently Asked Questions

๐Ÿ”Can peer reviewers use ChatGPT to write referee reports?

Most major publishers now say no. The U.S. National Institutes of Health prohibits peer reviewers from using generative AI to analyse or critique grant applications. Elsevier and Springer Nature instruct reviewers not to use AI tools to draft their reports. If a journal permits any use, it usually requires disclosure to the editor first.

๐Ÿ”’Why did journals update policies on generative AI in peer review?

Confidentiality is the main reason. Uploading an unpublished manuscript to a public model may expose it to training data or retained prompts. Hallucinated references and generic critiques also undermine accountability, so publishers want the human reviewer to own the report.

โš–๏ธWhat does COPE say about AI in peer review?

COPE says AI tools cannot be treated as authors or reviewers because they cannot take responsibility for the work. Any use in editorial or peer review must be transparent, and the human editor or reviewer remains accountable. COPE's position statement keeps responsibility with named people.

๐Ÿ“„Which journals and funders restrict generative AI in referee reports?

The U.S. National Institutes of Health, Elsevier, Springer Nature, Wiley, the JAMA Network and journals from the American Association for the Advancement of Science have all published or updated restrictions. Some funders, including the Australian Research Council, prohibit generative AI for external assessors.

โœ๏ธAre reviewers allowed to use AI for grammar and language editing?

It depends on the journal. Some policies allow language polishing if no manuscript text is exposed, while others say to disclose any AI assistance. Ask the editor before using any tool for grammar if the policy is silent.

๐ŸšจHow are journals enforcing generative AI bans in peer review?

Enforcement is limited. Editors see only the final report, and detection tools are unreliable on short, technical prose. Most policies rely on the reviewer's own disclosure and the editor's judgment.

๐ŸšซWhat happens if a reviewer violates an AI policy?

Violations can lead to removal from the journal's reviewer pool, notification to the reviewer's institution, and in grant review cases exclusion from future panels. Because confidential information may have been exposed, publishers treat it as a serious breach.

๐Ÿง Is AI-generated peer review considered more biased?

There is no broad evidence that AI improves review quality. Some researchers worry it may produce more generic, less critical reports. The core risk is not machine bias alone but reviewers deferring judgment to a model.

๐Ÿ“ฌCan grant peer reviewers use AI tools?

No. The NIH explicitly prohibits generative AI in grant peer review. Other funding bodies have similar restrictions, so reviewers should keep proposals offline unless the funder states otherwise.

๐ŸŽ“What should early-career researchers do before accepting a review?

Read the journal's policy, keep the manuscript out of public AI tools, disclose any AI use to the editor, and ask when uncertain. Document your review work so there is a clear record of human judgment.

๐ŸงชDo preprint servers have similar AI rules?

Some preprint servers discourage AI-generated comments and require disclosure. Because preprints are public, confidentiality is lower, but community norms still frown on undisclosed AI reviews.

๐Ÿค–Is AI detection reliable for referee reports?

No. AI text detection is unreliable, especially on short, technical, individual-sounding referee reports. Journals are unlikely to rely on it alone to accuse referees.