Academic Jobs - Home of Higher Ed Logo

Publisher AI Training on Academic Content Faces Scrutiny and Potential Litigation

Postar uma história
0Opinião
Native advertising — guest articles from $400See packages
a poster on a wall
Photo by Dedale on Unsplash

The moment usually arrives as an email. A researcher receives a form note from a journal publisher: the author agreement now covers machine learning, data mining, and large language model training. The word 'including' appears before a list of uses that did not exist when the journal first published the work. The author has one question: did I just license my scholarship to train an AI, and why didn't anyone ask first?

That question has moved from private inboxes to public disputes in the past two years. Academic publishers, facing flat subscription growth and the high cost of maintaining journal platforms, have begun selling access to scholarly content for AI training. The companies on the other side of those agreements include Microsoft and other developers building large language models, which need massive corpora of reliable text to improve their output.

The publishers describe these licenses as a new revenue stream that keeps journals financially sustainable. Authors and scholarly societies describe something closer to a surprise: content uploaded to editorial systems years ago is now being packaged as training data, often under contracts that authors never saw, with royalty terms that remain unclear and opt-out deadlines buried in portal settings.

How the deals work

Taylor & Francis, owned by Informa, became the most visible example in 2024 when The Bookseller reported that the company had signed an AI data deal with Microsoft. The publisher later said authors could opt out before their work would be used, but many researchers learned of the arrangement through the press rather than from the publisher. That sequence, more than the deal itself, is what turned a commercial agreement into a governance problem.

Wiley has discussed AI licensing revenue in public investor communications, and Cambridge University Press & Assessment has published principles saying it may license certain academic works for AI training unless authors opt out. Several smaller scholarly societies have struck their own terms through publishing partners. In most cases the publisher, not the individual author, is the contracting party with the technology company.

Payment structures are seldom public. Some publishers have said authors will receive a share of licensing revenue. Others have treated the arrangements as part of general operations without specifying author compensation. Because the money is aggregated and the agreements are confidential, authors have no easy way to know whether their particular article was included, which model used it, or how much a single paper contributed.

a stack of books sitting on top of a wooden table

Photo by Jotform on Unsplash

Why researchers are pushing back

The Authors Guild has warned publishers against treating AI clauses as routine boilerplate. The UK Society of Authors has called for authors to be asked for explicit consent before their work enters a training corpus, and a survey from the Authors' Licensing and Collecting Society found that most writers believe they should be asked and paid when their work is used to train AI systems.

That backlash has moved well beyond individual complaints. As AcademicJobs reported in its coverage of author backlash on AI licensing deals, several society journals discovered their publishers had not consulted editorial boards or had framed the deal as an operational matter outside academic governance. Editorial boards that resign over AI journal workflows have forced journals to clarify policies quickly.

Humanities scholars worry that a broad license could allow a model to paraphrase or reproduce an argument without attribution. Biomedical researchers worry about how training corpora handle tables, clinical trial reports, and annotated supplementary data files. Open-access advocates point out that many funder policies require works to be reusable under Creative Commons licenses, and a separate commercial training deal may sit uneasily with those promises to readers and taxpayers.

Learned societies have their own stake. Many own journals but license publishing operations to large houses; they are now asking whether those houses can strike AI deals unilaterally, and whether society members should have a say before the archive they built becomes a training asset.

What the law says and what it doesn't

Copyright is the first question. In the United States and the United Kingdom, copyright protects the specific expression in an article or book, not the underlying facts or ideas. A publisher that owns or controls the copyright in a work may generally license that work for a new use, provided the contract supports it. But many older publishing agreements do not mention large language models, machine learning, or text and data mining. Courts have not yet resolved whether a standard grant of 'all electronic rights' reaches AI training, and legal scholars differ on whether training is a reproduction, a transformation, or an activity outside the bundle of exclusive rights entirely.

The European Union has a narrower statutory carve-out. Article 4 of the EU Copyright Directive allows text and data mining for commercial purposes unless rightsholders have expressly reserved the right in a machine-readable way. A publisher or author who reserves rights can therefore opt out at the level of the material itself, but the practical effect depends on whether the reservation is honored downstream by data brokers and model developers.

US litigation has so far targeted technology companies rather than publishers. The Authors Guild and individual writers have sued OpenAI and others, arguing that using unauthorized collections and books amounts to copyright infringement. Those cases do not directly settle the publisher-licensing question, but their outcome will shape what a training license is worth. If courts hold that training on copyrighted works without a license is fair use, the value of these publisher agreements could weaken. If they hold the opposite, demand for clean licensed corpora will rise.

Potential litigation against publishers sits in another lane: contract law. Authors who believe a publisher exceeded the rights granted in a publishing agreement can sue for breach of contract, seek declaratory rulings, or issue pre-action letters demanding accountings of licensing revenue. Because many agreements are governed by different national laws and contain arbitration clauses, any dispute is more likely to resolve as a series of private claims than as one clean class action.

text

Photo by Brett Jordan on Unsplash

The contract problem hiding in plain sight

Look at a typical author agreement from 2012. It grants the publisher the right to reproduce, distribute, and sublicense the contribution in all forms and media now known or hereafter developed. A publisher's lawyer will read that clause as covering AI training. An author will read it as covering e-books and databases. Neither can be certain, because no court has interpreted the clause in this setting, and the contracts usually say nothing about training data, attribution, or author revocation.

Some newer contracts attempt to remove the ambiguity. Taylor & Francis and Cambridge University Press have introduced clauses that refer directly to AI training and set out opt-out choices. But the framing matters. An opt-out system places the burden on the author to monitor publisher announcements and respond by a deadline. That is a familiar pattern from open access transition deals, and it tends to amplify the voices that already understand licensing while quieter groups are enrolled by default.

Researchers at institutions with strong copyright offices have begun asking for master publishing agreements that reserve AI rights to the author unless separately negotiated. Some European universities now advise faculty to add a rider to journal contracts: the work may be published in the journal, but the author retains the right to control commercial machine learning uses. UK learned societies are testing a similar line with their publishing partners.

What to look for in an agreement:

  • Does the contract mention large language models, machine learning, or text and data mining?
  • Is the AI use treated as a separate right requiring separate payment, or folded into existing sublicensing language?
  • Does the opt-out apply to future work only, or to archives already in the publisher's system?
  • What happens to the clause if you decline, and how long do you have to respond?

What a university research office can do now

University research offices can take one concrete step this month. Add a single question to your pre-publication checklist: does this agreement mention AI training or machine learning, and what happens if you decline? For many authors the answer will be that the publisher has never raised the issue. For some it will be a checkbox they missed. Make it visible before the contract is signed, not after the training corpus has been built.

The dispute is only partly about copyright. It is also about whether scholars get to know, and have a say in, what happens to the record they spent years building. Publishers that make the arrangement legible in plain language, with real opt-outs and a fair share of revenue, will face far less of the scrutiny now heading their way.

Retrato do Prof. Sophie Martinez
Sobre o autor

Prof. Sophie MartinezVeja o autor

Academic Jobs In House Author

Discussão

De sorte em:

Seja o primeiro a comentar este artigo!

Você

Você será solicitado a entrar antes que seu comentário seja postado.

novo0 comments

Junte-se à nossa conversa!

Adicione seus comentários agora!

Tenha sua palavra

Nível de engajamento

Browse por Faculdade

Browse por assunto

Frequently Asked Questions

❓What does publisher AI training on academic content mean?

It means a publisher licenses journal articles, books, or other scholarly works to a technology company so the material can be used to train large language models such as OpenAI's GPT models. The publisher may sell or bundle access to the text, and the model learns patterns from the writing before producing new text.

📚Which academic publishers have signed AI training deals?

Taylor & Francis, part of Informa, signed an AI data deal with Microsoft in 2024. Wiley has discussed AI licensing revenue with investors, and Cambridge University Press & Assessment has published principles allowing some academic works to be licensed for AI training unless authors opt out.

💰Do academic authors get paid when their work is used for AI training?

Payment varies and is often not itemized. Some publishers have said authors will receive a share of licensing revenue, while others have not specified author compensation. Because the agreements are confidential, an author typically cannot tell whether a particular paper was used or what value a single article contributed.

🚪Can authors opt out of AI training under publisher deals?

Some publishers offer an opt-out, usually through an author portal or submission system. The opt-out may apply only to newly submitted work, may require action by a deadline, and may not cover older content already in the publisher's archive.

⚖️Is it legal for publishers to license academic content for AI training?

The answer depends on the contract and the country. In the United States and United Kingdom, a publisher can generally license works it owns or controls, but older contracts may not clearly reach AI training. In the European Union, Article 4 of the Copyright Directive allows commercial text and data mining unless rights are expressly reserved.

🗣️What are authors and scholarly societies saying about AI training deals?

The Authors Guild has warned publishers against treating AI licensing as routine boilerplate, and the UK Society of Authors has called for explicit author consent. Many learned societies want a say before journals they own are licensed to technology companies by their publishing partners.

🔓Do AI training deals conflict with open access or funder mandates?

They can sit uneasily together. Some funders require works to be distributed under Creative Commons licenses that permit reuse, while a separate commercial training agreement may impose additional constraints. Authors should check whether a publisher's AI terms conflict with the terms attached to the version of record or the accepted manuscript.

🏛️Have academic publishers been sued over AI training content?

The most prominent lawsuits have targeted technology companies such as OpenAI, not academic publishers. Publisher-facing disputes are more likely to proceed as breach of contract claims or pre-action accountings, because the central question is whether the publisher's existing author agreement covered machine learning.

🔍What should researchers look for in a publishing contract before signing?

Look for explicit mention of AI training, machine learning, text and data mining, or large language models. Check whether AI use is bundled into existing sublicensing language, whether payment is separate, whether the opt-out covers older work, and what happens if you decline. Add a rider if your institution requires your consent for commercial machine learning.

🏫How can university research offices help faculty with publisher AI licensing?

Research offices can add a single question to pre-publication checklists: does this agreement mention AI training or machine learning, and what happens if you decline? They can also maintain a short guide to publisher AI clauses and advise authors to resolve the issue before the work is accepted, not after.