The moment usually arrives as an email. A researcher receives a form note from a journal publisher: the author agreement now covers machine learning, data mining, and large language model training. The word 'including' appears before a list of uses that did not exist when the journal first published the work. The author has one question: did I just license my scholarship to train an AI, and why didn't anyone ask first?
That question has moved from private inboxes to public disputes in the past two years. Academic publishers, facing flat subscription growth and the high cost of maintaining journal platforms, have begun selling access to scholarly content for AI training. The companies on the other side of those agreements include Microsoft and other developers building large language models, which need massive corpora of reliable text to improve their output.
The publishers describe these licenses as a new revenue stream that keeps journals financially sustainable. Authors and scholarly societies describe something closer to a surprise: content uploaded to editorial systems years ago is now being packaged as training data, often under contracts that authors never saw, with royalty terms that remain unclear and opt-out deadlines buried in portal settings.
How the deals work
Taylor & Francis, owned by Informa, became the most visible example in 2024 when The Bookseller reported that the company had signed an AI data deal with Microsoft. The publisher later said authors could opt out before their work would be used, but many researchers learned of the arrangement through the press rather than from the publisher. That sequence, more than the deal itself, is what turned a commercial agreement into a governance problem.
Wiley has discussed AI licensing revenue in public investor communications, and Cambridge University Press & Assessment has published principles saying it may license certain academic works for AI training unless authors opt out. Several smaller scholarly societies have struck their own terms through publishing partners. In most cases the publisher, not the individual author, is the contracting party with the technology company.
Payment structures are seldom public. Some publishers have said authors will receive a share of licensing revenue. Others have treated the arrangements as part of general operations without specifying author compensation. Because the money is aggregated and the agreements are confidential, authors have no easy way to know whether their particular article was included, which model used it, or how much a single paper contributed.
Why researchers are pushing back
The Authors Guild has warned publishers against treating AI clauses as routine boilerplate. The UK Society of Authors has called for authors to be asked for explicit consent before their work enters a training corpus, and a survey from the Authors' Licensing and Collecting Society found that most writers believe they should be asked and paid when their work is used to train AI systems.
That backlash has moved well beyond individual complaints. As AcademicJobs reported in its coverage of author backlash on AI licensing deals, several society journals discovered their publishers had not consulted editorial boards or had framed the deal as an operational matter outside academic governance. Editorial boards that resign over AI journal workflows have forced journals to clarify policies quickly.
Humanities scholars worry that a broad license could allow a model to paraphrase or reproduce an argument without attribution. Biomedical researchers worry about how training corpora handle tables, clinical trial reports, and annotated supplementary data files. Open-access advocates point out that many funder policies require works to be reusable under Creative Commons licenses, and a separate commercial training deal may sit uneasily with those promises to readers and taxpayers.
Learned societies have their own stake. Many own journals but license publishing operations to large houses; they are now asking whether those houses can strike AI deals unilaterally, and whether society members should have a say before the archive they built becomes a training asset.
What the law says and what it doesn't
Copyright is the first question. In the United States and the United Kingdom, copyright protects the specific expression in an article or book, not the underlying facts or ideas. A publisher that owns or controls the copyright in a work may generally license that work for a new use, provided the contract supports it. But many older publishing agreements do not mention large language models, machine learning, or text and data mining. Courts have not yet resolved whether a standard grant of 'all electronic rights' reaches AI training, and legal scholars differ on whether training is a reproduction, a transformation, or an activity outside the bundle of exclusive rights entirely.
The European Union has a narrower statutory carve-out. Article 4 of the EU Copyright Directive allows text and data mining for commercial purposes unless rightsholders have expressly reserved the right in a machine-readable way. A publisher or author who reserves rights can therefore opt out at the level of the material itself, but the practical effect depends on whether the reservation is honored downstream by data brokers and model developers.
US litigation has so far targeted technology companies rather than publishers. The Authors Guild and individual writers have sued OpenAI and others, arguing that using unauthorized collections and books amounts to copyright infringement. Those cases do not directly settle the publisher-licensing question, but their outcome will shape what a training license is worth. If courts hold that training on copyrighted works without a license is fair use, the value of these publisher agreements could weaken. If they hold the opposite, demand for clean licensed corpora will rise.
Potential litigation against publishers sits in another lane: contract law. Authors who believe a publisher exceeded the rights granted in a publishing agreement can sue for breach of contract, seek declaratory rulings, or issue pre-action letters demanding accountings of licensing revenue. Because many agreements are governed by different national laws and contain arbitration clauses, any dispute is more likely to resolve as a series of private claims than as one clean class action.
Photo by Brett Jordan on Unsplash
The contract problem hiding in plain sight
Look at a typical author agreement from 2012. It grants the publisher the right to reproduce, distribute, and sublicense the contribution in all forms and media now known or hereafter developed. A publisher's lawyer will read that clause as covering AI training. An author will read it as covering e-books and databases. Neither can be certain, because no court has interpreted the clause in this setting, and the contracts usually say nothing about training data, attribution, or author revocation.
Some newer contracts attempt to remove the ambiguity. Taylor & Francis and Cambridge University Press have introduced clauses that refer directly to AI training and set out opt-out choices. But the framing matters. An opt-out system places the burden on the author to monitor publisher announcements and respond by a deadline. That is a familiar pattern from open access transition deals, and it tends to amplify the voices that already understand licensing while quieter groups are enrolled by default.
Researchers at institutions with strong copyright offices have begun asking for master publishing agreements that reserve AI rights to the author unless separately negotiated. Some European universities now advise faculty to add a rider to journal contracts: the work may be published in the journal, but the author retains the right to control commercial machine learning uses. UK learned societies are testing a similar line with their publishing partners.
What to look for in an agreement:
- Does the contract mention large language models, machine learning, or text and data mining?
- Is the AI use treated as a separate right requiring separate payment, or folded into existing sublicensing language?
- Does the opt-out apply to future work only, or to archives already in the publisher's system?
- What happens to the clause if you decline, and how long do you have to respond?
What a university research office can do now
University research offices can take one concrete step this month. Add a single question to your pre-publication checklist: does this agreement mention AI training or machine learning, and what happens if you decline? For many authors the answer will be that the publisher has never raised the issue. For some it will be a checkbox they missed. Make it visible before the contract is signed, not after the training corpus has been built.
The dispute is only partly about copyright. It is also about whether scholars get to know, and have a say in, what happens to the record they spent years building. Publishers that make the arrangement legible in plain language, with real opt-outs and a fair share of revenue, will face far less of the scrutiny now heading their way.
