Academic authors did not learn about the Taylor & Francis AI training deal from a contract addendum. They learned about it from business news. On 19 July 2024, The Guardian reported that the publisher, part of Informa, had licensed access to a large body of its journal and book content to Microsoft to improve AI systems. The initial payment was $10 million. Authors who had submitted and revised that content, and who had refereed other people's work, found out after the fact.
The reaction was quick and sharp. The Society of Authors responded that the scale and lack of notice raised hard questions about consent. Some academics said they would decline future peer review requests for the publisher. The deal itself was not unique; it became the visible case because the money was disclosed in an investor call and the authors were not.
The $10 million disclosure that broke the story
Informa's chief executive, Stephen Carter, described the partnership as a data access arrangement rather than a sale of individual articles. Taylor & Francis said the agreement would let Microsoft access a protected content repository to improve the relevance and performance of its AI systems. The financial terms surfaced in a publishing earnings cycle, which made the announcement travel farther than a routine licensing note.
The backlash had a simple base rate. Most researchers who publish in subscription journals transfer copyright or grant broad exclusive licences to the publisher. They are not paid for the article. They are often not paid for peer review. When a publisher then licenses the resulting corpus, those authors see an arrangement that assigns all the new value to the publisher and none of the control to them. That is the structural complaint beneath the individual deal.
Photo by Lala Azizli on Unsplash
What the consent gap actually looks like
In a 2023 survey of more than 1,700 published writers, the Authors Guild reported that roughly 90 percent said writers should be compensated when their work is used to train generative AI systems. The survey predates many of the highest-profile academic licensing deals, but the sentiment transferred directly. The Guardian's July 2024 report on the Taylor & Francis deal captured the same mismatch: publishers had a legal path to license the content, and authors had no comparable path to know about it.
Part of the confusion is legal. Under many subscription publishing agreements, the author transfers copyright to the publisher for the final published version. The publisher can then contract with a technology company without seeking author approval. The accepted manuscript, however, may sit in an institutional repository under a different set of rights. This split means the same article can have several lives: the publisher's final version, the author-accepted version in a repository, a funder-mandated open copy, and often a preprint posted before review. Each carries different rules.
Open access complicates the claim further. A Creative Commons Attribution licence, known as CC BY, allows anyone to reuse the work for any purpose, including commercial purposes, as long as credit is given. That reuse can include AI training. The Creative Commons FAQ treats machine training as a form of processing that a CC licence can permit. Authors who chose CC BY because their funder required it sometimes assume it stops commercial reuse; it does not. Publishers are quick to point this out, though it does not resolve the subscription journal cases where no open licence was involved.
Publishers' counterargument: this is a new licensing stream
From the publisher side, an AI training agreement is a database licence, not an author royalty event. Informa and other publishers have told investors that generative AI licensing offers recurring revenue from content they have already invested in curating, typesetting, indexing and protecting. The Bookseller reported that the Taylor & Francis deal was understood in trade terms as an access arrangement, comparable in mechanics to how libraries license full-text collections rather than how authors license individual works.
Publishers also argue that authors benefit indirectly through a more stable publishing platform and through discovery tools that depend on the same content. The argument has some weight for small societies that rely on publishing income to support journals. It has less weight for individual authors who cannot see what specifically was licensed, when, or under which terms. The gap involves money, but the larger problem is the absence of a standard notice-and-consent mechanism before a deal is signed.
Photo by Daria Nepriakhina 🇺🇦 on Unsplash
What this means for your lab
A collaborator I will call 'Dr. L' published three articles with the same subscription journal across a decade. When the Taylor & Francis story broke, she checked her author agreements from each year. The first two transferred copyright outright. The third, from a newer contract, granted the publisher an exclusive licence but included a machine learning clause that had not appeared in the earlier versions. She had signed all three without noticing the wording change. The publisher had every right to license those articles under the terms she accepted. That is the practical reality: most researchers do not track the licensing language in their own publication agreements.
Before the next deal lands in your field, check your own paper trail.
- Look for 'AI training', 'machine learning' or 'text and data mining' language, plus the catch-all phrase 'all media now known or hereafter developed' in the agreement you signed.
- If you published open access under CC BY, assume the licence permits commercial reuse, including AI training, unless the licence notice says otherwise.
- Ask the publisher whether authors receive notice or payment when the portfolio is licensed, and request a written answer.
- Keep a copy of your accepted manuscript in a repository where you control the version, because your funder may have given you rights the publisher did not mention.
What comes next: contracts, consortia, and a paper trail
Universities are already fighting separate battles over publisher contracts and open access terms. The dispute between UK libraries and Elsevier over subscription pricing shows how institutions can exert pressure when they negotiate as a group, though the AI clause is rarely the first item on the agenda. The same publishers under scrutiny for AI licensing are also managing paper-mill retractions at scale, which has made communication with authors more defensive across the board.
AI is also entering the submission pipeline from the author side. Editors have reported an increase in manuscripts that cite nonexistent or AI-generated sources, a problem the STM publishing trade body has begun addressing through integrity guidance. The issue is distinct from training data licensing, but the two share a root condition: machine-generated text and machine-read corpora are both expanding faster than the editorial norms around them.
For researchers, the immediate step is not to abandon journals. It is to read the agreement language before submission and to ask the question that most authors skipped: 'If you license this article to a technology company, will I be informed?' If the answer is no, that is a data point. If the answer is yes, get it in writing.
