Academic Jobs - Home of Higher Ed Logo

Academic Publishers Face Author Backlash Over AI Training Data Deals

Post a Story
84views
A large pile of various self-help and business books scattered on a surface
Photo by Shiromani Kant on Unsplash

Academic authors did not learn about the Taylor & Francis AI training deal from a contract addendum. They learned about it from business news. On 19 July 2024, The Guardian reported that the publisher, part of Informa, had licensed access to a large body of its journal and book content to Microsoft to improve AI systems. The initial payment was $10 million. Authors who had submitted and revised that content, and who had refereed other people's work, found out after the fact.

The reaction was quick and sharp. The Society of Authors responded that the scale and lack of notice raised hard questions about consent. Some academics said they would decline future peer review requests for the publisher. The deal itself was not unique; it became the visible case because the money was disclosed in an investor call and the authors were not.

The $10 million disclosure that broke the story

Informa's chief executive, Stephen Carter, described the partnership as a data access arrangement rather than a sale of individual articles. Taylor & Francis said the agreement would let Microsoft access a protected content repository to improve the relevance and performance of its AI systems. The financial terms surfaced in a publishing earnings cycle, which made the announcement travel farther than a routine licensing note.

The backlash had a simple base rate. Most researchers who publish in subscription journals transfer copyright or grant broad exclusive licences to the publisher. They are not paid for the article. They are often not paid for peer review. When a publisher then licenses the resulting corpus, those authors see an arrangement that assigns all the new value to the publisher and none of the control to them. That is the structural complaint beneath the individual deal.

three white and red labeled boxes

Photo by Lala Azizli on Unsplash

What the consent gap actually looks like

In a 2023 survey of more than 1,700 published writers, the Authors Guild reported that roughly 90 percent said writers should be compensated when their work is used to train generative AI systems. The survey predates many of the highest-profile academic licensing deals, but the sentiment transferred directly. The Guardian's July 2024 report on the Taylor & Francis deal captured the same mismatch: publishers had a legal path to license the content, and authors had no comparable path to know about it.

Part of the confusion is legal. Under many subscription publishing agreements, the author transfers copyright to the publisher for the final published version. The publisher can then contract with a technology company without seeking author approval. The accepted manuscript, however, may sit in an institutional repository under a different set of rights. This split means the same article can have several lives: the publisher's final version, the author-accepted version in a repository, a funder-mandated open copy, and often a preprint posted before review. Each carries different rules.

Open access complicates the claim further. A Creative Commons Attribution licence, known as CC BY, allows anyone to reuse the work for any purpose, including commercial purposes, as long as credit is given. That reuse can include AI training. The Creative Commons FAQ treats machine training as a form of processing that a CC licence can permit. Authors who chose CC BY because their funder required it sometimes assume it stops commercial reuse; it does not. Publishers are quick to point this out, though it does not resolve the subscription journal cases where no open licence was involved.

Publishers' counterargument: this is a new licensing stream

From the publisher side, an AI training agreement is a database licence, not an author royalty event. Informa and other publishers have told investors that generative AI licensing offers recurring revenue from content they have already invested in curating, typesetting, indexing and protecting. The Bookseller reported that the Taylor & Francis deal was understood in trade terms as an access arrangement, comparable in mechanics to how libraries license full-text collections rather than how authors license individual works.

Publishers also argue that authors benefit indirectly through a more stable publishing platform and through discovery tools that depend on the same content. The argument has some weight for small societies that rely on publishing income to support journals. It has less weight for individual authors who cannot see what specifically was licensed, when, or under which terms. The gap involves money, but the larger problem is the absence of a standard notice-and-consent mechanism before a deal is signed.

pile of assorted-title books

Photo by Daria Nepriakhina 🇺🇦 on Unsplash

What this means for your lab

A collaborator I will call 'Dr. L' published three articles with the same subscription journal across a decade. When the Taylor & Francis story broke, she checked her author agreements from each year. The first two transferred copyright outright. The third, from a newer contract, granted the publisher an exclusive licence but included a machine learning clause that had not appeared in the earlier versions. She had signed all three without noticing the wording change. The publisher had every right to license those articles under the terms she accepted. That is the practical reality: most researchers do not track the licensing language in their own publication agreements.

Before the next deal lands in your field, check your own paper trail.

  • Look for 'AI training', 'machine learning' or 'text and data mining' language, plus the catch-all phrase 'all media now known or hereafter developed' in the agreement you signed.
  • If you published open access under CC BY, assume the licence permits commercial reuse, including AI training, unless the licence notice says otherwise.
  • Ask the publisher whether authors receive notice or payment when the portfolio is licensed, and request a written answer.
  • Keep a copy of your accepted manuscript in a repository where you control the version, because your funder may have given you rights the publisher did not mention.

What comes next: contracts, consortia, and a paper trail

Universities are already fighting separate battles over publisher contracts and open access terms. The dispute between UK libraries and Elsevier over subscription pricing shows how institutions can exert pressure when they negotiate as a group, though the AI clause is rarely the first item on the agenda. The same publishers under scrutiny for AI licensing are also managing paper-mill retractions at scale, which has made communication with authors more defensive across the board.

AI is also entering the submission pipeline from the author side. Editors have reported an increase in manuscripts that cite nonexistent or AI-generated sources, a problem the STM publishing trade body has begun addressing through integrity guidance. The issue is distinct from training data licensing, but the two share a root condition: machine-generated text and machine-read corpora are both expanding faster than the editorial norms around them.

For researchers, the immediate step is not to abandon journals. It is to read the agreement language before submission and to ask the question that most authors skipped: 'If you license this article to a technology company, will I be informed?' If the answer is no, that is a data point. If the answer is yes, get it in writing.

Portrait of Dr. Nathan Harlow
About the author

Dr. Nathan HarlowView author

Academic Jobs In House Author

Discussion

Sort by:

Be the first to comment on this article!

You

You’ll be asked to sign in before your comment is posted.

New0 comments

Join the conversation!

Add your comments now!

Have your say

Engagement level

Browse by Faculty

Browse by Subject

Frequently Asked Questions

❓What triggered the author backlash over academic publishers and AI training data?

In July 2024, news reports revealed that Taylor & Francis, owned by Informa, had licensed a large body of journal and book content to Microsoft for AI training. The initial payment was reported at $10 million. Academic authors said they had not been consulted, prompting criticism from groups including the Society of Authors and the Authors Guild.

❓Did Taylor & Francis tell authors about the Microsoft AI deal before signing it?

Multiple reports said authors found out from press coverage rather than from the publisher. Taylor & Francis framed the agreement as a data access arrangement, but the absence of notice became the central complaint.

❓Do academic journal authors usually own copyright in their published articles?

In many subscription journals, authors transfer copyright or grant an exclusive licence to the publisher. Once that happens, the publisher can license the final published version without a separate author sign-off.

❓Can open access articles be used for AI training without permission?

Yes, depending on the licence. A CC BY licence allows reuse for any purpose, including commercial reuse, as long as attribution is given. The Creative Commons FAQ treats machine training as a permitted form of processing under CC BY.

❓What are academic publishers saying about AI training deals?

Publishers describe these agreements as database licensing, not article-level sales. They point to investment in curation, indexing and distribution, and they argue that licensing supports the publishing platform.

❓Can authors opt out of AI training once a journal article is published?

It depends on the contract and the licence. For subscription articles where copyright was transferred, the publisher generally controls licensing. Authors who retained rights may have room to object, but should check the specific agreement.

❓Which academic publishers have faced author backlash over AI training data?

Taylor & Francis became the most visible case after the Microsoft deal. Other publishers have attracted questions from authors' organisations, but the Taylor & Francis disclosure is widely cited as a flashpoint.

❓Does the Authors Guild survey address academic author views on AI compensation?

A 2023 Authors Guild survey found roughly 90 percent of published writers said authors should be compensated when their work trains generative AI. The survey covered authors broadly, not just academics, but the sentiment carried into academic publishing debates.

❓What does a CC BY licence mean for machine learning?

CC BY permits anyone to copy and redistribute the material in any medium or format, and to adapt it for any purpose, with attribution. That includes text and data mining and training generative models, unless a separate contract says otherwise.

❓What should researchers do before submitting their next manuscript?

Read the publication agreement for AI, machine learning, or text and data mining language. Ask the publisher whether authors will be notified about portfolio-level licences. Keep the accepted manuscript in an institutional repository where possible.

❓Are there proposed legal or policy changes for AI licensing in scholarly publishing?

Authors' organisations and some university library consortia have called for consent, notice and compensation mechanisms. The policy debate is active, though standard contract language has not yet settled.