AI Copyright Fight Puts Training Data Under the Legal Microscope

A major copyright battle involving OpenAI, Microsoft and The New York Times is moving toward a potentially consequential ruling over whether companies can legally use copyrighted works to train artificial intelligence systems.

In filings before U.S. District Judge Sidney Stein in Manhattan, both sides have asked the court to rule in their favor on the central question of whether AI training qualifies as “fair use” under copyright law. The decision could influence a growing number of lawsuits involving publishers, authors and technology companies.

The dispute stems from a lawsuit filed by The New York Times, which alleges that OpenAI and Microsoft used millions of its articles without permission in developing AI systems. Authors, including John Grisham, Jonathan Franzen and George R.R. Martin, have brought a separate case alleging that their books were similarly used for AI training.

The cases were later consolidated in New York, alongside a wider wave of copyright disputes targeting the use of protected works in the development of generative AI.

At the heart of the fight is the concept of “transformative” use. Courts have been divided over whether converting copyrighted material into training data represents a sufficiently different purpose to qualify for protection under the fair-use doctrine.

Two earlier rulings in California offered some support for the technology companies, although neither completely settled the issue.

Judge William Alsup previously described Anthropic’s use of books to train AI as highly transformative. Judge Vince Chhabria likewise ruled in favor of Meta in a separate case involving authors, while cautioning that AI training could fall outside fair use in circumstances where generated material threatens to overwhelm markets for human-created works.

Authors challenging OpenAI have seized on that concern. They argue that generative AI is already putting pressure on the market for books and could weaken the financial incentive for writers to produce original work.

News organizations involved in the litigation have made a similar argument, saying AI-generated answers can divert audiences away from publishers’ websites while competing with the original journalism used to develop those systems.

OpenAI has taken the opposite position. The company argues that its training process does not seek to reproduce protected expression. Instead, it says, the system learns broad statistical relationships in language that allow it to generate new material.

The company has characterized AI training as an exceptionally transformative technological use and disputes claims that its systems have caused meaningful harm to writers.

Microsoft has also rejected the argument that large language models replace copyrighted books or the people who create them. The company maintains that evidence gathered during the litigation does not show that either AI training or AI-powered products serve as substitutes for the copyrighted works at issue.

The ruling in the consolidated litigation could therefore extend well beyond the parties currently before the court. With numerous copyright owners pursuing similar claims against AI developers, the court’s treatment of fair use could help shape the legal boundaries of AI development and the use of copyrighted material in future training systems.

Print Friendly, PDF & Email
Scroll to Top