News
Microsoft Asks Court to Rule AI Book Training Is Fair Use
Microsoft has filed a motion for summary judgment in a New York copyright lawsuit, arguing that training language models on books qualifies as fair use. The company cites an analysis of 8.2 million Copilot conversations in which experts found just 24 responses resembling passages from the disputed books.
Contents
Microsoft has filed a motion for summary judgment in federal court in New York in a lawsuit over the copyrighted materials used to train its artificial intelligence. The company wants the court to rule once and for all that using copyrighted books to train language models falls within the bounds of fair use.
The case is part of a consolidated multidistrict litigation combining lawsuits filed by writers and publishers against Microsoft and OpenAI. The plaintiffs allege that both companies built their language models on illegally obtained copies of copyrighted works, downloaded from sites hosting pirated book collections.
The transformative use argument
The central legal argument in Microsoft's filing is that training a large language model on copyrighted books constitutes fair use as a matter of law, not merely a factual question for a jury to decide. The company emphasizes that the training process is deeply transformative, since it serves a purpose entirely different from that of the original works.
Training a language model on copyrighted books constitutes fair use as a matter of law - from Microsoft's motion for summary judgment to the U.S. District Court for the Southern District of New York
Books were created to be read, but their use here served a technological purpose, building a model - from Microsoft's legal memorandum in the MDL case
Numbers as evidence
A key element of the motion is a forensic analysis Microsoft commissioned from an independent expert. Researchers searched 8.2 million real user conversations with Copilot for passages resembling the disputed books. The result, according to the company, is unambiguous: just 24 responses contained even 30 consecutive words matching the text of the copyrighted works.
Microsoft also states that the plaintiffs themselves made 5.3 million attempts to force Copilot to reproduce book passages using adversarial methods, deliberately crafted prompts. Fewer than one percent of these attempts produced even a 30-word match. For 202 of the 212 books named in the lawsuit, no trace of content reproduction was found in the logs.
The Meta and Anthropic precedents
Microsoft bases its argument partly on earlier rulings in similar cases against Meta and Anthropic, in which federal judges found that training AI models on legally obtained books is highly transformative. The company argues these precedents should also apply to its case, despite differences in how the training materials were acquired.
The plaintiffs, however, dispute both Microsoft's methodology and the comparison to those cases, pointing out that the key difference lies in the source of the data. They argue that Microsoft, unlike Anthropic in the case it won, relied on illegally copied materials rather than legally purchased copies of books.
Stakes for the industry
Judge Stein's ruling could set a precedent for dozens of similar cases pending in the US against other large language model developers. A summary judgment in Microsoft's favor would mean AI companies could avoid jury trials in similar cases by relying on the transformative use argument as a matter of law rather than a disputed fact.
For Polish companies and creators, the issue has only indirect relevance, since European legal frameworks differ from the American fair use doctrine. EU copyright law and text and data mining provisions include separate exceptions, and the legality of practices such as mass scanning or destroying physical books for training purposes remains far less settled in Europe than in US federal court rulings.
The next step in the case will be the plaintiffs' response to Microsoft's motion and Judge Stein's decision on whether the case can be resolved without a jury trial. OpenAI has already filed similar motions in a parallel thread of the same multidistrict case involving news articles, suggesting a coordinated strategy by both companies ahead of upcoming trial deadlines.

