Microsoft has escalated its legal defense against publishers and authors accusing the company and OpenAI of unlawfully using copyrighted works to develop artificial intelligence systems.

In court filings submitted on September 4, 2026, Microsoft asked a federal judge to grant summary judgment in its favor, arguing that training large language models on copyrighted material is a transformative use protected by fair-use principles.
The company also pointed to an analysis of 8.2 million Copilot conversations, arguing that the chatbot rarely reproduces substantial portions of the copyrighted works at issue.
The filings could become an important moment in the growing legal battle over whether AI companies can use copyrighted books, journalism and other creative works to train commercial AI systems.
What Microsoft Filed in Court
Microsoft submitted two motions for summary judgment in the Southern District of New York.
A summary-judgment motion asks the court to resolve claims without sending the dispute through a full trial when the moving party argues that the relevant facts and law support judgment in its favor.
Microsoft’s central position is that using copyrighted works to train AI models is fundamentally different from simply republishing those works.
The company argues that model training converts the material into a system capable of generating new outputs rather than functioning as a substitute copy of the original books or articles.
Microsoft Cites 8.2 Million Copilot Conversations
One of the most notable pieces of evidence Microsoft highlighted is an analysis involving 8.2 million Copilot chat logs.
According to Microsoft’s description of the analysis, the conversations were specifically selected because they contained keywords associated with the publishers’ websites, meaning the sample was designed to be particularly likely to contain material from the plaintiffs.
Microsoft says that 59,545 conversations contained at least 16 words in common with news content used to ground Copilot.
An expert examining Center for Investigative Reporting material reportedly identified 51 instances of substantial overlap within the dataset.
The company also cited analysis in the authors’ litigation, saying only 24 Copilot responses contained at least 30 matching words from the books examined. Only 10 of 212 books evaluated reportedly showed any matches.
Why Microsoft Thinks the Numbers Matter
Microsoft’s argument is straightforward: if Copilot were functioning primarily as a mechanism for reproducing copyrighted books and journalism, substantial copying should appear much more frequently.
Instead, the company says the analysis shows that substantial reproduction is rare.
Microsoft therefore argues that occasional reproduction does not change the fundamentally transformative purpose of AI model training.
However, the figures come from Microsoft’s characterization of expert analyses submitted in the litigation. They do not mean a court has accepted Microsoft’s conclusions about copyright or fair use.
Publishers Say the Problem Is Bigger Than Output Matching
The opposing side sees the case differently.
Publishers and authors argue that the copyright problem is not limited to whether Copilot frequently reproduces long passages.
Their broader claim is that Microsoft and OpenAI used copyrighted works to create commercial AI products that can answer questions that users might otherwise research by visiting the publishers’ websites or purchasing books.
The publishers therefore argue that AI systems can compete with the original works even when they do not reproduce large portions of them verbatim.
The New York Times has rejected Microsoft’s characterization of the evidence, maintaining that the companies used its journalism to develop commercial products that can substitute for the newspaper’s reporting.
The Bigger Legal Question: Is AI Training Fair Use?
At the heart of the case is one of the most important unresolved questions in generative AI law:
Can an AI company legally use copyrighted works to train models without obtaining permission from copyright holders?
Microsoft says yes, arguing that AI training is highly transformative.
The company says the purpose of training is not to distribute copies of books or articles but to create a new technological system capable of generating responses across a wide range of tasks.
That argument is particularly important because a ruling against Microsoft could have implications far beyond Copilot.
It could affect how AI companies acquire training data, negotiate licensing agreements and defend themselves against copyright lawsuits.
The Output Question Could Be Just as Important
There are actually two connected copyright questions in these lawsuits.
The first concerns training:
Can copyrighted material be copied into datasets and used to train an AI model?
The second concerns outputs:
If an AI model produces text substantially similar to a copyrighted work, who is responsible?
Microsoft’s latest filing focuses heavily on the first question while also using Copilot output analysis to argue that the second problem occurs relatively rarely.
This distinction could become critical in future AI copyright decisions.
Microsoft Wants the Case Ended Without a Trial
By seeking summary judgment, Microsoft is asking the court to resolve the claims at the current stage rather than allowing them to proceed to a full trial.
If the court accepts Microsoft’s arguments, some or potentially substantial portions of the copyright litigation could end before a jury hears the case.
If the court rejects the motions, the legal battle can continue toward trial or further proceedings.
The outcome therefore matters not only for Microsoft and OpenAI but also for the broader AI industry.
U.S. Government Has Also Entered the Copyright Debate
The Microsoft filing arrives shortly after the U.S. government submitted a statement of interest supporting OpenAI’s position in the separate New York Times copyright dispute.
The Justice Department argued that AI training using copyrighted material can qualify as fair use and emphasized the importance of AI development to technological progress, economic competitiveness and national security.
The government’s position is not a court ruling, but it gives AI companies another significant argument in favor of permissive treatment of AI training under U.S. copyright law.
Why This Case Matters for the Future of AI
The stakes extend well beyond Copilot.
If courts broadly accept Microsoft’s interpretation, AI companies could have more freedom to train models using copyrighted material without negotiating licenses for every individual work.
That could accelerate AI development and reduce the cost of building large training datasets.
But if publishers and authors prevail, AI companies could face substantially greater licensing obligations.
That could lead to a different AI ecosystem in which model developers pay publishers, authors and other rights holders for access to high-quality copyrighted material.
The decision could therefore influence everything from AI search engines and chatbots to coding assistants, research agents and future AI-generated content platforms.
Microsoft Copilot Copyright Case: What Happens Next?
The immediate next step is for the court to consider Microsoft’s summary-judgment arguments alongside the responses from publishers and authors.
The court will ultimately have to determine how existing copyright principles apply to AI model training and whether the evidence supports Microsoft’s fair-use defense.
For now, neither side has won the broader legal question.
Microsoft’s 8.2-million-conversation analysis is evidence supporting its position, but it is not a judicial finding that Copilot does not infringe copyright.
Likewise, the publishers’ allegations remain claims that must be evaluated by the court.
The Bottom Line
Microsoft’s latest court filing marks another major escalation in the AI copyright battle.
The company is telling the court that AI training is transformative fair use and that real-world Copilot conversations show substantial reproduction of copyrighted works is relatively uncommon.
Publishers and authors counter that the legal harm goes beyond verbatim copying: AI systems can use copyrighted journalism and books to create commercial products that compete with the original works.
That makes this case about more than Microsoft Copilot.
It is ultimately a test of how copyright law will work in the age of generative AI—and whether the existing rules can accommodate technology that learns from enormous collections of human-created works.
Read More:- McKinsey State of AI 2026: AI Adoption Is Rising, But ROI Lags
FAQ
What did Microsoft file in the Copilot copyright case?
Microsoft filed motions for summary judgment in a federal copyright dispute involving publishers and authors, arguing that its use of copyrighted material to train AI models is protected by fair use.
How many Copilot conversations did Microsoft analyze?
Microsoft cited analysis involving 8.2 million Copilot conversations produced during discovery.
What did the 8.2 million conversations show?
Microsoft said the analysis found relatively few substantial overlaps with the copyrighted works examined. It cited 59,545 conversations containing at least 16 matching words with news content and 24 responses containing at least 30 matching words from books in a separate analysis.
Is Microsoft saying Copilot never reproduces copyrighted content?
No. Microsoft’s argument is that substantial reproduction happens relatively rarely and that occasional reproduction does not undermine its broader fair-use argument.
What is Microsoft’s main legal argument?
Microsoft argues that training AI models on copyrighted works is highly transformative because the purpose is to create a new AI system rather than redistribute the original books or articles.
What do publishers argue?
Publishers argue that Microsoft and OpenAI used copyrighted journalism to build commercial AI products that can compete with or substitute for the original publications.
Has the court ruled that Microsoft’s AI training is fair use?
No. Microsoft’s fair-use position is an argument in the litigation, not a final judicial ruling.
Why is this case important for AI?
The case could help establish how U.S. copyright law applies to AI training data and could influence licensing, model development and the economics of generative AI.




