
Microsoft has filed legal documents arguing that its Copilot AI chatbot rarely reproduces meaningful portions of copyrighted content from news publishers and authors, as the company defends itself against copyright infringement lawsuits.
The technology company submitted 8.2 million Copilot chat logs as part of litigation discovery, selecting conversations it says were most likely to contain content from news outlets and publishers. According to Microsoft’s analysis, only 59,545 of these conversations contained at least 16 words matching news content used in the AI model’s training. An expert examining the dataset on behalf of news publishers found 51 instances of “substantial overlap” with materials from the Center for Investigative Reporting. In the separate authors’ lawsuit, experts identified only 24 Copilot responses with at least 30 matching words across 8.2 million conversations, and only 10 of 212 evaluated books contained any matches.
The New York Times, a primary plaintiff in the case, rejected Microsoft’s characterization of the findings. The newspaper’s legal team stated that discovery evidence demonstrates Microsoft and OpenAI “stole” copyrighted content to create commercial products that compete directly with journalism and threaten publishers’ business models.
Microsoft contends that using copyrighted material for AI training qualifies as fair use since the resulting products serve substantially different purposes than the original works. The company argues that occasional text reproduction does not undermine the transformative nature of large language model development. Microsoft filed this submission on Friday as part of a motion seeking summary judgment to conclude the case at an early stage.
The consolidated lawsuit involves claims from major news publishers and the Authors Guild against Microsoft and OpenAI. If the judge does not grant summary judgment, the case will proceed to full litigation.
Article Attribution | Read More at Article Source
Article summary produced by Claude AI