Microsoft defends Copilot, claims minimal use of New York Times articles
Microsoft claims its AI assistant, Copilot, rarely reproduces text from articles, defending against copyright lawsuits from publishers like The New York Times. This assertion is critical, as a rulingโฆ
Microsoft has declared that its artificial intelligence assistant, Copilot, almost never reproduces full sentences or substantive chunks of text from news articles and books. This assertion comes in new legal filings as the tech giant defends itself against copyright infringement lawsuits brought by major publishers, including The New York Times, and groups of book authors. The company argues that its system is designed to synthesize information rather than copy it, suggesting that the risk of users accessing protected content through the chatbot is negligible. This claim stands in direct contrast to the core allegations made by the plaintiffs, who argue that the training of large language models inherently involves the unauthorized copying and distribution of their creative works.
The dispute has become one of the most significant legal battles defining the boundaries of artificial intelligence and intellectual property. Publishers are furious that AI companies scraped billions of words from the internet to train their models without permission or compensation. They view this as a massive theft of their revenue and creative labor. Microsoft and other tech firms counter that their use of copyrighted material constitutes fair use, a legal doctrine that allows limited use of copyrighted material without requiring permission from the rights holders. The New York Times, in particular, has been aggressive in its pursuit, arguing that AI summaries deprive them of subscriptions and ad revenue by providing the core value of their journalism for free. The stakes are enormous, as a ruling against Microsoft could force a complete overhaul of how AI models are built and operated globally.
To support its defense, Microsoft revealed during the discovery process that it analyzed 8.2 million interactions with Copilot. The data aims to prove that the chatbot rarely outputs verbatim text from source materials. Instead, the AI typically generates original summaries or answers based on the patterns it learned during training. Microsoft contends that this synthesis is transformative and does not serve as a market substitute for the original works. However, critics and the plaintiffs remain skeptical of these internal metrics. They argue that even partial reproductions or highly accurate summaries can degrade the value of the original content. The debate centers on whether the act of ingestion for training is the primary violation, or if the output is what matters. Legal experts suggest that the court will need to determine if the AIโs output is sufficiently different from the source material to qualify as fair use.
The outcome of these cases will likely reshape the entire digital media landscape. If Microsoft loses, it could face billions in damages and be forced to remove copyrighted material from its training data, potentially degrading the quality of its AI products. It may also require the company to pay licensing fees to publishers for future use of their content. Conversely, a victory for Microsoft would set a powerful precedent that allows tech companies to continue using existing internet content for AI development with minimal restriction. This could accelerate innovation but leave creators and journalists with little leverage to demand compensation. The legal proceedings are ongoing, with both sides gathering more evidence and preparing for trial. The resolution will determine whether the current AI boom can continue with its existing business models or must adapt to a new regulatory reality that prioritizes creator rights over technological speed. For now, the industry watches closely, knowing that the verdict will define the economic relationships between tech giants and the media that feeds them.
Read Full Story at The Verge โ


