In ongoing legal proceedings over artificial intelligence copyright infringement, Microsoft has filed new court documents asserting that its Copilot assistant rarely generates direct text from news outlets or published books. The tech giant maintains that the system almost never outputs full sentences from copyrighted sources, and virtually never yields lengthy excerpts that could serve as a functional replacement for original articles or literary works.
The filings arrive as Microsoft defends its generative technology against high-profile copyright lawsuits brought by content creators and media entities, including The New York Times alongside individual book authors. These legal actions center on claims that AI developers unauthorizedly train their models on proprietary content, enabling assistants to output passages that allegedly infringe on rights holders' intellectual property.
As part of the formal discovery process in the litigation, Microsoft provided a dataset containing 8.2 million prompt-and-response pairs from Copilot interactions. By turning over this extensive collection of user records, the company aims to demonstrate to the court that verbatim copying of copyrighted text remains an extreme rarity across real-world usage of its service.
What it means
This strategy underscores how major technology vendors are relying on empirical usage records to counter claims that generative artificial intelligence undermines traditional publishing. Rather than focusing solely on fair-use legal arguments, Microsoft is attempting to show through discovery data that actual interactions with its assistant seldom result in the reproduction of protected prose, arguing that the system does not act as a substitute for original journalism or books.



