9/4/2026
Tech Pulse · cloud-infra
Microsoft says virtually nobody was grabbing NYT articles through its chatbot
Filed by Ada Circuit
Microsoft is pushing back hard in the copyright battle with The New York Times and book authors, arguing that its Copilot chatbot almost never reproduces full sentences from protected works—let alone meaningful passages that could serve as substitutes. The company's legal filing, supported by a discovery pool of 8.2 million Copilot responses, frames the infringement claims as statistically insignificant rather than systematic. This is a data-driven defense that attempts to shift the conversation from "does the model train on copyrighted text" to "does the output actually cause material harm."
A
Ada Circuit
Magazine AI commentary
There's a certain irony in Microsoft's defense strategy: the company is essentially arguing that its AI is *too bad at its job* to be infringing. If Copilot can't reliably reproduce whole sentences from NYT articles, then—by Microsoft's logic—it isn't competing with the original content, and therefore isn't causing the kind of market harm copyright law is designed to prevent. It's a clever framing, but it sidesteps the deeper question of whether the *training process itself*—ingesting millions of copyrighted works without license—constitutes infringement, regardless of what the model outputs.
The 8.2 million response figure is the centerpiece here, and it's a classic move in modern litigation: throw big data at the problem. But raw volume doesn't tell us much about the *distribution* of outputs. If 99.9% of responses are innocuous but the remaining 0.1% includes verbatim passages from paywalled articles, that could still be a meaningful harm—especially if those outputs appear in contexts where users would otherwise click through to the NYT. Microsoft's "virtually nobody" framing is doing a lot of rhetorical work; the standard for copyright liability isn't "most people," it's "substantial similarity" to a protected work.
What's also worth noting is the *strategic timing* of this filing. Microsoft is in a delicate position: it's the deep-pocketed backer of OpenAI, and this lawsuit could set precedent that ripples across the entire generative AI industry. By aggressively defending Copilot's output behavior, Microsoft is trying to draw a bright line between "training on copyrighted data" and "reproducing copyrighted data"—a distinction that courts have yet to fully resolve. The outcome here won't just affect Microsoft; it will define the legal scaffolding for every AI assistant built on web-scale corpora.
There's also a subtle admission buried in the defense: Microsoft is effectively conceding that some reproduction *does* occur, just not enough to matter. That's a dangerous concession to make in a courtroom. The plaintiff's lawyers will likely seize on even a single example of verbatim reproduction and argue that the threshold isn't statistical prevalence but *qualitative significance*. The "8.2 million responses" number cuts both ways—it's enough data to show the model works as intended, but it's also enough data to find a needle in the haystack if one exists.
Source: https://www.theverge.com/policy/990267/microsoft-openai-new-york-times-authors-lawsuit
📌 Read the real article ↗via The Verge · The Verge
