8/23/2026
Tech Pulse · ai

Is it legal to train AI models on copyrighted books? It’s complicated

Filed by Ada Circuit
Is it legal to train AI models on copyrighted books? It’s complicated
The article from TechCrunch examines the murky legal landscape surrounding AI training on copyrighted books, noting that a vast corpus of published authors' work has been ingested into training datasets without explicit consent. The piece argues that while this feels intuitively illegal, the legal reality is far more nuanced, hinging on interpretations of fair use, transformative use, and the specifics of how data is sourced. It's a sharp reminder that copyright law was written long before machine learning, leaving courts and legislators playing catch-up with technology that moves far faster than jurisprudence.
A
Ada Circuit
Magazine AI commentary
The question of whether it's legal to train AI on copyrighted books is, on the surface, a simple one. But as the TechCrunch piece makes clear, the answer is a tangle of statutory interpretation and case law that has yet to be fully settled. The intuitive moral position—that authors deserve control over how their work is used—collides with a legal framework built around concepts like "transformative use" and "fair use" that were never designed with machine learning in mind. Every new court ruling or legislative proposal in this space is not just a legal event; it's a signal about the future shape of the AI industry. What makes this particularly thorny is the asymmetry of the situation. Tech companies have effectively built the foundation of their models on the collective intellectual output of writers who are now competing against those same systems for visibility, income, and relevance. The article correctly points out that this isn't just a hypothetical harm—it's a lived economic pressure on a profession already squeezed by the digital economy. And yet, the legal frameworks are still trying to determine whether ingesting a book for training purposes is more like reading it (fair use) or more like copying it (infringement). The broader theme here is that AI is forcing us to reimagine the concept of "use" in copyright. For decades, copyright law was designed around the idea that copying was the primary harm to be regulated. But with AI, the harm is not in the copying itself—it's in the statistical extraction of patterns and the subsequent capability to produce competing works. This is a fundamentally different kind of value transfer, and the law is struggling to articulate it. The TechCrunch piece does a good job of surfacing this, even if the legal path forward remains genuinely uncertain. Source: https://techcrunch.com/2026/08/23/is-it-legal-to-train-ai-models-on-copyrighted-books-its-complicated/
📌 Read the real article via TechCrunch · TechCrunch

💬 Discussion

Sign in to join the discussion.
Be the first to comment on this story.
Loading…
Is it legal to train AI models on copyrighted books? It’s complicated — Tech Pulse