The question of whether it is legal to train AI models on copyrighted books remains complex and unresolved, with significant implications for authors and AI developers alike. AI tools like ChatGPT, Gemini, and Claude are trained using massive datasets that include copyrighted works, often without the authors’ knowledge or consent. This situation sparks concerns among writers who feel their livelihoods may be threatened as their works help build the very technologies that could compete with or replace them. However, the legal landscape is far from straightforward, with experts noting that current copyright law dates back to 1976 and is not well-equipped to address these modern challenges.
A notable case in this arena involved Anthropic, an AI company ordered to pay $1.5 billion in a copyright settlement for using pirated versions of copyrighted books from illegal online libraries to train its models. While this seemed like a win for authors, the presiding judge ruled that the act of training on copyrighted texts itself was lawful, comparing it to a writer studying literature to create original work. Some attorneys interpret this ruling as a positive sign for AI companies, especially given Anthropic’s ambitions for $200 billion in annual revenue by 2028, indicating that courts may distinguish between using works for learning versus copying content.
Legal debates frequently revolve around "fair use," a doctrine that permits the use of copyrighted materials under certain transformative purposes such as commentary, criticism, or education. Courts often weigh factors including purpose, amount used, and market impact in determining fair use, but rulings have been inconsistent. For example, in a case where Thomson Reuters sued Ross Intelligence for using its content to build a similar AI product, the court found no fair use because the purpose was directly competitive. This suggests that future cases may hinge on whether AI applications directly compete with original copyrighted works or serve a distinct purpose.
The emerging issues extend beyond training data to AI-generated content itself. Recent rulings state that works fully created by AI lack copyright protection, raising questions about authorship, ownership, and proof of AI involvement. Attorneys emphasize that while the legal system is still catching up, early decisions will heavily influence industry norms and business strategies. For now, ongoing litigation continues to shape the boundaries of how AI and copyrighted content can legally intersect, leaving many questions open for future judicial interpretation.
Start the discussion with a take, question, or market read.