Newly revealed court documents from the ongoing copyright lawsuit filed by The New York Times against OpenAI and Microsoft expose internal acknowledgments that training AI models on scraped content constituted "the largest theft of labor in human history," according to a senior Microsoft executive. These filings, part of a three-year legal battle, detail how both companies unlawfully bypassed paywalls and amassed massive datasets from paywalled articles, deliberately removing copyright notices before using the material to train AI models like GPT-3. The lawsuit highlights concerns that these AI systems directly compete with the original news sources by substituting traditional web traffic, which poses a profound economic threat to publishers.
An internal Microsoft presentation from January 2024, cited in the filings, reveals that the AI-generated responses caused user click-through rates for The New York Times to plummet by as much as 93% compared to standard Bing searches, describing the decline as a damaging "doom loop." Microsoft’s CEO Satya Nadella testified earlier this year that paywalled content should require licensing for AI training, and he expressed that had he known OpenAI scraped behind paywalls, Microsoft would have compelled the company to retrain its models. OpenAI’s own leaders admitted the technology represents an “existential threat” to publishers, acknowledging the chatbots are “largely substitutive” and could displace the demand for original journalism.
The unsealed filings also shed light on the sheer scale of the scraped data: OpenAI’s mid-training datasets included over 91,000 copies of articles from The New York Times and other news organizations, while the Common Crawl dataset contained more than 2 million documents from nytimes.com alone. OpenAI and Microsoft collaborated by sharing training data through initiatives named Project Taxi and Project Mango, collectively incorporating over 160,000 unique works from the news publishers. Evidence shows OpenAI staff devised methods to covertly circumvent digital paywalls to expand their training data without alerting the content providers, further undermining fair-use claims.
These revelations raise critical questions about the legality and ethics of using copyrighted material for AI training, as well as the impact on the sustainable business model of news media. While courts have been generally sympathetic to AI firms invoking fair use, several admissions in the filings undercut that defense, especially because the AI output is competitive rather than transformative. The New York Daily News’s legal counsel accused OpenAI and Microsoft of knowingly engaging in wrongdoing, underscoring the escalating conflict between tech giants harnessing AI and the content creators who are being economically displaced in the process.
Start the discussion with a take, question, or market read.