The New York Times and The Daily News have accused OpenAI of concealing evidence in an ongoing copyright infringement lawsuit over the use of their journalism as training data for ChatGPT. The publications claim OpenAI misled the court about its ability to search both its training datasets and customer chat logs for copyrighted content. Despite OpenAI's previous assertions that searching those datasets was technically difficult and could compromise user privacy, a deposition in April revealed the company had internally conducted such searches and built a database of roughly 78 million de-identified ChatGPT conversations for this purpose, along with a “Bloom” filter tool to detect content regurgitation.
The plaintiffs argue this discovery is critical because they have long sought access to ChatGPT's training and output data to prove unauthorized copying of their work, yet OpenAI has impeded these efforts by providing a heavily redacted and unreliable sample of 20 million chat logs, after negotiating down from a requested 120 million. They also allege OpenAI violated a court preservation order by deleting billions of ChatGPT outputs after the lawsuit began and substituting millions of logs in the submitted sample. These actions have prompted the news organizations to file a motion asking the judge to sanction OpenAI, ban use of the unreliable sample as evidence, and award legal fees.
OpenAI has denied these allegations, dismissing them as attempts to intrude on private user conversations as the lawsuit weakens. The company insists it is committed to protecting user privacy and adhering to fair use principles, claiming that the plaintiffs’ efforts are driven by a failing case rather than credible legal grounds. OpenAI spokesperson Drew Pusateri stated that the accusations of hiding evidence are false and emphasized their ongoing defense grounded in privacy and fair use law.
This latest development intensifies the legal battle over AI training data transparency and copyright compliance, illustrating the growing tension between publishers seeking control over how their content is used and AI firms defending proprietary processes and user confidentiality. The outcome may have broader implications for how AI companies handle training data and transparency in the future, as courts grapple with these novel challenges.
Start the discussion with a take, question, or market read.