New York Times says OpenAI hid evidence in ChatGPT copyright trial
The New York Times and The Daily News have accused OpenAI of concealing evidence in their long-running copyright lawsuit, claiming the company misrepresented its ability to search its ChatGPT chat logs and training datasets for their copyrighted journalism. The allegation matters because it strikes at the heart of a landmark case testing whether training generative AI models on published news content, and reproducing it in outputs, breaches copyright law — and because it could expose OpenAI to court sanctions for its conduct during discovery.
The claims stem from an April deposition of OpenAI privacy engineer Vinnie Monaco, who allegedly revealed the company had already searched its training corpus for copyrighted works and had amassed a database of roughly 78 million de-identified conversations to gauge its own infringement, alongside a "Bloom" filter under "Project Giraffe" that logged regurgitation in outputs. The plaintiffs, who had accepted a reduced 20-million-log sample that arrived heavily redacted and deemed "unusable" by the court, also allege OpenAI deleted billions of outputs in breach of a preservation order. They are asking the judge to bar OpenAI from relying on the sample and to make it pay their legal fees. OpenAI's spokesperson Drew Pusateri denied the allegations, calling them "blatantly false" and accusing the Times of trying to invade users' privacy as its case weakens.
- NYT accuses OpenAI of hiding evidence it could search its data.
- A deposition allegedly exposed internal infringement searches and an 78-million-log database.
- OpenAI denies it, defending user privacy and fair use.