Source-led article
New York Times Accuses OpenAI of Evidence Concealment in Copyright Lawsuit

The New York Times, joined by The Daily News, has significantly escalated its two-year-old copyright infringement lawsuit against OpenAI. The publishers have filed a new motion for sanctions, accusing the artificial intelligence firm of deliberately concealing evidence, including internal search tools and datasets, that could identify copyrighted journalism within ChatGPT’s outputs and training data.
The core of the dispute revolves around OpenAI’s alleged use of copyrighted journalistic content to train its generative AI models and subsequently reproduce that content in user responses. Throughout the legal proceedings, OpenAI has consistently maintained that it lacked the technical capability to search its own vast training corpus for specific copyrighted works. It also argued that retrieving and processing its extensive collection of ChatGPT conversations would be unduly burdensome and raise user privacy concerns.
Allegations of Hidden Tools and Data
However, recent revelations from an April court-ordered deposition appear to contradict OpenAI’s previous claims. Vinnie Monaco, an OpenAI data privacy engineer, reportedly disclosed that the company had, in fact, conducted internal searches and evaluations of its training corpus to detect copyrighted journalism works.
Monaco’s testimony further revealed that prior to the New York Times filing its lawsuit, OpenAI had already accumulated a database of approximately 78 million de-identified ChatGPT conversations. This database was allegedly used internally to assess the extent of potential infringement on others’ works. Moreover, shortly after the lawsuit commenced, OpenAI is accused of implementing a “Bloom” filter as part of a toolset called “Project Giraffe.” This system was designed to detect and record instances of content regurgitation in ChatGPT’s outputs.
Disputed Evidence and Discovery Process
The plaintiffs initially requested a sample of 120 million chat logs from OpenAI but eventually negotiated a reduced sample of 20 million. When this sample was submitted to the courts last December, the New York Times and The Daily News claimed it was so heavily redacted as to render it “unusable,” a sentiment reportedly echoed by the court. The publishers also allege that OpenAI deleted billions of ChatGPT outputs after the lawsuit was filed, a direct violation of a court preservation order, and substituted millions of logs in the requested sample. These actions, they argue, made it unnecessarily difficult to obtain information that OpenAI had already gathered.
Ian B. Crosby, lead counsel for the plaintiffs, stated, “If OpenAI genuinely believed that copying our clients’ journalism was fair and legal, it wouldn’t have hid the truth about having done it.”
Key Developments in the Lawsuit
| Point | Detail |
|---|---|
| Accused Party | OpenAI |
| Accuser | The New York Times and The Daily News |
| Core Allegation | Copyright infringement through training AI models on journalistic content |
| New Allegation | Concealment of evidence, including internal search tools and datasets |
| Key Witness | Vinnie Monaco, OpenAI data privacy engineer |
| Disputed Evidence | 78 million de-identified ChatGPT conversations, “Bloom” filter, “Project Giraffe” |
| Legal Action | Motion for sanctions against OpenAI |
Implications for the Case
The New York Times and The Daily News are now requesting that the judge impose sanctions on OpenAI for allegedly withholding evidence and obstructing the discovery process. Their demands include preventing OpenAI from using the 20 million chat log sample as evidence due to its unreliability. They also seek for the court to accept as fact that ChatGPT logs would have demonstrated significant regurgitation and grounding of the plaintiffs’ content, and to prevent OpenAI from arguing otherwise. Furthermore, they are asking OpenAI to cover the legal fees incurred as a result of pursuing this evidence.
OpenAI’s Response
Drew Pusateri, a spokesperson for OpenAI, has denied these allegations. He accused the New York Times of attempting to access private user conversations as their case weakens. Pusateri stated, “As the Times’ case weakens and they’ve been forced to drop claims against us, they’re persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations.” He affirmed OpenAI’s commitment to defending user privacy and the principles of fair use.
The outcome of this motion for sanctions could significantly impact the trajectory of the broader copyright infringement lawsuit, potentially setting precedents for how AI companies handle copyrighted material in their training data and outputs.
Source: https://techcrunch.com/2026/07/09/new-york-times-says-openai-hid-evidence-in-chatgpt-copyright-trial/