OpenAI Inc. told a federal judge that news organizations are seeking sanctions against it for alleged discovery misconduct in a sweeping copyright lawsuit because publishers have “no evidence whatsoever” that ChatGPT systematically regurgitates their articles.
OpenAI said Friday that it has followed court orders to preserve users’ chatbot conversations “to the letter” and “gone to herculean efforts” to turn over all of its training data and tens of millions of ChatGPT outputs.
Accusations by the New York Times Co. and other publishers that the AI giant hid its ability to search training data and output logs for their articles for two years is an attempt “to rewrite history” and avoid proving their case, OpenAI told the US District Court for the Southern District of New York.
“Plaintiffs’ real gripe appears to be that these materials show their claims premised on regurgitation are baseless,” OpenAI’s opposition brief said.
The Times and other publishers including The Daily News moved for sanctions last month, arguing OpenAI’s willful deception about its technical capabilities has prevented them from establishing the foundation of their case — the extent to which OpenAI scrapes, trains, and copies its content to respond to user prompts.
OpenAI deleted user chats against a court order, and rendered sample consumer logs unusable due to redactions, according to the motion. The publishers seek to prohibit OpenAI from relying on the produced output logs “for any purpose,” a finding that the logs include or would have shown substantial regurgitation of their articles, and attorneys’ fees associated with the alleged obstruction.
OpenAI argued that the court’s preservation order didn’t require “wholesale” preservation or for OpenAI to stop all deletions, so users could still delete chats that were more than 30 days old as of the date of the court’s directive.
That resulted in about 5% of the logs the publishers requested being deleted, but OpenAI says they were replaced with other logs that didn’t bias the sample. The sample is “critically” important to examine the publishers’ claims, and precluding it would “significantly hamper OpenAI’s ability to mount a defense.”
And it’s “a technical reality” that OpenAI’s existing tools can’t execute searches across all training datasets and output logs for publishers’ articles, OpenAI said. While an internal OpenAI project searched for Times’ content across a “rough-and-ready” sample of 78 million conversations, that search “differs fundamentally” from the publishers’ request.
OpenAI urged the court against handing the publishers an “unwarranted windfall.”
Though plaintiffs didn’t seek sanctions against co-defendant Microsoft Corp., the tech company also opposed the publishers’ bid Thursday, arguing the requested restrictions against OpenAI would unduly prejudice Microsoft.
Susman Godfrey LLP, Rothwell Figg Ernst & Manbeck PC, Loevy & Loevy, Klaris Law PLLC, and Ruttenberg IP Law represent the news plaintiffs.
OpenAI is represented by Williams & Connolly LLP, Keker Van Nest & Peters LLP, Latham & Watkins LLP, and Morrison & Foerster LLP.
Microsoft is represented by Faegre Drinker Biddle & Reath LLP and Orrick, Herrington & Sutcliffe LLP.
The case is In Re: OpenAI Inc. Copyright Infringement Litig., S.D.N.Y., No. 25-md-03143, opposition filed 8/14/26.