OpenAI Accused of Hiding Evidence in NYT Copyright Fight — Sanctions Sought

OpenAI is in trouble. The New York Times and other news organizations are asking a court to hit the company with sanctions, accusing it of lying for years about whether it could search ChatGPT logs for evidence of copyright infringement.

The allegations are laid out in a sanctions motion filed Thursday. At the center of it: OpenAI allegedly pretended it couldn’t search large samples of ChatGPT logs — when it had already done exactly that before the lawsuit even started.

The evidence in question is pretty critical. It could show whether ChatGPT users were able to get the chatbot to regurgitate paywalled news articles. That’s the core of the copyright case. If OpenAI was concealing that capability, it’s a big deal.

Here’s what came out. A privacy engineer named Vincent Monaco was deposed and eventually slipped that OpenAI had two large samples — 10 million and 78 million logs — that were already de-identified and could have been made available. OpenAI never disclosed them over two years of discovery.

Meanwhile, the news orgs were stuck searching a heavily redacted 20-million-log sample that OpenAI hit with 19 billion redactions. The court called it “unusable.”

OpenAI’s response? Standard stuff. Spokesperson says the Times’ case is weakening, the sanctions motion is a distraction, and they’re protecting user privacy. The Times’ lead counsel called it flat-out lying.

A lot is riding on this. If the court agrees sanctions are warranted, it could seriously damage OpenAI’s fair use defense.