ChatGPT Was Built on Concealed 'Mass Piracy', Authors Tell Court

Book authors have asked a New York federal judge to rule that OpenAI built its AI models on “mass piracy”. Pointing to internal documents, a summary judgment motion alleges the AI giant downloaded books from LibGen, hid the evidence by renaming datasets, and designed its models to supplant human writers. OpenAI filed the mirror-image motion, stressing that its data harvesting qualifies as fair use.
Over the past three years, authors have filed a series of lawsuits accusing AI companies of training their models on pirated books.
Some of those cases have already produced rulings, with a bittersweet victory for Meta in California for example.
In New York, several other cases were bundled into a single proceeding where Judge Sidney Stein is overseeing claims against OpenAI and Microsoft.
This includes the Authors Guild s class action, a case filed by a group of nonfiction writers who were the first to name Microsoft as a defendant, and the Tremblay and Silverman lawsuit , which started in California in 2023 and survived a partial dismissal before moving to New York.
This week, these authors filed a motion for summary judgment. Ahead of any trial, they want Judge Stein to rule that OpenAI copied their work without permission, and that this can t qualify as fair use.
The motion covers 194 titles and asks for a finding of liability, not damages.
OpenAI s GPT models pose an existential threat to those who write and publish books, the brief states, while adding that AI-generated books of all types are already flooding the market.
The authors start by accusing OpenAI of obtaining the book copies through unauthorized sources. While the filing is heavily redacted, OpenAI stands accused of using torrented copies downloaded from LibGen,
OpenAI did not even buy the books it used. Instead, it began by torrenting [REDACTED] books from the notorious and illegal pirate library Library Genesis, also known as LibGen, the motion reads.
At the time, LibGen had already been featured in the U.S. Trade Representative s list of notorious piracy markets . According to the authors, OpenAI was well aware of the controversial nature of the site.
OpenAI took steps to conceal their piracy from the public, the motion notes, pointing to the paper that introduced GPT-3.