Three of the United States’ largest publishing houses and bestselling author Scott Turow filed a class-action lawsuit against Google on July 10 in the U.S. District Court for the Southern District of New York, accusing the technology giant of committing mass copyright infringement to build its flagship artificial intelligence platform, Gemini. The plaintiffs allege that Google illegally copied millions of copyrighted books, textbooks, and academic journals without permission, stripped identifying metadata to cover its tracks, and created an AI system that directly competes with and devalues the original works.
The complaint, spearheaded by Hachette Book Group, Cengage Learning, and Elsevier, alongside Turow and his publishing entity S.C.R.I.B.E., marks a significant escalation in the publishing industry’s battle against AI developers. The publishers argue that Google weaponized legitimate business partnerships—specifically its Google Books search database and the Google Play eBook store—to secretly amass a training dataset that the company internally acknowledged could be “highly problematic.”
“Google reproduced millions of copyrighted works without permission, without providing any compensation to authors or publishers, and with full knowledge that its conduct violated copyright law,” the complaint states. The legal filing further accuses Google of deliberately removing Copyright Management Information (CMI), including titles and author names, to conceal the origin of the training data, a direct violation of the Digital Millennium Copyright Act (DMCA).
The Alleged Scheme: From Snippets to Training Data
The lawsuit details a complex relationship between publishers and Google that dates back years. Publishers voluntarily provided copyrighted texts to Google for the specific, limited purpose of making books searchable via Google Books. This service was designed to show users only short snippets and bibliographic data, not full texts. The plaintiffs assert that Google breached this agreement by feeding the complete, full-length digital copies into the maw of its Gemini AI models.
The complaint alleges that Google’s data collection was not limited to its own platforms. The company is accused of scraping content from “known pirate sources” and even bypassing paywalls on academic websites to obtain scholarly articles published by Elsevier, the parent company of prestigious journals like The Lancet and Cell. Hachette, the third-largest book publisher in the U.S., and Cengage, a major educational textbook provider, claim their entire catalogs were ingested without a licensing deal.
To underscore the scale of the alleged infringement, the suit references an internal Google document. According to the filing, the document warned that utilizing copyrighted books for AI training could be “highly problematic for Google” and potentially expose the company to “$10Bs-$100Bs in potential fines.”
Economic Harm and the “$0.39 Murder Mystery”
Beyond the act of copying, the publishers argue that Gemini’s output constitutes a direct market substitute for their products, causing severe economic damage. The lawsuit illustrates this threat with a stark example, claiming that Gemini can generate a 100-page murder mystery in twenty minutes “for a mere $0.39.”
“The scale and speed at which Gemini can create books and compete with human writers is unprecedented, and it can only do that because Google copied Plaintiffs’ and the Class’s works to train its AI,” the suit claims. The plaintiffs argue that Gemini lacks effective guardrails, allowing it to produce verbatim excerpts, replacement textbook chapters, and stylistic knockoffs that mimic specific authors’ creative choices without any credit or payment to the original creators.
The financial context for the lawsuit is immense. The filing notes that Google reported $100 billion in quarterly revenue in October 2025, a figure the plaintiffs attribute largely to the growth of its AI business, which includes Gemini’s over 650 million monthly active users. The publishers contend that Google could have easily afforded to license the content properly, as HarperCollins did in a 2024 deal with Microsoft, but chose to infringe willfully instead.
A Shifting Legal Landscape
This lawsuit is the latest volley in a multifront war between the creative industries and Silicon Valley. The same core group of plaintiffs—Elsevier, Cengage, Turow, and Hachette—filed a similar class-action suit against Meta earlier this year over its Llama AI models. Other authors have targeted OpenAI and Apple.
However, the legal precedent remains murky and geographically fractured. Two early court rulings in California recently favored AI companies, finding that training on copyrighted works can qualify as “fair use” under U.S. law. Conversely, a separate group of writers secured a historic $1.5 billion settlement against Anthropic in 2025 over pirated training data for its Claude chatbot, though a judge later rejected the deal as insufficiently comprehensive, leaving the door open for further litigation.
By filing in the Southern District of New York rather than California, the publishers appear to be seeking a more favorable judicial interpretation of copyright law, which has not been significantly updated since the pre-internet era. The outcome of this case could establish a binding precedent on whether the “fair use” doctrine extends to the industrial-scale ingestion of copyrighted material for commercial generative AI products.
The plaintiffs are asking the court to certify the case as a class action, seeking not only monetary damages but also a court-supervised process to identify all infringed works and destroy the AI models trained on illegal data. Google did not immediately respond to requests for comment on the litigation.