Source-led article
Major Publishers File New Lawsuit Against Google Over AI Training Data

A consortium of prominent publishers, including Hachette, Cengage, and Elsevier, along with author Scott Turow and S.C.R.I.B.E., has initiated a class-action lawsuit against Google. The suit alleges that Google utilized their copyrighted materials to train its Gemini AI platform without obtaining the necessary permissions. This action marks another significant legal challenge for AI developers regarding intellectual property rights.
The plaintiffs further claim that Google intentionally altered or removed copyright information from these works. This, they argue, was an attempt to conceal that its Gemini models were trained on “stolen materials.” This lawsuit adds to a growing number of legal battles where copyright holders are confronting AI companies like Google, Meta, OpenAI, and Anthropic over the unauthorized use of their content for AI training.
Key facts
| Aspect | Details |
|---|---|
| Plaintiffs | Hachette, Cengage, Elsevier, Scott Turow, S.C.R.I.B.E. |
| Defendant | |
| Allegation | Unauthorized use of copyrighted works for training Gemini AI models; alleged removal/alteration of copyright information. |
| Filing Court | U.S. District Court for the Southern District of New York |
Historical Relationship and Alleged Misuse
The lawsuit highlights a long-standing relationship between Google and the publishing industry. Publishers and authors have historically provided Google with copyrighted works for Google Books, primarily to make them searchable. This service typically offers short snippets of books along with bibliographic data, not full access. The plaintiffs contend that Google then used copies of these books, as well as those from the Google Play store, to train Gemini, despite never receiving authorization for such use.
The publishers assert that Google “illegally copied works from all these scope-limited programs for AI training, knowing it lacked authorization to do so.” This points to a potential breach of trust and established agreements concerning the use of their intellectual property.
Internal Concerns and Potential Fines
The lawsuit also references an internal Google document. This document allegedly indicates that using copyrighted books for AI training could be “highly problematic for Google” and might lead to “10Bs-$100Bs in potential fines.” This suggests that Google may have been aware of the legal risks associated with its AI training practices. Google has not yet issued a public response to these specific allegations.
Broader Implications for AI and Copyright
This legal action is part of a larger, ongoing debate about “fair use” in the context of AI training. While some early court decisions in California have sided with AI companies, ruling that using copyrighted works for AI training could fall under fair use, the legal landscape remains complex and unsettled. The lawsuit against Google has been filed in a different jurisdiction, the U.S. District Court for the Southern District of New York, allowing a new judge to consider the arguments. This case could establish a different precedent given the nuanced relationship between Google and the publishers involved.
For Indian AI developers and businesses, this lawsuit underscores the critical importance of intellectual property compliance when developing and deploying AI models. The outcome could influence global standards for AI training data acquisition and highlight the need for clear licensing agreements or alternative data sourcing strategies to avoid similar legal challenges. It also brings into focus the ongoing efforts to update copyright laws to address the complexities introduced by advanced AI technologies.
Source: TechCrunch AI – https://techcrunch.com/2026/07/14/google-faces-another-ai-training-lawsuit-from-major-publishers/