Source-led article

Publishers File Class Action Against Google Over Gemini AI Training Data

AI News India//3 min read
Illustration showing books and data flowing into a stylized AI brain or neural network, with the Google logo subtly incorporated.
Illustration showing books and data flowing into a stylized AI brain or neural network, with the Google logo subtly incorporated.
Featured image from the source article

Google is facing a proposed class-action lawsuit from a consortium of publishers and a prominent novelist, who claim the company unlawfully used millions of copyrighted books and journal articles to train its Gemini AI model. The lawsuit, filed on July 10 in the U.S. District Court for the Southern District of New York, names the Hachette Book Group, Cengage Learning, Elsevier, novelist Scott Turow, and his company S.C.R.I.B.E. as plaintiffs.

The Association of American Publishers publicly announced the legal action, asserting that works supplied to Google Books, Play Books, and Scholar were utilized for AI training without the necessary permissions. The plaintiffs further allege that Google incorporated copyrighted materials obtained through extensive web scraping, including content from pirate sites and paywalled digital libraries.

Core Allegations of Copyright Infringement

At the heart of the lawsuit is the question of whether Google’s use of these copyrighted materials to train its commercial AI model, Gemini, constitutes unauthorized reproduction. The complaint details four primary counts: three counts of unauthorized reproduction under the Copyright Act, specifically addressing content sourced from Google’s own services, materials acquired via web scraping, and copies made during the AI training process. The fourth count alleges that Google violated the Digital Millennium Copyright Act (DMCA) by removing copyright management information from these works.

The plaintiffs are seeking substantial monetary damages, a court injunction to prohibit further unauthorized use, a comprehensive accounting of all copyrighted works used to train Gemini, and court orders compelling Google to delete any unauthorized copies. As of the publication of this report, Google has not issued an official public statement regarding the complaint.

Internal Concerns and Google’s Public Stance

The lawsuit references what it describes as internal Google documents, suggesting an internal awareness of potential legal ramifications. One document cited reportedly characterized the use of books from Google Play Books for AI training as “highly problematic,” with potential fines estimated to be in the range of “billions of dollars.” Another quote, attributed to Gemini’s lead engineer, allegedly states, “we don’t do deals for data we already have or already possess.” It is important to note that these documents are not publicly available, and the quotes are presented within the plaintiffs’ legal filing.

This case underscores an ongoing and significant debate within the AI industry concerning intellectual property rights and the applicability of the “fair use” doctrine. Google previously published a policy paper arguing that training AI on publicly available web data qualifies as a “transformative, non-expressive use” protected under fair use. However, the current lawsuit specifically focuses on materials allegedly acquired through direct agreements with publishers or via web scraping, distinguishing it from content gathered from the public web that might be subject to robots.txt exclusions like Google-Extended.

Implications for AI Development and Content Across India

Although this lawsuit originates in the U.S., its resolution could establish crucial precedents with far-reaching implications for AI development and content industries globally, including in India. With the rapid expansion of AI adoption in India, the legal outcomes of such cases will significantly influence how AI models are trained and how intellectual property rights are upheld. Indian AI companies and content creators will closely monitor this case for clarity on data sourcing practices, licensing requirements, and the precise scope of fair use in the context of generative AI. The ruling could catalyze the development of new licensing frameworks or stricter regulations on data acquisition for AI training, impacting future collaborations between AI developers and content owners.

Case Overview

| Aspect | Detail

Datos clave

Punto Detalle
Fuente Search Engine Journal
Fecha 2026-07-17T19:51:57+00:00
Tema Google Faces Class Action Over Books Used To Train Gemini via @sejournal, @MattGSouthern