A coalition of leading publishers and author Scott Turow have launched a class action lawsuit, alleging Google used millions of copyrighted works to train its Gemini AI models. The case highlights growing legal risks for generative AI platforms.
Three of the largest publishing houses in the United States-Hachette Book Group, Inc., Cengage Learning, Inc., and Elsevier Inc.-alongside best-selling author Scott Turow, have filed a class action lawsuit against Google. The suit accuses Google of willfully infringing on the copyrights of millions of books and articles to develop its Gemini large language models.
The plaintiffs, representing both themselves and a proposed class of authors and publishers, argue that Google copied a vast array of works spanning fiction, nonfiction, children's literature, memoirs, poetry, educational materials, and scholarly articles. According to the complaint, these works were originally provided to Google for limited use in services like Google Books, but were later used without authorization to train Gemini AI.
The lawsuit details how Google allegedly obtained and duplicated copyrighted content, including material scraped from the web-sometimes from pirate sources or behind paywalls-and then stripped copyright management information to obscure the origins of the data. The complaint claims Google knowingly used these works to build a generative AI system capable of producing content that could directly substitute for the originals.
Internal Google communications cited in the filing reveal concerns about the legal and business risks of using “Publisher Provided [] copyrighted books” from Google Play Books for AI training. One internal warning referenced potential fines in the range of “$10Bs-$100Bs.” Other internal analyses highlighted restrictive licenses, publisher sensitivities, and heightened risks around fair use defenses. The complaint also notes that copyright owners explicitly told Google their works could not be used for AI training outside the original, limited-purpose agreements.
Initially, the publishers planned to intervene in the ongoing In re Google Generative AI Copyright Litigation, but chose to file this separate suit to ensure all their claims-including those not covered by the existing class action-could be pursued. The plaintiffs are represented by Oppenheim + Zebrak, LLP and Keller Rohrback. The full complaint is available here.
This legal action comes amid intensifying scrutiny of how tech giants use copyrighted material to train AI. For context, regulatory pressure on Google has been mounting in other areas as well, such as when European authorities prepared to rule on the company's search competition practices, as reported in this related coverage.
Google, founded in 1998 and now a subsidiary of Alphabet Inc., reported over $300 billion in annual revenue in 2025. The company employs more than 180,000 people worldwide and operates a range of AI products, including Gemini, which was launched in 2024 as a direct competitor to other leading generative AI platforms. Gemini's rapid adoption has made it a focal point in ongoing debates over copyright, data use, and the future of AI-driven content creation.