NEW YORK / RankWire.AI / – Hachette Book Group, Cengage Learning and Elsevier have filed a lawsuit against Google concerning its Gemini artificial intelligence platform. Author Scott Turow and his company, S.C.R.I.B.E., have joined the proposed class action. The complaint was submitted on July 10 to the U.S. District Court for the Southern District of New York. The plaintiffs claim Google replicated millions of copyrighted books and journal articles without authorization during the development and training of Gemini. As of July 15, the court had yet to rule on the allegations or certify the class.

According to the complaint, Google sourced material from Google Books, Google Play Books and Google Scholar. Publishers and authors provided works for specific functions such as search, sales, and research access. The plaintiffs argue these agreements did not permit extensive commercial AI training. They also contend Google downloaded large datasets scraped from the web containing copyrighted content. The filing states some of this material originated from known pirate sources and paywalled services.
The 57-page complaint outlines four claims under federal law. Three relate to alleged reproduction via Google services, web scraping, and the development or training of Gemini. The fourth cites the Digital Millennium Copyright Act. The plaintiffs allege Google removed or altered copyright management information from training datasets. The filing also references internal discussions about using publisher-supplied books. One assessment estimated potential fines between $10 billion and $100 billion. These allegations have not yet been tested in court.
Class includes owners of registered works
The proposed class encompasses holders of registered U.S. copyrights in qualifying books and journal articles. Eligible books must have an International Standard Book Number, known as an ISBN. Eligible articles should bear a Digital Object Identifier or International Standard Serial Number. The class includes works allegedly copied from Google services or obtained through web scraping. It also covers works purportedly reproduced during Gemini’s development or training phases.
Membership is also limited by copyright registration timing. One criterion requires registration within five years of publication and prior to Google’s alleged reproduction or distribution. Another mandates registration within three months of publication. The complaint excludes government entities, Google affiliates, certain court participants, and individuals who properly opt out. The court must approve the class designation before proceeding with broader group litigation.
Complaint demands damages and an accounting
The plaintiffs seek statutory damages or actual damages corresponding to proven infringements. They also request Google’s profits attributable to any confirmed copyright violations. Their remedies include an injunction, legal costs, and a jury trial. The complaint does not specify a total damages amount. It calls for Google to disclose Gemini training data, data collection practices, and known model capabilities through a court-ordered accounting.
This accounting would identify copyrighted works used in training Gemini and detail how Google collected, copied, processed, and encoded those materials. The plaintiffs also request court-supervised destruction of unauthorized copies under Google’s control. Earlier, Hachette and Cengage sought to join separate Google generative AI lawsuits in California. The New York case expands the scope to include Elsevier, Turow, and S.C.R.I.B.E., focusing on claims related to Google services, web scraping, and Gemini training.
