A group of textbook authors sued OpenAI in Manhattan’s federal court Friday, alleging the tech giant has been scraping common pirating sites for academic books and using the volumes to train its generative artificial intelligence platforms like ChatGPT.
The authors, who say that textbooks make up a significant portion of the most valuable material used to train large language models, allege that OpenAI feeding their work into its platforms without permission or compensation threatens the textbook industry and the production of up-to-date academic material at large. Schools and students are always looking to cut education costs, the authors say, and if learning institutions feel they can turn to “good enough” AI-generated alternatives, no one will fund the critical work of actually producing educational materials.
“[ChatGPT] and other generative AI agents threaten to displace human-authored textbooks with cheaper or no-cost AI-generated textbooks or textbook equivalents, at the expense of losing the valuable authorial voice, research, expertise, and perspective provided by textbook authors,” says the suit, which levies four counts of copyright violation against the tech giant. “Defendants bypassed licensing [and compensation] efforts … thus threatening to displace the very sources on which they relied.”
Essentially, the suit suggests, if AI makes textbooks obsolete by violating author copyright and not compensating writers when their work is used in AI platforms, there won’t be updated educational information to be found in the future – either in print or on AI chatbots – because no one will be paid to write it.
When asked for comment on the suit, an OpenAI spokesperson said the company’s models were trained on publicly available materials and were “grounded in fair use” in a statement.
“Our models are trained on publicly available data and grounded in fair use, which helps hundreds of millions of people improve their daily lives and delivers benefits such as empowering human creativity, science, and medical research,” the spokesperson said.
The group of authors – which includes professors Michael Sullivan, Kenneth Saladin, Zvi Bodie, Alan Marcus, Kent Patton and Dana Loewy – are far from the first to levy copyright infringement allegations against AI platforms over claims of stolen work.
Merriam-Webster Dictionary and Encyclopedia Britannica sued OpenAI and a New York Times journalist filed a class action against Grammarly’s AI over similar copyright infringement claims. And, some of the world’s largest book publishers – including Hachette Book Group, Cengage Learning and Elsevier, sued Meta’s AI over copyright infringement in July, alleging it was “weakening the incentive to create” by stealing the work. Hachette also filed a joint suit with McGraw-Hill and MacMillan Publishing against Meta’s AI, also over copyright infringement claims, in May.
Those suits read similarly to this collection of authors’ suit, which seeks a court order preventing OpenAI from continuing to allegedly infringe upon the authors’ copyright and demands financial damages.
The textbook authors’ Friday suit alleges that OpenAI could have purchased their work via a legal licensing scheme – allowing the authors to remain properly compensated so they can continue their work producing academic material in fields they are specialists in – but didn’t, even though the tech giant knew its actions were illegal. It even went so far as to remove the copyright markers on the textbooks and to train ChatGPT to not tell users the text was copyrighted if they asked, the suit says.
Reproducing excerpts of their textbooks to users when asked, the authors argue, hurts the textbook industry: It makes it harder for authors to make a living and produce this work, and it also means the information people are getting via AI frequently won’t be as good as what they’d get via an actual textbook.
“The primary goal of textbooks is to teach the content of a field or subfield in a knowledgeable, coherent, and approachable way that serves the purposes of instructors and students in that field or subfield,” the suit says. “Over decades of writing these textbooks, each author develops a unique voice that becomes their distinctive style. A particular textbook is selected and adopted [by a school] not only because of the title or subject matter but because the author has a style that works for the intended learners.”
Authors write these books once they’ve developed deep expertise in the subject matter and years of experience teaching a certain subject at a certain grade level. This makes them uniquely positioned to explain a subject to students in a way that will help them learn best, the suit says.
“Most textbook authors are experienced classroom instructors, skilled writers, and subject matter experts in their fields. They each bring their unique combination of knowledge, experience, and skills to writing their textbooks,” the suit says. “Textbooks in any particular academic area cover the field in a variety of ways, reflecting the unique backgrounds, interests, and styles of their authors, as well as the needs of their intended audience of instructors and students.”
For example, the suit argues, Sullivan’s Precalculus, now in its 12th Edition, teaches the content of the undergraduate precalculus course, including graphs, functions, and analytic geometry, in a way that is accessible to students and that prepares those students to go on to study calculus.
And, sometimes, the AI tools tell users they’re providing information in the style of a specific author (ie, explaining to the user that it will explain precalculus concepts in the way Sullivan would in his textbook), making the copyright infringement claim even more blatant and threatening the livelihood of the textbook industry further.
OpenAI has gained “substantial” financial benefit from these actions, the suit says. The authors say they’re looking to claw some of that money back and preserve their industry so they’re able to continue providing sound educational materials to students and schools.