{"id":982489,"date":"2026-08-05T14:37:14","date_gmt":"2026-08-05T14:37:14","guid":{"rendered":"https:\/\/www.europesays.com\/us\/982489\/"},"modified":"2026-08-05T14:37:14","modified_gmt":"2026-08-05T14:37:14","slug":"why-is-anthropic-destroying-books-kathryn-james","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/us\/982489\/","title":{"rendered":"Why is Anthropic destroying books? | Kathryn James"},"content":{"rendered":"<p class=\"dcr-1s160rg\">Should we destroy all the books in the world?<\/p>\n<p class=\"dcr-1s160rg\">An answer to this question can be found in the court documents of Bartz v Anthropic PBC. The northern California district court case, decided in late July this year, highlighted the improbably named \u201cProject Panama\u201d, one of the AI company Anthropic\u2019s efforts to improve its large language model Claude. \u201cWhat is Project Panama?\u201d <a href=\"https:\/\/www.courtlistener.com\/docket\/69058235\/554\/21\/bartz-v-anthropic-pbc\/\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">court exhibit 21<\/a> asks, in an internal memo. The answer: \u201cProject Panama is our effort to destructively scan all the books in the world.\u201d The memo advises discretion: \u201cWhy use a codename? \u2026 [B]ecause we don\u2019t want it to be known that we are working on this.\u201d<\/p>\n<p class=\"dcr-1s160rg\">Destructive scanning was Anthropic\u2019s solution to a problem: in order to \u201ctrain\u201d Claude, Anthropic had to procure a large, high-quality language dataset, <a href=\"https:\/\/futurism.com\/artificial-intelligence\/ai-companies-destroying-rare-books\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">preferably one created before 2022<\/a> and the corrupting influence of generative AI on contemporary text. Claude needed as many language combinations as possible, to improve its ability to predict language outcomes. Anthropic needed data, lots of it, of very high quality. Books, as it happens, remain one of the best sources for complex, high-quality, long-form text. As the <a href=\"https:\/\/copyrightalliance.org\/wp-content\/uploads\/2025\/06\/Bartz-v.-Anthropic-Order.pdf\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">court\u2019s decision<\/a> relates, Anthropic hoped that books\u2019 \u201cwell-curated facts, well-organized analyses, and captivating fictional narratives\u201d would help \u201cClaude write as accurately and as compellingly as Authors\u201d.<\/p>\n<p class=\"dcr-1s160rg\">Anthropic had a choice: it could have secured copyright permission to use existing e-books. This would have required the \u201clegal\/practice\/business slog\u201d, as Anthropic\u2019s co-founder and CEO <a href=\"https:\/\/www.nytimes.com\/2025\/09\/05\/technology\/anthropic-settlement-copyright-ai.html\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">phrased it<\/a>, of managing copyright. Rather than engage with the texts\u2019 owners, Anthropic first chose to use pirated sources instead, a decision informing the company\u2019s <a href=\"https:\/\/www.washingtonpost.com\/technology\/2025\/09\/05\/anthropic-book-authors-copyright-settlement\/\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">$1.5bn out-of-court settlement<\/a> with authors. When that approach seemed too complicated (or, as the court decision phrases: \u201cAnthropic became \u2018not so gung ho about\u2019 training on pirated books \u2018for legal reasons\u2019\u201d), Anthropic turned to destructive scanning. As it happened, the judge ruled that using proprietary material to \u201ctrain\u201d an LLM did not, in and of itself, constitute an infringement of copyright. To the court, it would seem, \u201ctraining\u201d a corporate product is equivalent to training any human, teaching how to read in order to learn how to write.<\/p>\n<p class=\"dcr-1s160rg\">Should we be surprised that destroying printed texts seemed easier to Anthropic than working with their human authors? Either way, the decision to use destructive scanning turned the issue into one of logistics. Under US copyright law, the \u201cfair use\u201d doctrine allows you to make \u201ctransformative\u201d use of copyrighted works without the owner\u2019s permission. Anthropic took printed books and scanned them, \u201ctransforming\u201d or remediating them into a new, electronic format. They then disposed of the original printed copy: the \u201cdestructive\u201d part of destructive scanning. Along the way, Anthropic\u2019s vendors had already <a href=\"https:\/\/www.404media.co\/ai-companies-are-buying-tons-of-old-books-because-theyre-free-of-ai-slop\/\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">sliced the spines and edges of the books<\/a>, to scan them more easily before destroying them. \u201cOne replaced the other,\u201d as Judge William Alsup wrote, <a href=\"https:\/\/deadline.com\/wp-content\/uploads\/2025\/06\/anthropic.pdf\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">noting<\/a>: \u201cThere is no evidence that the new, digital copy was shown, shared, or sold outside the company.\u201d<\/p>\n<p class=\"dcr-1s160rg\">Anthropic hired an experienced logistics manager, sourced the books from vendors, hired staff and housed the books in a warehouse, and found a digitization vendor to take on the project of the destructive scanning itself. The case exhibits show warehouses of books neatly stacked and labelled on shelves, staff moving between them. As the court\u2019s decision <a href=\"https:\/\/deadline.com\/wp-content\/uploads\/2025\/06\/anthropic.pdf\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">relates<\/a>, Anthropic\u2019s vendors \u201cstripped the books from their bindings, cut their pages to size, and scanned the books into digital form \u2013 discarding the paper originals\u201d. In the court\u2019s images, stacks of books await destructive scanning. Not seen: the clean-up project of shredding and disposing of the books (or, \u201call the books in the world\u201d, reformatted as recycling or landfill).<\/p>\n<p class=\"dcr-1s160rg\"><a href=\"https:\/\/cup.columbia.edu\/book\/bookishness\/9780231195133\/\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">Bookishness<\/a>, or the range of meanings attached to <a href=\"https:\/\/www.hachettebookgroup.com\/titles\/leah-price\/what-we-talk-about-when-we-talk-about-books\/9781541673908\/?lens=basic-books\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">the book as cultural object<\/a>, has always lived alongside the book\u2019s role as textual instrument. Witness the <a href=\"https:\/\/www.nytimes.com\/2020\/09\/18\/technology\/no-trump-did-not-hold-the-bible-upside-down-at-lafayette-square.html\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">example of images<\/a> of Donald Trump, holding up a copy of the Bible at St John\u2019s during the protests of 1 June 2020. Yet unlike many other countries, the US has very few legal provisions relating to the regulation or export of American cultural heritage, and few to none governing the treatment of books.<\/p>\n<p class=\"dcr-1s160rg\">In the context of the court\u2019s decision in Bartz v <a href=\"https:\/\/www.theguardian.com\/technology\/anthropic\" data-link-name=\"in body link\" data-component=\"auto-linked-tag\" rel=\"nofollow noopener\" target=\"_blank\">Anthropic<\/a> PBC, destructive scanning is an effective mechanism to strip authorial involvement from printed texts, in order to use the content of those works to improve the functions of an LLM. It is both legal and less regulated than strip mining. What are the consequences if, as seems likely, this practice is adopted by other generative AI companies, now and in the future? How many warehouses of destructively scanned books would be too many? There is no endangered list for printed works, and little regulation of what might constitute survival of the rare or unique. Still further, there is no formal understanding of the human-generated textual object, in and of itself, as a category of cultural asset or heritage that might require protection. If we take seriously the 2022 threshold, as a moment when AI-generated text began to make significant entry into the textual record, should we start to think of the \u201cwholly human author\u201d as an emergent category of collections preservation and stewardship?<\/p>\n<p class=\"dcr-1s160rg\">There is more to say here (what to make, for instance, of the \u201cforever\u201d research library Anthropic states it intends to create), but let me close by observing that Bartz v Anthropic PBC is also a powerful statement on the importance of books or long-form text \u2013 and of readers. Judge William Alsup noted: \u201cFor centuries, we have read and re-read books. We have admired, memorized, and internalized their sweeping themes, their substantive points, and their stylistic solutions to recurring writing problems.\u201d We should worry that Anthropic decided it was easier to scan and destroy physical books than to deal with the \u201clegal\/practice\/business slog\u201d. We should worry that the current understanding of fair use allowed Anthropic to decide that it was easier to buy and destroy \u201call the books in the world\u201d than to pay the creators of those works.<\/p>\n<p class=\"dcr-1s160rg\">But perhaps the most telling aspect of this case is that it was so important to Anthropic to have access to a dataset of complex, long-form, uncorrupted text. Let\u2019s ask ourselves why that text should seem so profitable. One response to Bartz v Anthropic PBC might be to refuse to devalue our lives as readers and writers, to claim ownership of the cultural spaces in which our thought and words are created, shared and preserved. The risk with generative AI, as this single instance with Anthropic indicates, is that we cede the means of production of our large language lives: that we turn from creators to consumers, and hand the generative promise of our work to large language models and their proprietors.<\/p>\n","protected":false},"excerpt":{"rendered":"Should we destroy all the books in the world? An answer to this question can be found in&hellip;\n","protected":false},"author":3,"featured_media":982490,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_share_on_mastodon":"0"},"categories":[21],"tags":[691,738,158,67,132,68],"class_list":["post-982489","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-technology","tag-united-states","tag-unitedstates","tag-us"],"share_on_mastodon":{"url":"https:\/\/pubeurope.com\/@us\/117043412552866724","error":""},"_links":{"self":[{"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/posts\/982489","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/users\/3"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/comments?post=982489"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/posts\/982489\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/media\/982490"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/media?parent=982489"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/categories?post=982489"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/us\/wp-json\/wp\/v2\/tags?post=982489"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}