In a video, a thick book is pushed into a hydraulic paper cutter, the platen descends, the blade slices, and the spine is cleanly removed. The once-connected pages become a neat stack of paper. Without context, such footage might seem satisfying and stress-relieving. But when you learn that AI companies are using this method to process physical books in bulk, including obscure and out-of-print titles, the comfort quickly fades.
In the quest for more training data, AI companies have moved from the public internet and pirated e-books to the physical world of paper books. Books published before 2022 are considered prime because they contain no AI-generated content and haven't been tampered with by modern data poisoning tools. Obscure and out-of-print books that can't be found online or in digital form offer content that web scraping can't capture, making them even more valuable.
Anthropic's exposed "Panama Project" explicitly stated that they didn't want the public to know about it. ISBNdb, a book database company, openly advertises that it can source books for AI companies, purchasing anywhere from 1,000 to 1 million physical books from used bookstores, libraries, and out-of-print catalogs, while hiding buyer identities and purchase lists through confidentiality agreements.
AI companies are aware that this looks bad. The "Panama Project" is a case in point.
AI Companies Turn to Physical Books
The first hint of AI companies targeting physical books emerged from a lawsuit against Anthropic. In August 2024, three authors sued Anthropic, alleging unauthorized use of their works to train Claude. Such disputes have been common since OpenAI's ChatGPT, but the focus was initially on whether scraping the entire internet for training constitutes infringement and whether authors' consent and royalties are required.
In this lawsuit, physical books weren't the initial focus. To make Claude more knowledgeable and better at writing, Anthropic first scanned the public internet. But web content varies in quality, so books—written by authors, edited, and proofread by publishers—caught their attention. The easiest method was to obtain e-books directly. Court documents revealed that Anthropic downloaded over 7 million books from sources like Books3, LibGen, and PiLiMi, many from pirate sites. Instead of deleting them, the company stored these files in a long-term "central library" for future model training and research.
Subsequently, an internal project codenamed "Panama Project" was launched. Simply put, it involved Anthropic purchasing physical books in bulk, scanning them, and feeding them to AI. At the time, the court was concerned about the legality of how Anthropic obtained and used these materials.
In June 2025, Judge William Alsup issued a ruling favorable to Anthropic. He reasoned that a large language model reading a book is not to recite or sell it to users, but to learn language, knowledge, and expression, then generate new content. This use is highly "transformative," akin to a human reading a book for knowledge and then creating new work, thus falling within fair use. As for the purchased physical books, Anthropic paid for each one. The company scanned a book and then destroyed the physical copy, leaving only a digital replacement, without creating an additional copy for sale. The judge found this did not infringe copyright.
However, the 7 million e-books downloaded from pirate libraries were problematic. The judge held that training models might be fair use, but it doesn't justify pirated acquisition. How you use a book is one thing; how you obtained it is another. In the U.S. legal system, precedent is significant. This ruling opened a path for AI companies: they can learn from human-written works, provided they obtain them legally.
Thus, in last year's Anthropic lawsuit, although the practice of buying and scanning physical books was publicized, the court and public focused more on the familiar debate about the legal boundaries of training AI on human knowledge. It wasn't until this year, with two reports, that the severity of the issue became apparent.
Out-of-Print Books Disappear Forever?
In January, The Washington Post unearthed more details of the "Panama Project" from over 4,000 pages of newly unsealed court documents. Anthropic spent tens of millions of dollars over about a year to purchase millions of physical books. Suppliers removed spines, fed loose pages into high-speed scanners, and sent the remaining paper to recycling companies. One project plan alone aimed to process 500,000 to 2 million books in six months.
The scale was staggering, but two lines from Anthropic's internal documents were even more damning: "The Panama Project is our destructive scanning operation of all books in the world." and "We don't want the public to know we're doing this." Together, these statements made the whole affair look terrible. Anthropic, which has always emphasized AI safety, ethics, and responsibility, was secretly running a project that destroyed millions of books. Cutting spines could be explained as an efficient scanning method, but "not wanting the public to know" suggests Anthropic itself knew this wouldn't stand up to scrutiny.
How many books did Anthropic actually destroy? The Washington Post didn't obtain the full list, and it's unknown how many were common bestsellers versus out-of-print titles. The shock at the industrial-scale book destruction made it hard to assess the specific losses.
Six months later, a 404 Media report focused concerns on out-of-print books. The report featured ISBNdb, a company claiming to have the world's largest book database. On its website, it openly sold physical book sourcing services to AI clients, claiming it could purchase 1,000 to 1 million books at a time, with sources including used bookstores, libraries, and out-of-print catalogs. The company even marketed confidentiality as a selling point: each project signs strict NDAs, and client identities, purchasing strategies, and target titles are never disclosed.
More ironically, ISBNdb itself knew this business was unsavory. In its promotional material, it directly warned clients: "AI companies destroying 2 million books" is not a headline that wins sympathy.
Another interesting point: AI companies have recently shown a particular preference for physical books published before 2022. The reason is simple: after the generative AI boom, the internet became flooded with AI-generated content, and some people deliberately used data poisoning tools to interfere with model training. In contrast, physical books published before 2022 are almost certainly written by humans, having gone through editing, proofreading, and publishing processes, and haven't been tampered with by modern poisoning tools. AI companies, having created more and more internet junk, are now seeking out old books unpolluted by AI.
This demand has reached the secondhand book market. A bookseller specializing in rare and low-circulation titles told 404 Media that since April, his weekly orders have surged from about 20 to hundreds. The books purchased are thematically diverse, in different languages, with little connection to each other, the only commonality being an ISBN number. The bookseller couldn't confirm the buyers were AI companies, but his inventory included many foreign-language, obscure, and out-of-print books. On one hand, these orders helped clear out years of unsold stock; on the other, he felt uneasy: "I don't like what these books are used for, and I don't like uncommon books being turned into pulp."
As mentioned, the process of scanning physical books is crude, which even the secondhand bookseller, whose business improved, found distressing. From the internet to pirated e-books to bulk-purchased physical books, AI companies' reach for training resources keeps extending. Now, obscure and even out-of-print books are not spared.
Out-of-print doesn't mean unique. There's no evidence that the last copy of any book has been destroyed by an AI company. The public can't see purchase lists, doesn't know which companies are buying, and doesn't know if anyone checks how many copies remain before books are fed into the cutter. But the risk is obvious.
Extremely Limited Concessions
The 2025 ruling (at least in the U.S.) has given Anthropic some legal protection. The judge held that if Anthropic legally buys a physical book, scans it into a digital file, and destroys the physical original, there is still only one copy. The digital file isn't sold or increased in circulation; it's essentially a format conversion, which can be fair use. Under this logic, cutting the book even becomes part of proving legality. If the original book remained on the market and Anthropic had an extra digital copy, it might constitute unauthorized reproduction; destroying the physical book and keeping only the digital file reduces legal risk.
James Grimmelmann, a professor of digital and information law at Cornell Tech, believes Anthropic's shift from pirated libraries to purchasing and scanning physical books is a "smart choice" and reflects a more restrained, lawful approach.
Destroying physical books is less problematic for bestsellers printed in hundreds of thousands of copies; losing one copy doesn't affect others' ability to read and purchase. But obscure, out-of-print, and historically significant editions can't be counted that way. Dawn Albinger, president of the Australian and New Zealand Association of Antiquarian Booksellers, points out that the value of old books isn't just in the printed text. Pages may contain annotations by historical figures, signatures, inscriptions, and provenance information from previous owners, and rebinding might hide earlier manuscripts. This information is attached to that specific physical book. Even if the text is scanned, once the original paper, binding, and the relationships between parts are destroyed, future generations can't re-examine them. In other words, AI companies get a digital file suitable for training, but may destroy an artifact that can't be fully replicated digitally.
After the controversy expanded, the tech industry quickly made concessions. Elon Musk said on X that he had instructed the SpaceXAI team to preserve rare books in libraries and use a more laborious non-destructive scanning method, not simply cut spines. He didn't disclose how many books SpaceXAI bought or what qualifies as "rare," making this more of a stance than a policy.
ISBNdb retreated faster. Nine days after the 404 Media report, ISBNdb deleted its physical book sourcing page for AI companies, removing promotional material about confidential sourcing and destructive scanning. The company then claimed the page was just a market demand test, the service never actually launched, and ISBNdb never purchased, scanned, or sold any books for AI training. After previously advertising the ability to purchase 1 million books at once, the entire business suddenly became "a test." Whether this explanation is convincing is one thing; it at least shows that anonymously helping AI companies bulk-purchase physical books has become a reputational risk companies are unwilling to bear.
However, Musk's promise and ISBNdb's page deletions don't replace real rules. Who is responsible for determining whether a book is rare before scanning? Is the criterion publication year, edition, and number of copies in existence, or signatures, annotations, and binding? Should AI companies disclose their bulk purchase lists? If a book is out of print, should non-destructive scanning be mandatory, with the physical copy preserved in a library? Currently, these questions have no answers.
Ironically, ISBNdb repeatedly emphasized that pre-2022 physical books are valuable precisely because they preserve human knowledge unpolluted by AI, written by authors and edited by editors. The more AI companies acknowledge this content is irreplaceable, the harder it is to explain why the physical carriers can be treated as disposable consumables.