A book company has removed a webpage advertising printed books as ideal material for artificial intelligence training, after the technique behind it, known as destructive scanning, provoked a public outcry once it came to light.
Destructive scanning is a process that has been used for decades. It involves cutting the spine off books and feeding the pages through a scanner at speed before the paper is discarded or recycled.
What the AI industry has brought to this long-standing practice is the industrial scale of the process, with tech companies at times ordering millions of books to train their models.
A Database Backs Down
Commercial bibliographic database ISBNdb posted an update on 28 July on its website stating that it had removed a recent landing page titled "Printed Books Sourcing for Your AI LLMs Dataset Needs".
"It was part of exploring demand, and we've chosen to pivot away from that direction", the company said, adding that its main services were unaffected and running as usual.
The backtrack came in the wake of an article by technology publication 404 Media, which picked up on the initial post by ISBNdb and its claim that the "world's best AI training data is sitting on a shelf".
The deleted landing page noted that AI models are trained on bodies of text, and that books represent the best source of quality material for this purpose.
Books printed before 2022 are particularly prized by AI labs, since AI-generated text became increasingly common in books, articles and online content after that point.
The industry regards books as the gold standard because they contain richer language and more information than text found on the internet. Such training does not typically involve the AI model memorizing or storing the text itself. Rather, the model is simply exposed to the writing, refining the statistical process that ultimately allows it to produce coherent language.
The now-deleted ISBNdb landing page, according to 404 Media and verified via the Wayback Machine internet portal, positioned the company as a "streamlined partner" for sourcing printed books "in bulk, tailored to your LLM training needs, delivered at the scale AI demands".
It argued that non-digitized books remain unavailable online, meaning physical acquisition is the only option, particularly for "older, rare, and specialist volumes (especially pre-digital)".
The vendor offered bulk sourcing of up to one million titles per order, along with help acquiring "domain-specific" books across a wide range of subjects and fields.
Destructive Scanning Takes the Spotlight
It linked to an article on its own website about "the case that validated the strategy", a reference to a 2025 US ruling in which a federal judge found that Anthropic's purchase and destruction of millions of physical books, as part of its training strategy, was legal.
Prior to this, much of the AI industry, Anthropic included, had trained its models on millions of books sourced from piracy sites, a practice the court ruled would have to be settled through damages to affected authors.
During the legal proceedings, unsealed documents revealed that Anthropic had been working on an "escape hatch" in case the litigation turned out unfavorably. Titled "Project Panama", the plan was described in an internal Anthropic document, quoted by the Washington Post, as an "effort to destructively scan all the books in the world".
The company chose to keep the scheme quiet, not out of concern over its legality, but because of how the public might perceive the mass destruction of books for AI training purposes.
In its blog post about the case, ISBNdb acknowledged that "the optics problem is real".
"'AI company destroys two million books' is not a headline that generates sympathy", it said.
Critics of destructive scanning argue that it amounts to an act of mass cultural vandalism, a charge not softened by the argument that preserving a text's information in AI form offsets the loss of the physical book.
Others point out that, given the growing scale of the practice among AI labs, out-of-print, rare and specialist books are becoming harder to obtain, particularly once the copies used are destroyed.
Aware of these criticisms and the negative public perception surrounding the issue, SpaceX CEO Elon Musk said he had asked the SpaceXAI team, responsible for generative AI chatbot Grok, to "preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning".