in Tech

AI Is Learning From Books by Destroying Them?

I use AI a lot. I find the technology fascinating and useful. But every now and then I come across a story that makes me wonder whether we have completely lost sight of what we are doing. This is one of those stories.

404 Media reports that Amazon is buying large quantities of physical books, scanning them to create AI training data and destroying the books in the process. The journalists even placed a tracker inside a shipment of rare books and followed it to an Amazon facility in Las Vegas where, according to employees, books are cut apart so they can be scanned more efficiently. Amazon itself confirmed that it purchases books through commercial channels to help develop and improve its products and services.

If this report is accurate — and there are now quite a few similar stories appearing elsewhere — I find this astonishing.

The basic process is brutally simple. Cut off the spine, separate the pages, scan them at high speed and turn the text into training data. The physical book is effectively gone afterwards. And apparently this does not only involve cheap mass-market paperbacks. 404 Media says the shipment it followed contained books that were rare in the sense that relatively few copies were in circulation.

This is not an isolated story either. Court documents previously revealed that Anthropic bought and destructively scanned millions of books for AI training. The Washington Post reported on the project, while Ars Technica described how bindings were removed, pages scanned and the original books discarded. Booksellers in Europe and Australia have meanwhile reported unusual bulk purchases of obscure and sometimes rare titles, although in many of those cases they cannot prove who the ultimate buyer is or what happens to the books afterwards.

I understand why AI companies want books. Compared with much of the web, books contain carefully edited, structured, human-written information. Older books also have another attractive property: they predate the current explosion of AI-generated text. For companies desperately looking for clean training data, that makes them extremely valuable. 404 Media previously reported that printed books are actively being marketed to AI companies for exactly this reason.

But surely there has to be a better way. A book is not simply a convenient container for a sequence of tokens. Especially with older or uncommon books, the physical copy can itself be part of our cultural and intellectual history. Once a scarce edition has been cut apart and recycled, having its text somewhere inside a gigantic training dataset is not quite the same thing.

What makes this even stranger is the contradiction. We are destroying human-made objects containing carefully collected human knowledge so machines can learn from that human knowledge.

Maybe all these reports will eventually turn out to be less dramatic than they currently appear. I hope so. But if AI really needs books this badly, I would much rather see us invest in ways of digitising and preserving them at the same time. Teaching machines should not require destroying the things from which they learn.