3 min read

AI companies are buying books by the thousands

AI companies are buying books by the thousands for industrial scanning, raising fresh copyright questions and threatening rare editions.

Image: ITzine

Printed books are becoming raw material for model training. According to an investigation by 404 Media, intermediaries are buying anywhere from thousands to millions of copies, sending them through scanning operations, and often dismantling the books into individual pages. The pages are then discarded.

The arrangement keeps the AI companies that ultimately receive the text largely out of sight. It gives them access to human-written material without requiring a direct rights negotiation for every book or putting their names on each purchase. Used-book sellers are already seeing the shift: stores that once sold about 20 copies a week now report sales in the hundreds, while orders on marketplaces including Alibris and Biblio have also become noticeably more frequent.

How the book supply chain works

The buyers are reportedly purchasing books in bulk with little concern for genre, author, or edition rarity. The objective is not to participate in the book market, but to acquire a large supply of human-written text that can be scanned quickly and assembled into a training corpus.

Recommended reading

Yandex delivery robots launch in Astana

One notable participant in the chain is ISBNdb. The database was previously associated with bookstores and libraries; it now offers bulk purchasing for clients in the artificial intelligence sector, according to 404 Media.

The process is straightforward. A used book leaves a warehouse and soon reaches an industrial scanner. The book can be cut apart and its pages fed through the equipment continuously, making the process cheaper than careful, nondestructive scanning. For model training, the approach supplies text written by people rather than generated material.

The practice also raises legal questions. Documents from the case involving Anthropic described a similar process: the company bought printed books, had contractors scan them, and in some cases had the copies dismantled for faster processing. A court ruled that using lawfully acquired books for training was permissible under fair use.

That ruling did not resolve every issue. Anthropic separately faced a lawsuit over a library of 7 million pirated books that the company was alleged to have stored. A similar dispute is now unfolding around Google, with publishers accusing the company of illegally using millions of copyrighted books to train Gemini.

Against that backdrop, bulk purchases of physical books offer AI companies another way to obtain data, particularly when they want material with a clearer provenance. The source does not identify most of the end buyers because intermediaries handle the transactions.

Rare books may be hardest to replace

The supply chain could have a disproportionate effect on older books that have not been reissued and already rarely appear for sale. Once scanned, some copies do not return to circulation, while their digital versions are placed in closed databases used only for model training.

If these purchases continue to grow, the books most at risk may be precisely those with the fewest surviving copies—not the mass-market titles that are easiest to replace.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via ITzine

/ Keep reading