Rare books become AI feedstock

Amazon is reportedly buying large numbers of rare books, removing their spines, and scanning the pages so the text can be used for AI training.
The report comes from 404 Media, which said it placed a tracking device inside a rare book and later found that the book had arrived at an Amazon facility in Las Vegas. The facility is known as VGT3 and identifies itself with an image of a dinosaur holding a book in its claws. Amazon told 404 Media that it “purchases books through commercial channels to improve the products and services customers use.”
The story has symbolic force because Amazon began as an online bookseller. But the deeper issue is practical: AI companies are looking beyond the open web for high-value text, and rare books represent material that may never have been digitized or widely circulated online.
Why old paper matters to new models

Large language models, or LLMs, are AI systems trained on enormous collections of text in order to predict, generate, and organize language. For years, much of that training material came from the internet: websites, forums, public documents, digitized books, and other large-scale text collections.
That supply is no longer simple. Many models have already consumed much of what is easily available online. The article notes that companies such as Amazon need vast amounts of text, while Anthropic has faced controversy over pirated books used in training. Rare books are attractive because they can offer material that is out of print, difficult to find, or absent from the web.
Key reasons these books matter include:
- They may contain text not already absorbed by earlier models;
- they can reduce repetition in training data;
- works published before 2022 were not written by an LLM;
- older human-written text can help avoid overreliance on AI-generated material.
That last point is increasingly important. When models train too heavily on AI-generated text, they can face what researchers call “model collapse.” In plain terms, this means a model’s output may degrade if it keeps learning from synthetic text produced by systems like itself, narrowing its language patterns and amplifying errors.
Legal purchase, unresolved questions

Amazon’s statement stresses that it buys books through commercial channels. That matters: the report does not describe these books as pirated. Still, buying a physical copy of a book does not automatically settle every question about using its contents to train commercial AI systems.
There is also a preservation issue. Scanning books is not inherently bad; libraries and archives have long digitized fragile materials to protect and share them. The controversy lies in the method and purpose. Removing a book’s spine can make scanning faster, but it can also permanently damage the physical object. When rare books are treated primarily as extractable data, cultural value and commercial efficiency come into conflict.
The public record remains limited. The report identifies the Las Vegas facility, the tracking experiment, and Amazon’s response, but it does not establish how many rare books Amazon has purchased, which titles are involved, which products or models receive the resulting data, or what happens to the physical books afterward.
A new phase of the AI data race

The episode points to a broader shift in the AI industry. Early competition focused on computing power, model size, and massive web scraping. The next contest is increasingly about scarce, high-quality, and less duplicated data.
For ordinary users, this matters because AI systems reflect the data used to build them. If companies turn to rare, offline, or culturally significant materials to improve models, the debate will expand beyond copyright into preservation, transparency, and public interest.
The likely direction is more pressure for clearer rules: how companies document data sources, whether they disclose broad categories of training material, how they handle culturally valuable physical works, and whether authors, publishers, libraries, or other rights holders should share in the value created from digitized text. Amazon’s reported book-scanning operation shows that the AI race is no longer only about chips and algorithms. It is also about access to knowledge that has not yet been absorbed by the machine-readable web.
