AI laboratories are purchasing large quantities of rare and out-of-print books, digitizing them, and then destroying the physical copies. The practice targets older books printed before the era of large language models—their contents carry no contamination from artificially generated text, making them valuable training material for AI systems.
Anthropic has reportedly deployed industrial-scale methods for this process, using hydraulic cutting machines to remove pages from books and scanning them with industrial-grade equipment. The approach prioritizes efficiency in extracting textual data from works that would otherwise be difficult to access digitally.
This practice raises questions about the fate of rare literary artifacts and out-of-print works. While digitization preserves content that might otherwise remain inaccessible, the destruction of original copies removes physical cultural heritage. The sourcing of training data for large language models continues to evolve in ways that reshape how publishers, authors, and archivists think about knowledge preservation and intellectual property.

