
A growing controversy has emerged over reports that some of the world’s largest artificial intelligence (AI) companies are purchasing physical books, scanning them to train AI models, and then destroying the original copies—raising fresh concerns over copyright, cultural heritage, and the preservation of knowledge.
According to recent international media reports, several technology firms have been acquiring large volumes of books, particularly older and out-of-print editions, to extract high-quality text for training generative AI systems. Court documents made public in the United States indicate that AI startup Anthropic internally referred to one such initiative as “Project Panama,” which involved destructively scanning millions of books by cutting off their bindings, digitizing every page, and discarding the physical copies afterward.
The reports suggest that AI developers increasingly prefer printed books because they provide reliable, human-written content that is free from AI-generated material now flooding the internet. Some companies have reportedly used intermediaries to anonymously purchase books in bulk from second-hand booksellers and online marketplaces.
Experts have warned that while destructive scanning is a common digitization technique, the practice becomes controversial when rare, out-of-print, or historically significant books are permanently removed from circulation. They argue that if unique editions are destroyed without proper archival preservation, humanity risks losing valuable cultural and historical resources.
The issue has also intensified the global debate over copyright and AI. Several leading AI companies, including Anthropic, Meta, and Google, are already facing lawsuits from authors and publishers over the alleged use of copyrighted books to train AI models without permission. Publishers argue that AI firms should obtain licenses instead of relying on disputed legal interpretations of “fair use.”
