A report published this week by Anna's Archive — a shadow library and book preservation project — details a troubling pattern in which artificial intelligence companies have been acquiring physical books in bulk and destroying them as part of their training data pipelines. The piece, surfaced on Hacker News, describes a practice where AI firms purchase used, out-of-print, and in some cases rare books, scan or otherwise extract their content for model training, and discard the physical copies rather than donating or reshelving them. Anna's Archive does not name specific companies in every case but characterizes the behavior as systemic enough to constitute a measurable threat to physical book stock.
The scale matters. Anna's Archive has indexed roughly 30 million books as of mid-2026, but the organization estimates that tens of millions of additional titles exist only in physical form, with no surviving digital copy anywhere in its catalog or in the Internet Archive's holdings. Out-of-print titles published between approximately 1930 and 1995 represent a particular vulnerability: they fall into a copyright status gray zone that made them commercially unattractive to digitize during the Google Books era, yet they are old enough that surviving print runs are small. The report argues that commercial destruction of even modest quantities of these books — volumes that may exist in only a few dozen copies worldwide — represents an irreversible cultural loss.
The Internet Archive, which has faced its own legal pressures from publisher lawsuits over its Controlled Digital Lending program, currently hosts around 4 million fully accessible scanned books, with additional millions in restricted access. Library of Congress holdings run to approximately 17 million books, but digitization of that collection remains incomplete and unevenly funded. The Anna's Archive report frames the AI industry's appetite for text data as a new and largely unregulated pressure on a physical book supply that preservation institutions have always assumed would be stable enough to digitize gradually.
For readers who track resource scarcity and the long-term accessibility of non-digital knowledge, the preservation angle here carries practical weight that general technology coverage tends to understate. Physical books have historically functioned as a parallel information infrastructure — one that operates without electricity, internet connectivity, or platform terms of service. The destruction of low-circulation physical copies does not merely reduce cultural heritage in an abstract sense; it narrows the realistic options for knowledge access in any scenario where digital infrastructure is degraded, restricted, or selectively censored. Preppers who maintain personal reference libraries — covering topics like medicine, agriculture, engineering, or law — have implicitly understood this redundancy argument for years. The Anna's Archive report adds a concrete, documented mechanism by which that redundancy is being actively eroded, not by neglect or natural decay, but by industrial demand.
Anna's Archive is calling for accelerated community-driven scanning efforts and for institutions holding rare physical volumes to prioritize digitization before commercial acquisition channels can reach that inventory. The organization accepts book donations and has published technical guidance for individuals and small libraries who want to contribute scans to its preservation index.





