Are AI companies scanning and destroying millions of books, including rare titles? - GoGoSpoiler

Are AI companies scanning and destroying millions of books, including rare titles?


Fact Check: Are AI Companies Buying and Destroying Millions of Books?

Claim: Artificial intelligence companies are purchasing millions of physical books—including rare and obscure titles—in bulk, scanning them, and destroying them in the process.

Rating: Mostly True

What’s True

Newly unsealed internal corporate documents from Anthropic, one of the leading artificial intelligence firms, confirm that the company purchased millions of books in bulk under a confidential initiative called “Project Panama.” The documents reveal that the company prioritized high-quality and “less common” titles to “destructively scan” them for AI training data. A federal judge ordered the release of these files as part of an ongoing copyright infringement lawsuit.

What’s Undetermined

While social media rumors and initial reports suggested that multiple AI companies were engaging in similar book-destruction campaigns, it remains unconfirmed whether firms other than Anthropic have carried out comparable projects. Furthermore, while Anthropic’s internal memos referenced targeting “less common” books, the exact definition of this phrase regarding market rarity or circulation numbers requires further clarification.


Background and Investigation

In mid-2026, social media users across platforms like Instagram, Facebook, and Reddit began sharing rumors that artificial intelligence developers were buying up physical books—specifically antique, rare, and obscure editions—only to chop them apart, scan them, and reduce them to pulp to feed machine-learning models.

The rumor gained significant traction following investigative reports by digital media outlets, which cited a database company called ISBNdb. Initial blog posts from ISBNdb claimed that AI firms were sourcing millions of physical books to prevent “model collapse”—a degradation in AI output quality that happens when algorithms train excessively on synthetic data rather than human-generated text.

However, ISBNdb later retracted those blog posts and clarified that it had never purchased, scanned, or sold physical books for AI training, explaining that the page had merely been a market-interest test.

The Reality Behind Anthropic’s “Project Panama”

Despite the confusion surrounding third-party data brokers, court documents provide clear evidence regarding at least one major AI developer. In a copyright lawsuit filed against Anthropic by a group of authors in 2024, a judge ordered the unsealing of internal company files.

These documents exposed the existence of “Project Panama,” a covert operation led by Tom Turvey—a former Google executive who previously worked on partnerships for Google Books. An internal memo dated April 2024 explicitly stated the project’s objective: “Project Panama is our effort to destructively scan all the books in the world.”

Because high-speed document scanners require pages to be fed individually, the books had to undergo a destructive process where their bindings were sliced off.

Targeting “Less Common” Titles

Further examination of the unsealed records—including email correspondence between Anthropic leadership and bulk book distributors like Wonder Book and Spanish publisher Océano—revealed that the company intentionally sought out large inventories of physical literature. An email from August 2024 confirmed that “less common books are a great place to start,” aligning with claims that the company focused on obscure or specialized texts. An internal memo from October 2024 confirmed that the company had already acquired millions of volumes to fuel its data collection efforts.

While the documents definitively prove that Anthropic engaged in the bulk purchasing, scanning, and destruction of millions of books to train its Claude AI models, the extent to which other artificial intelligence corporations have adopted this specific method remains unverified.



Reference

Leave a Comment