AI Book Scanning Controversy: Company Removes Website Claims About Bulk Book Destruction Service

A significant controversy has erupted in the publishing and artificial intelligence industries following revelations about companies allegedly destroying millions of physical books after scanning them for AI training data. The practice, which critics have likened to modern-day book burning, has sparked intense debate about the ethics of AI development and the preservation of literary works. At the center of this storm, a third-party book database company has now quietly removed controversial content from its website after facing public backlash.

ISBNdb, a company that maintains databases of book information using International Standard Book Numbers, has deleted portions of its website that previously advertised services to help AI companies source physical books in bulk for large language model training purposes. The removed content had explicitly marketed the company as a “streamlined partner for sourcing printed books in bulk, tailored to your LLM training needs, delivered at the scale AI demands.” This revelation was first reported by 404 Media, which has been investigating the intersection of book destruction and AI training data acquisition.

The Controversial “Destruction Narrative” Defense

Perhaps most striking was ISBNdb’s attempt to reframe the destruction of physical books as something positive. In a now-deleted blog post titled “Reframing the Destruction Narrative,” the company made remarkable philosophical claims about the process. They argued that “the book is not destroyed. Its value has migrated. The paper returns to the material cycle; the knowledge enters the intellectual one.” This framing attempted to position book scanning and subsequent destruction as a form of transformation rather than loss, suggesting that converting physical texts into AI training data somehow preserved their essential value.

This argument has been met with significant skepticism from literary scholars, librarians, and book preservation advocates. The suggestion that an AI chatbot’s ability to generate summaries or information based on scanned content equals the experience of reading an actual book strikes many as fundamentally flawed. Books contain nuance, context, and cultural significance that critics argue cannot be fully captured by data extraction processes. Moreover, once physical copies are destroyed, rare or out-of-print works may become permanently inaccessible in their original form.

Company Claims Service Never Actually Existed

Following the public outcry, ISBNdb has now claimed that the controversial service was never actually operational. In a statement to 404 Media, the company asserted that “the page was a test of market interest; no such service was ever brought to life.” However, this explanation has done little to quell concerns about the broader practice of acquiring books specifically for AI training purposes. The company’s multiple dedicated web pages and blog posts about the service suggest a level of investment that goes beyond simple market research.

The deleted content remains accessible through the Internet Archive’s Wayback Machine, allowing researchers and journalists to document what was originally published. This preservation ironically demonstrates the importance of maintaining accessible archives—the very thing critics accuse AI companies of undermining through their practices of buying, scanning, and destroying physical books.

Mysterious Bulk Book Purchases Continue to Raise Concerns

Despite ISBNdb’s claims that their service never launched, evidence suggests that someone is actively purchasing large quantities of books in patterns consistent with AI training data acquisition. According to reporting by Guardian Australia, multiple secondhand booksellers have noticed unusual purchasing patterns from third-party rebuyers. These sellers describe “waves of orders that did not fit with usual customer patterns,” with particular interest in old and obscure titles that might not typically generate significant sales volume.

The practice of using intermediary companies to purchase books adds a layer of opacity to the process, making it difficult to trace exactly which AI companies are ultimately benefiting from these acquisitions. This obfuscation appears deliberate, as the optics of major technology companies directly participating in mass book destruction would be particularly damaging to their public image. The use of third-party rebuyers allows AI firms to maintain plausible deniability while still acquiring the training data they seek.

The controversy touches on broader questions about copyright, fair use, and the rights of authors whose works may be ingested into AI systems without compensation or consent. As large language models require increasingly vast datasets to improve their capabilities, the tension between technological advancement and cultural preservation continues to intensify. Publishers, authors’ groups, and cultural institutions are now grappling with how to respond to practices that may fundamentally alter the relationship between written works and their readers in the digital age.

Expert Opinion: This incident represents a critical inflection point in the AI industry’s relationship with creative content. As regulatory scrutiny intensifies and public awareness grows, we can expect AI companies to face increasing pressure to develop transparent, ethical frameworks for training data acquisition. The long-term viability of current practices appears unsustainable, and companies that fail to address these concerns may face both legal challenges and reputational damage that could significantly impact their market position.