A sharp opinion piece published in The Guardian, written by Kathryn James, is putting fresh pressure on Anthropic over its reported practice of physically destroying books in order to digitize them for use in training Claude. The piece, which drew significant attention online, frames the issue not merely as a legal dispute but as a question of cultural stewardship and institutional accountability.
What the Allegation Says
James argues that Anthropic's approach to building its training corpus involved acquiring physical books and scanning them in ways that rendered the originals unusable. The destruction of those materials, she contends, is a separate and underappreciated harm layered on top of the already contested question of whether using copyrighted text to train AI systems is lawful. The Guardian piece connects the physical act of destruction to a broader pattern of tech companies treating cultural artifacts as raw inputs for commercial products. For a deeper look at the legal dimensions of this practice, our earlier reporting on Anthropic destroying millions of books to train Claude covers the ongoing litigation in detail.
Key Facts
- Kathryn James published the op-ed in The Guardian, questioning Anthropic's book digitization practices.
- The piece alleges physical books were destroyed as part of the scanning and digitization process.
- Anthropic faces multiple active lawsuits from authors and publishers over its training data practices.
- The op-ed adds a cultural preservation angle to what has largely been framed as a copyright argument.
- Anthropic has not issued a detailed public response specifically addressing the destruction of physical copies.
The timing of the piece matters. Anthropic is navigating a crowded legal landscape, with authors and publishers filing suits over the use of their work without compensation or consent. These cases are still working through the courts, and public opinion pieces like James's have a way of shaping the broader conversation even when they carry no legal weight. The company has positioned itself as a safety-focused lab, and scrutiny of its data sourcing sits awkwardly alongside that branding.
The destruction of books is not a side effect of progress. It is a choice, and it deserves to be named as one.Kathryn James, The Guardian
Wider Context for Anthropic's Training Practices
Anthropic is far from alone in facing questions about what went into its models. The debate over training data has become one of the defining regulatory and legal fault lines in the AI industry. Governments are watching closely, and the issue has surfaced at the highest levels of international policy discussion, including conversations at the G7 summit where AI regulation was a central topic. Whether any of this translates into binding rules remains to be seen, but the pressure on labs to document and justify their data pipelines is clearly growing.
For Anthropic specifically, the question of training data intersects with its commercial ambitions. The company has been expanding aggressively into enterprise markets, and reputational risk from ongoing litigation is a genuine concern. How the company chooses to respond to public scrutiny, including pieces like James's, will likely matter as much as the legal outcomes themselves. Readers following these developments can stay up to date through the latest Claude AI news as new details emerge.
James's op-ed does not offer new documentary evidence, and it is worth noting that the full picture of how Anthropic sourced and processed its training data remains contested. What the piece does effectively is widen the frame. Copyright is a legal category. The destruction of physical books, if confirmed at scale, touches something older and less easily litigated: the idea that cultural objects carry value beyond their informational content. That argument may not win in court, but it is resonating with readers and, increasingly, with journalists and policymakers paying attention to how AI companies have operated behind closed doors.