securecomm Get started

When AI Meets Antiquity: The Hidden Cost of Training Data on

July 27, 20265 min read

Key takeaways

  • Rare books provide unique linguistic and historical data that can improve AI model performance, but their physical destruction raises ethical concerns.
  • Shredding scanned books eliminates valuable material evidence such as marginalia, binding, and tactile qualities that cannot be captured digitally.
  • Legal ambiguities around copyright and moral rights can expose AI companies to litigation when they dispose of original works.
  • Collaborative digitisation, controlled licenses, synthetic data generation, and transparent auditing offer pathways to balance AI development with preservation.
  • The debate underscores a broader question about what society values: rapid digital access versus the safeguarding of cultural artifacts.

By [Your Name]July 27, 2026*

---

Artificial intelligence has become the engine of modern innovation, powering everything from chatbots to drug discovery. Yet behind the glossy demos and headline‑grabbing breakthroughs lies a quieter, more controversial process: the acquisition of massive text corpora to train large language models (LLMs). In an effort to improve accuracy, fluency, and cultural breadth, several AI firms have begun scanning, digitising, and—controversially—shredding rare and out‑of‑print books. The practice, recently highlighted in a tweet titled “AI companies are shredding rare books,” forces us to confront a clash between two seemingly disparate realms: cutting‑edge technology and centuries‑old heritage.

The Drive for Data Diversity

LLMs learn by detecting patterns across billions of words. The richer and more varied the training set, the better the model can understand nuance, dialect, and historical context. For most commercial applications, public domain texts—Wikipedia, news archives, and modern web pages—are sufficient. However, when developers aim for deep cultural competence—say, a model that can answer questions about 17th‑century poetry or medieval scientific treatises—they quickly run out of readily available material.

Rare books, first‑edition manuscripts, and limited‑run scholarly monographs become attractive targets. They contain language styles, idioms, and knowledge that simply do not exist in contemporary datasets. By feeding these texts into an LLM, a company can claim a competitive edge: a model that “understands” the subtleties of early modern English, or that can generate historically accurate prose for entertainment and education.

From Preservation to Destruction

The paradox is stark. Libraries and archives have long devoted resources to preserving fragile volumes, often storing them in climate‑controlled vaults and limiting physical access to protect them from wear. Yet some AI firms, eager to accelerate data collection, have resorted to mass digitisation followed by physical disposal—sometimes literally shredding the originals after scanning.

Why Shred? 1. **Space Constraints** – Rare books occupy valuable storage; once a high‑resolution scan is obtained, the physical object is deemed redundant. 2. **Legal Ambiguity** – In some jurisdictions, once a work is digitised, the original may be considered a “copy” that can be discarded without violating copyright, even if the work is still under protection. 3. **Speed of Acquisition** – Bulk scanning projects can be completed faster when physical handling is minimized; shredding eliminates the need for long‑term custodial care.

The Consequences - **Irreversible Loss** – Even the best scans cannot capture the tactile qualities, marginalia, binding techniques, or subtle colour variations that scholars study. - **Cultural Erasure** – Many rare books are the sole surviving witnesses to marginalized voices, indigenous knowledge, or early scientific experimentation. Destroying them removes a tangible link to those histories. - **Precedent for Future Collections** – If the industry normalises shredding after digitisation, libraries may feel pressured to surrender their holdings, accelerating a wave of loss.

Ethical and Legal Crossroads

The practice sits at the intersection of intellectual property law, archival ethics, and corporate responsibility.

- Copyright – While many rare books are in the public domain, some are still under copyright or have complex rights holders. Scanning and repurposing them without clear permission can expose companies to litigation. - Moral Rights – Authors (or their estates) retain moral rights in many jurisdictions, including the right to the integrity of the work. Destroying the original could be viewed as a violation. - Professional Standards – Organizations such as the International Council on Archives (ICA) and the American Library Association (ALA) have explicit guidelines against the destruction of unique items after digitisation.

A Path Forward: Balancing Innovation with Stewardship

The tension does not have to be binary. Several strategies can allow AI developers to benefit from rare texts while preserving them for future generations.

1. **Collaborative Digitisation Projects** Partnerships between AI firms and cultural institutions can fund high‑quality digitisation while keeping the physical items intact. In return, the institution receives a copy of the digital files and retains ownership.

2. **Controlled Access Licenses** Instead of outright ownership, AI companies can negotiate limited‑use licenses that permit training on the text but prohibit commercial redistribution of the raw data.

3. **Synthetic Data Generation** Recent advances enable the creation of *synthetic* corpora that mimic the linguistic style of rare books without needing the originals. This approach respects the source while mitigating legal risk.

4. **Transparent Auditing** Public registries that log which works have been used for training can foster accountability. Stakeholders—including scholars, authors, and the public—can verify that no irreversible destruction has occurred.

The Bigger Question: What Do We Value?

At its core, the debate forces us to ask: What is the ultimate purpose of knowledge? If the goal is to make information universally accessible, then digitisation is a powerful tool. But accessibility should not be achieved at the expense of the artifacts that embody cultural memory.

Preserving rare books is not merely a nostalgic hobby; it is a safeguard against the homogenisation of thought. When an LLM can generate Shakespeare‑like verses without ever having touched a first‑edition folio, we gain convenience but lose the material dimension of that literature—its ink, its paper, its marginal notes.

Conclusion

The rush to feed AI models with ever‑more data must be tempered by a respect for the physical heritage that underpins human knowledge. Shredding rare books may accelerate model performance, but it also erodes the very foundations of the cultures those books represent. By adopting collaborative, transparent, and ethically grounded approaches, the tech industry can continue to innovate without consigning irreplaceable artifacts to the shredder.

---

If you’re a librarian, archivist, or AI practitioner, consider how your data‑collection practices impact the long‑term health of our collective memory. The choices we make today will shape the cultural landscape for generations to come.

Sources: https://xcancel.com/HedgieMarkets/status/2081534588485296565

More field notes

Start smaller than feels respectable.