Skip to main content
AI News

AI Companies Destroy Old Books to Train Models: Legal and Ethical Concerns Rise

Reports suggest some AI firms are physically destroying old books to digitise content for training, raising copyright and cultural preservation concerns.

TLThe Lemuran Team10 August 20262 min read
Old book scanning into digital text, with subtle AI circuit patterns in a library setting

Executive Summary

AI companies are destroying old books to use their contents as training data for generative models. This practice raises legal and ethical questions around copyright, data sourcing, and the preservation of cultural materials.

What Changed

Recent reports indicate that AI firms are physically destroying old books, converting them into digital text for model training. The destruction is a method to access and digitise content that may not be available in other formats. See: Mashable: AI companies keep destroying old books. Here's why.

Why It Matters

Using old books as training data can improve model breadth but may violate copyright laws and ethical guidelines. The destruction of physical books also has implications for cultural preservation and transparency in dataset sourcing.

Business Impact

Organisations relying on AI models trained with such data may face legal risks and reputational challenges. There is increased scrutiny from regulators and the public over the provenance of training datasets.

Developer Impact

Developers may need to audit datasets more carefully and ensure compliance with copyright and data usage policies. They should be aware of the potential for tainted or unauthorised data in model training.

AI Agent Impact

AI agents trained on data from destroyed books may inherit biases or inaccuracies present in the original materials. The lack of transparency in data sourcing can complicate model evaluation and trust.

RAG Impact

Retrieval-augmented generation (RAG) systems using such data may propagate unverified or copyrighted content, raising further legal and ethical concerns.

Prompt Engineering Impact

Prompt engineers may need to consider the provenance and reliability of model outputs, especially if the underlying data includes content from destroyed or unauthorised sources.

AI builders and product leaders should review data sourcing practices, prioritise transparency, and consult legal counsel regarding copyright. Consider alternative, authorised datasets and document data provenance for compliance.

Dataset auditing tools and copyright compliance checkers are relevant. No specific new tools are mentioned in the source.

Comparison with Previous Versions

Previously, AI training often relied on web-scraped or licensed digital content. The destruction of physical books for data marks a shift in sourcing practices and raises new legal and ethical issues.

Ready to get started?

Let's build something great with AI.

Book a free 30-minute consultation. No commitment, no sales pressure.