Skip to content
AI IntelligenceAug 18, 2026AI Intelligence
Article

Amazon is digitizing and physically destroying rare, out-of-print books to extract unique text data for training large language models.

Since public web corpora are largely exhausted, proprietary physical archives have become a critical bottleneck for model differentiation. This practice intensifies debates over intellectual property rights and the long-term sustainability of AI data pipelines.

Data Cube AI EditorialSource: TechCrunch AI
01

Source Brief

Amazon is digitizing and physically destroying rare, out-of-print books to extract unique text data for training large language models. Since public web corpora are largely exhausted, proprietary physical archives have become a critical bottleneck for model differentiation. This practice intensifies debates over intellectual property rights and the long-term sustainability of AI data pipelines.