Loading...

Amazon is reportedly purchasing and physically destroying rare, out-of-print books in order to digitize their contents for use in AI training datasets. The company, which built its early business selling books online, is now acquiring hard-to-find texts that have never been digitized and converting them into training material for large language models.
The rationale is straightforward: LLMs have largely exhausted the supply of text available on the open internet, making rare physical books an increasingly valuable source of novel training data.
Key points from the reporting:
The destruction of irreplaceable physical texts to feed AI pipelines is drawing criticism well beyond the preservation community. There are also unresolved questions about copyright, author compensation, and whether rightsholders have consented to this use of their work.
This story is a signal, not just a headline. The AI models your clients are using, and the ones you may be reselling or building services on top of, are only as good as the data they were trained on. As major players race to acquire novel training data through increasingly aggressive means, the quality and legal standing of that data is becoming a competitive and compliance issue.
For MSPs and telecom resellers deploying AI-powered tools, understanding the upstream data provenance of the models you rely on matters. Regulatory scrutiny around AI training data is increasing, and compliance obligations are already evolving rapidly in adjacent areas of the industry.
The practical takeaway: when evaluating AI vendors, start asking harder questions about model training practices. Your enterprise clients, particularly those in regulated industries, may soon be asking those questions of you.
Watch for potential legislative or legal action targeting AI training data acquisition practices, especially in the EU where existing frameworks are more aggressive. If you are advising clients on AI adoption, this is a good time to build vendor due diligence on data sourcing into your standard evaluation process.
For the full story, read the original article on TechCrunch AI.