Category: Technology | Published: 2026-07-28
There is a quiet gold rush happening in second-hand bookshops, library sales, and warehouse clearances across the world. The buyers are not collectors or publishers. They are AI companies, and what they want is not the stories inside the books. They want the text.
The reason comes down to a fundamental question about how AI learns, and a growing problem with the most obvious source of training data: the internet.
The Problem With Training AI on the Modern Web
To understand why old books have suddenly become valuable, you need to understand how AI learns in the first place.
Large language models are trained on enormous quantities of text. They read it, identify patterns in how words, sentences, and ideas relate to each other, and gradually build a model of language and knowledge that allows them to generate coherent, contextually appropriate responses. The more text they train on, and the better quality that text is, the more capable the resulting model tends to be.
For years, the internet was the obvious source. It contains an almost incomprehensible volume of text: articles, books, forums, academic papers, social media posts, comment sections, product descriptions, and everything else that human beings have put online since the early 1990s. Training on that material produced the first generation of capable AI systems.
But the internet is changing. A large and growing proportion of the text published online today was not written by a human. It was generated by AI. Marketing content, news summaries, blog posts, product listings, social media threads: all of these categories now contain significant volumes of AI-generated material, and that proportion is growing.
For AI companies trying to train the next generation of models, this creates a serious problem. Training on AI-generated text is not the same as training on human-authored text, and the consequences compound over time.
What Model Collapse Actually Means
Researchers have a name for what happens when AI systems are repeatedly trained on synthetic content: model collapse.
The concept is not complicated once you understand how AI learns from data. When a model trains on human-written text, it absorbs the full range of human knowledge, expression, and reasoning — the breadth and diversity of how people actually think and write. When it trains on AI-generated text instead, it is effectively learning from a compressed and filtered version of that original knowledge, one that has already been processed and flattened by an earlier model.
Repeat that process across successive model generations and the compounding effect becomes significant. Errors and biases get reinforced rather than corrected. Diversity of expression narrows. The range of knowledge represented in the model shrinks. Each generation learns from a slightly degraded version of what came before, and that degradation accumulates.
The practical result is models that become progressively less accurate, less nuanced, and less useful, not because the underlying technology has gotten worse, but because the training data has.
Printed books, particularly those published before large language models became widespread, offer something the modern web increasingly cannot: text that is guaranteed to be human-authored. A book printed in 2018 cannot contain AI-generated content, because the tools to generate it at scale did not yet exist. The text inside is a fixed record of human knowledge and expression from a period before generative AI changed the publishing landscape.
Why Pre-2022 Books Specifically
One company positioned to meet this demand is ISBNdb, which has traditionally helped booksellers, libraries, and distributors manage book inventories. It now offers AI companies the ability to acquire physical books in bulk specifically for AI training purposes, and it has been explicit about why books published before 2022 command particular interest.
2022 is approximately when large language models became widely accessible to the public and when AI-generated text began appearing at scale online. Books published before that point represent a corpus of text that predates the contamination problem entirely.
ISBNdb's pitch to AI companies is straightforward: books are dense, edited, and authoritative. They go through drafting, editing, fact-checking, and publishing processes that most web content never sees. They represent structured, considered human knowledge rather than the sprawling, variable, and increasingly synthetic content of the modern web.
The company's summary of its proposition is blunt: the world's best AI training data is sitting on a shelf.
The Copyright Question and the Spine Problem
The shift to physical books also has a legal dimension. Much of the AI training data debate over the past few years has centred on whether scraping copyrighted material from websites constitutes infringement. Court decisions have begun to draw distinctions between different sources and methods of obtaining text, and AI companies are looking for approaches with cleaner legal standing.
Purchasing books from the secondary market, ISBNdb argues, provides a clearer chain of custody than web scraping and does not deprive creators of income they would otherwise have received, since the sale happens between buyers and sellers in the used book market rather than involving publishers or authors directly. That argument remains contested, and the wider copyright dispute between publishers, authors, and AI developers is far from settled.
There is also a practical and reputational complication in how AI learns from physical books at scale. High-volume digitisation typically requires removing the spine of each book so that individual pages can pass through an automated scanner. The process is significantly faster and cheaper than scanning intact books, but it destroys the physical copy permanently.
ISBNdb acknowledges this openly and with some candour, noting that the optics problem is real. Headlines about AI companies destroying millions of books to extract their content are not going to generate public goodwill, and the company has reportedly suggested that clients frame the process as digitally preserving knowledge rather than destroying physical copies. That framing has not been universally accepted.
Human Knowledge as a Scarce Resource
There is a broader economic story here worth sitting with. For most of the AI boom, the conversation about valuable resources has focused on computing power, semiconductors, and energy. These are expensive and limited, and control of them confers competitive advantage.
What the book-buying trend reveals is that genuinely human-created knowledge is becoming scarce in a different but equally significant way. Not because less of it exists, but because it is becoming harder to identify and separate from the growing volume of AI-generated material that surrounds it.
The understanding of how AI learns, and specifically the recognition that it learns differently and less well from synthetic data, has given pre-AI human text a commercial value it did not previously have. Old books, journals, academic publications, and other carefully produced human-authored material are now being treated as premium training assets rather than legacy content.
There is an irony in this. The industry that produced the technology changing how information is created is now paying to recover information from the period before its own technology changed everything.
What This Means for Your Business
For businesses thinking about AI, this story has a few direct implications.
First, data quality is becoming as important as data quantity in how AI learns and performs. Organisations evaluating or deploying AI systems should expect increasing scrutiny of where training data comes from and whether it can be trusted. Provenance and authenticity are moving from philosophical concerns to practical differentiators.
Second, the same logic that makes pre-2022 books valuable applies to proprietary business knowledge. Well-written documentation, technical manuals, research reports, specialist publications, and accumulated institutional knowledge are assets that your organisation may already possess and that AI systems can be trained on. The quality of that underlying material will have a direct bearing on how useful and trustworthy any AI built on top of it will be.
Third, original human expertise is not being devalued by AI. If anything, this story suggests the opposite. As AI-generated content floods every category of online information, carefully researched, clearly reasoned, and accurately written material becomes harder to find and more valuable to have.
If you want to understand how to use AI effectively in your business using data you can actually trust, our AI Consultancy page is a practical starting point for that conversation.