On a bright summer day in Haarlem, Netherlands, Pieter de Vries, an antiquarian bookseller, received an unexpected email that piqued his curiosity. The message came from a woman named Natalia, who claimed to represent a company called “2077AI.” She was seeking to place a significant order of books, accompanied by a spreadsheet detailing over 3,000 ISBNs, and included instructions for matching titles, preparing quotes, and estimating shipping costs to China.
De Vries skimmed through the email, quickly dismissing it as spam or a phishing attempt. It wasn’t until weeks later, when a Dutch journalist reached out to him regarding the request, that he discovered its ties to the world of artificial intelligence. “I was stunned!” de Vries told Fortune, sharing both the email and the spreadsheet of 3,001 titles primarily published between 2020 and 2021 by notable academic houses like Emerald Publishing, Elsevier, Wiley, Routledge, and Oxford University Press. The subjects varied, covering areas from business and education to engineering, public policy, and medicine. “It was very unusual,” he remarked, explaining why he initially thought it was a scam.
De Vries found that he wasn't alone in his reaction; other antiquarian booksellers in the Netherlands received similar requests and also disregarded them, viewing them as dubious. While this inquiry had minimal impact on de Vries' business, he noted that it could pose a challenge for booksellers who focus on the bulk sale of secondhand books. “I specialize in rare and old books,” he explained, emphasizing that bulk orders are rare in his line of work. “These items are costly and unique, so I don’t engage in mass trading.”
The fascination with books by AI entities is revealing itself in understood yet alarming ways. The initial email didn’t mention AI or clarify the intended use for the books, but it comes at a time when AI firms are increasingly seeking high-quality texts to enhance the training of large language models. The previous summer, scrutiny over the practices employed by AI companies grew after reports emerged detailing Anthropic’s acquisition of millions of physical books, from which they stripped bindings, scanned the pages, and discarded the originals to establish a searchable digital library for their models as part of what they termed “Project Panama.” This method, labeled “destructive scanning,” involves disassembling books to feed the content into high-speed scanners.
Ultimately, the legal issues surrounding the practice found resolution when a federal judge determined that the use of legitimately purchased books for AI model training fell under fair use provisions within copyright law. Claims regarding Anthropic’s downloading of materials from online libraries like LibGen and PiLiMi were also settled.
An Anthropic representative noted that their Claude models rely on a composite of publicly available web data, paid datasets, and proprietary data. They acquire books through standard commercial channels, ensuring that none of their data acquisition programs involve rare or antique volumes.
While court documentation illustrates the aftermath of book procurement, details about how these firms initially accumulate millions of physical volumes remain somewhat elusive. Earlier this year, 404 Media highlighted how ISBNdb, a service that typically manages metadata for International Standard Book Numbers, had begun promoting bulk book sourcing services tailored for AI labs, with orders ranging from 1,000 to 1 million volumes. Their website indicated that this service was designed to meet "LLM training needs" and catered to the scale demanded by AI.
However, the pages advertising this service have since been taken down. ISBNdb later clarified to Fortune that the service was never officially launched, asserting they had not acquired or scanned books for AI companies, and the webpages were more of a conceptual exploration than a business initiative. Similar stories have surfaced in other regions, including Germany and Switzerland, where reports indicated that secondhand booksellers received peculiar bulk orders that did not align with typical buying patterns. A Canadian entity, Zoom Books, mentioned in these allegations, denied any wrongdoing, stating that their acquisitions were part of standard recycling and trading practices.
This email provides a rare insight into the physical supply chains that could eventually fuel AI models, shedding light on the complexities of an industry often dominated by algorithms. Fortune attempted to reach out to 2077AI for comments regarding the intent behind the procurement request and the role of the books in AI training, but they did not respond before publication.


