I Build AI Infrastructure for a Living. The Dutch ‘Spam’ Book Order Exposes a Secret No Tech Firm Will Admit.

(SeaPRwire) –

By: Ethan Gallagher

I’ve spent 12 years designing the physical infrastructure that powers large-scale AI models, and I can count on one hand the number of times my colleagues have brought up where training data actually comes from. We debate GPU cluster sizing, token throughput, inference latency — all the shiny, digital parts of the job. The Dutch antiquarian bookseller in Haarlem who dismissed a 3,000-book order as phishing? He just dragged the industry’s dirtiest open secret into the light. This isn’t a random scam attempt. It’s a tiny, visible crack in a global, largely unregulated physical supply chain that feeds millions of books into AI training pipelines every year, with zero public accountability. For all the talk of AI being a “weightless” revolution, it runs on very tangible, often wasted, physical resources. Most of those resources never make the press release.

The on-the-record facts around AI’s use of physical books are narrow and carefully curated. Last summer, The Washington Post uncovered Anthropic’s “Project Panama” via court records. The company bought millions of physical books, cut off their bindings for high-speed scanning, then discarded the originals — a process called destructive scanning. A federal judge later ruled using legally purchased books for AI training counts as fair use, and the related lawsuit settled. Separate claims over Anthropic’s downloads of books from LibGen and PiLiMi online libraries were also resolved in the settlement. Anthropic’s official statement says its Claude models train on a mix of public web data, commercially acquired datasets, and internal data. It adds that all books come through regular commercial markets, and no rare or antiquarian books are acquired or destroyed in its programs. What Anthropic doesn’t address is why it relies on destructive scanning of physical books at all. Let’s do the math: Licensing digital copies of academic books from publishers like Elsevier or Wiley costs hundreds of dollars per title for institutional access. Bulk licensing for AI training is either unavailable or priced out of reach for even well-funded labs. Destructive scanning? A high-speed scanner can process 1,000 pages an hour, and a used academic book costs $10 to $30 on the secondhand market. Once you scan it, you throw the physical copy away — no ongoing licensing fees, no publisher tracking, no paper trail tying the digital copy to a specific purchase. The settlement also resolved claims over Anthropic’s use of pirated book sites like LibGen and PiLiMi. Those sites have gaps in recent, niche academic titles, though, and scan quality is inconsistent. Physical books give labs clean, complete, verifiable copies that don’t come with the stigma of piracy — as long as no one asks where the books came from. The fair use ruling gives legal cover for the end product, but it says nothing about how the books are sourced in the first place. That’s the gap companies are actively exploiting.

The court records only cover what happens after books reach AI labs, not how they get there. The official line from potential middlemen is uniformly dismissive. Earlier this year, 404 Media reported ISBNdb — a company best known for maintaining ISBN metadata, the unique identifiers in book barcodes — advertised bulk book sourcing services for AI labs. The service offered orders from 1,000 to 1 million books, tailored to “LLM training needs” and “delivered at the scale AI demands.” Those webpages have since been taken down. ISBNdb says the service was never launched, just an exploratory concept, and it never bought or scanned books for AI firms. Similar bulk order reports popped up in Germany and Switzerland, where secondhand booksellers got unusual requests for highly specialized titles that didn’t fit traditional collecting or resale patterns. Canadian firm Zoom Books, named in some of those reports, told Swiss broadcaster SRF the purchases are part of its regular recycling and trading model. The company behind the Dutch order, 2077AI, did not respond to requests for comment. It sought 3,001 titles, mostly published between 2020 and 2021 by academic presses including Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, with instructions to estimate shipping costs to China. None of these denials hold up to basic industry logic. ISBNdb’s entire business is built on ISBN metadata. It has no reason to draft detailed web copy for bulk book sourcing unless it had confirmed interest from paying AI clients. AI labs don’t want to put their names on bulk book orders for two simple reasons: public backlash over copyright, and the bad press of admitting they throw away millions of books after scanning them. So they use layers of middlemen, shell companies, and generic procurement requests that look like spam to small, niche sellers. The 3,001-title list sent to de Vries is no random collection. It’s all peer-reviewed academic nonfiction from top publishers, published in a tight two-year window. That’s exactly the kind of curated dataset labs use to fill gaps in LLM knowledge of recent, specialized research — content that’s scarce on the open web and spotty on piracy sites. Shipping to China makes perfect operational sense too. High-volume destructive scanning operations have lower labor and equipment costs there, with less regulatory scrutiny of how scanned content is used. The Dutch antiquarian booksellers who dismissed the order as spam? They’re the ones who don’t handle bulk inventory, so the request made no sense to them. Booksellers who specialize in high-volume used book sales? They’re already filling these orders quietly. The pay is above market rate, and the clients ask no questions about how the books will be used.

The AI training data supply chain is not a niche experiment or a hypothetical problem. It’s a fully operational, global, multi-hundred-million-dollar network that operates almost entirely outside public view. It will grow more opaque, not less, as copyright scrutiny and public pressure ramp up. Small used book sellers, independent academic dealers, and local libraries will be the first casualties of this unregulated hoarding, and no regulator in any major economy has even started tracking the scope of the issue.

Author bio: Ethan Gallagher, a 12-year Silicon Valley hardware architect specializing in AI infrastructure and large-scale training data pipeline strategy for tech firms.