I'm a second-generation bookseller. My family runs Houston's largest used & rare bookstore and I'm building an AI tool for used bookstores. We got hit by exactly these orders, including a single order for 70 obscure books that made us pause online sales entirely. So I dug in. I don't think this connects the way people want it to.
Every ingredient of the viral story is true. ISBNdb really does advertise sourcing books for AI companies. Anthropic really did destructively scan millions of books though they were mostly acquired through things like library deaccessions, not bookstore inventory. And booksellers really are getting bizarre bulk orders. But "rare" here isn't what you're picturing. It's The Insider's Guide to Metro Denver (1995) and how-to-use-WordPerfect 1991 manuals. Obscure, but often not precious.
When we combed through our orders, nearly every book bought from us was either unavailable on Amazon or listed there at 5, 10, 20x our price on the platform they were purchased on. And the shipments are going to FBA prep companies: which is what you do with a book you're about to resell, not one you're about to cut apart. There is more than circumstantial evidence that this is a well-funded company running algorithmic arbitrage, buying underpriced or out-of-stock books to relist on Amazon at markup. This evidence includes an official denial that they are selling to train AI but I don't want to mention them for several reasons. You can Google it.
That's more mundane than the shredder story. The actual risk is worse, though, because it's quieter: books that don't sell via FBA eventually get liquidated. If we really hold the last copies of some of these (and comparing platforms, it looks like we might) we're feeding them into a portal they may not come back from. No AI company required.
Everyone assumes the internet preserved everything. It didn't. Many of these books have almost no metadata anywhere, just a scattered listings across proprietary databases then nothing. That's what my project is for: building better bibliographic records for these books than exist anywhere else, so we at least have a map of what we stand to lose. So the Library of Alexandria is still burning, just more slowly, and kore from “who cares” than from some cinematic villain.
(I am open to being corrected.)
Full writeup: https://t.co/tHXMH0L3us
On a beaucoup trop peu parlé du vol du torque en or de la Dame de Vix, il y a trois jours, au musée de Châtillon-sur-Seine. La perte de ce chef-d’œuvre de l’art celte m’a personnellement bien plus ému que celle du diadème de l’impératrice Eugénie lors du cambriolage du Louvre, car le torque a tellement plus de valeur esthétique et historique. Créé au Ve siècle av. J.-C. pour une princesse gauloise, il témoigne de la sophistication d’une orfèvrerie jonglant avec virtuosité entre les traditions grecque, étrusque et celte. On y admire en particulier le petit cheval ailé posé sur un entrelacs de fils d’or, lui-même soutenu par une patte de lion. Ces 480 grammes d’or, par la cupidité qu’ils suscitent, auront peut-être eu raison de la forme artistique dont ils n’étaient que le support. Tout cela risque en effet de finir en lingot. Pour nous consoler, il ne nous restera que des photographies et des reproductions. Où sont les barbares, au Ve siècle av. J.-C. ou au XXIe siècle ?
btw anthropic's internal document on this literally said "we don't want it to be known that we are working on this.”
it was called project panama.
here's exactly what happened:
1: anthropic concluded that books were the cheapest way to build a world-class model because they gave claude curated facts, structured arguments, compelling stories, and writing “an editor would approve of.”
2: once anthropic decided it needed books at enormous scale, its first solution was piracy.
it downloaded 7m+ books from online libraries including libgen. the judge later wrote that although anthropic had legal ways to buy them, it chose piracy to avoid what dario amodei called the “legal/practice/business slog.”
3: that piracy created a massive legal risk.
so in february 2024, anthropic hired tom turvey, the former head of partnerships for google books, to find a legally safer way of obtaining “all the books in the world.”
4: turvey first contacted major publishers about licensing their catalogs.
those attempts didn’t produce agreements, so anthropic chose a route that required no publisher permission: buying millions of physical books through distributors and used-book retailers.
5: within about a year, anthropic spent tens of millions acquiring and scanning millions of books, including many rare and 1/1 titles. one vendor proposal targeted 500,000 to 2 million books in six months.
6: to scan that many books within months, the vendors physically dismantled them.
a hydraulic cutter removed each spine. the pages were trimmed to size, fed as loose sheets through high-speed industrial scanners, and converted into searchable PDFs. the paper remains were then sent for recycling.
7: these PDFs were fed into claude as training data.
the complete collection became a private, searchable anthropic library that the company planned to “store forever.” the scans aren’t available to the public and were never open-sourced.
The new Career Scenario "Brighter Together: Our Grand Concert" is here!
Follow & repost for a chance to win a $50 Gift Card, or 5,000 carats!
🥕 Play Now: https://t.co/HlXI094XKj
🥕 Details: https://t.co/M74LNocwrg
#Umamusume#GrandConcert
captured an enemy knight and he started talking about some “prisoner of war” nonsense like brother we’re the wretched blades of skar we’re straight up killing u