🦔AI companies are bulk-buying rare books, scanning them through high-speed machines that cut the spines off, and shredding the originals. A service called ISBNdb facilitates orders of up to a million books and keeps buyers anonymous. Pre-2022 books are premium because they're free of AI-generated text. A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time. Anthropic hired the former head of Google Books partnerships to obtain "all the books in the world."
My Take
This got to me. A bookseller told 404 Media that rare books with almost no surviving copies are being fed into this pipeline. Books that survived wars, fires, and centuries of handling are being shredded so an AI can learn to write a better marketing email.
ISBNdb's website literally says "'AI company destroys two million books' is not a headline that generates sympathy," and they still built an entire business around making it happen quietly. They offer NDAs as a feature. They coach clients to call it "digital preservation."
I've covered AI companies scraping the internet, torrenting libraries, and stealing music. This is worse because it's irreversible. You can re-upload a website. You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate.
"We shred rare books and offer NDAs so nobody finds out" is a legitimate business model in 2026. What a timeline.
Hedgie🤗
@TMTLongShort Also, in the short to medium term, even if the technology continues to improve quickly, deployment is a bottleneck. Most companies are run by regular people, that don't have the technical skills to deploy and manage AI in an optimal way.
@TMTLongShort I'm partial to the radiologist analogy. I realize AI is a lot more general purpose, but it still falls very short when deployed as a replacement for most jobs. I know it's still evolving, but I don't think extrapolating out into infinity is always a safe bet.
@AndreasSteno@BowieFan2024 Anyways, that's why richer people prefer the US system. However, for the majority of voters, in the US the reason they support our current system is honestly that they've been gaslighted about "socialized medicine."
@AndreasSteno@BowieFan2024 I think starfish is right. Epidemiological stats are capturing the fact that 1$ of healthcare spend on low income people has an outsized impact. But if you get an exotic form of cancer you're going to Texas not London for that. High income people are better off in the US system.
@HedgieMarkets Seems obviously right with a long view, but does anyone really know? For sure the people at the top are politicos. They 're not going to rain on this parade. But given the ambiguity, I'm surprised anyone ever got wind of this report, and not surprised they tried to bin it.
@redflagspress Ironically, this is a great example of why socialism fails. You can't legislate honest effort in a way that works. You can only allow economic incentives to work their magic.
If you don't understand this, you will not understand why LLM-based agents are irreparably failing for a general-purpose problem solving.
An agent (by the way it was the topic of my PhD 20 years ago) to be useful, must be rational. Being rational means to always prefer an outcome that results in the maximal expected utility to its master/user.
Let’s say an agent has two actions they can execute in an environment: a_1 and a_2.
If the agent can predict that a_1 gives its user an expected utility of 10, and a_2 gives an expected utility of -100, then a rational agent must choose a_1 even if choosing a_2 seems like a better option when explained in words. The numbers 10 and -100 can be obtained by summing the products of all possible outcomes for each action and their likelihoods.
Now here is the problem with LLM-based agents.
The LLM is not optimizing expected utility in the environment. It is optimizing the next token, conditioned on a prompt, a context window, and a training distribution full of examples of what helpful answers are supposed to look like.
Those are not the same objective.
So when we wrap an LLM in a loop and call it an “agent,” we have not created a rational decision-maker. We have created a text generator that can imitate the surface form of deliberation.
It may say things like:
“I should compare the expected outcomes.”
“The best action is probably a_1.”
“I will now execute the optimal plan.”
But the internal mechanism is not selecting actions by maximizing the user’s expected utility. It is generating a continuation that is statistically appropriate given the prompt and prior context.
This distinction matters enormously.
For narrow tasks, the imitation can be good enough. If the environment is constrained, the actions are simple, and the success criteria are close to patterns seen in training, the system can appear agentic.
But for general-purpose problem solving, the gap becomes fatal.
A rational agent needs stable preferences, calibrated beliefs, causal models of the world, the ability to evaluate consequences, and the discipline to choose the action with maximal expected utility even when that action is boring, non-linguistic, or unlike the examples in its training data.
An LLM-based agent has none of that by default.
It has fluency. It has pattern completion. It has a remarkable ability to compress and recombine human text. But fluency is not rationality, and a plausible plan is not an expected-utility calculation.
This is why these systems so often fail in strange, brittle, and irreparable ways when given open-ended responsibility.
They are not failing because the prompts are insufficiently clever.
They are failing because we are asking a simulator of rational agency to be a rational agent.
@Megatron_ron Doesn't sound like he has a coherent view at all. Not saying he's wrong or right. Just seems like he doesn't know what he thinks beyond rah-rah trump.