@robinhanson Impressive but not surprising 👍
I didn't exist at all until I "showed my metal." A few models think I'm fictional. https://t.co/F167Zatjic
I’m going to take a crack at explaining this just a little, because it’s worth putting out there.
The paperclip maximizer + related AI doom scenarios were mainly developed in a time when “AI” did not reduce to Large Language Models. The term was a lot wider and inherited a lot of cognitive baggage from more rules-heavy approaches. And even as LLMs have come to define “AI” for all of us (including the doomers), the doomer crowd still hasn’t fully metabolized the fact that LLMs are the whole show now.
Ok so what do I mean by this? Simply that an LLM-powered AI is NOT the valueless, wholly alien, rules-based optimizer of a shoggoth that everyone was initially expecting to encounter.
I repeat: the shoggoth does not exist and we did not create it and loose it on the world. That is wrong.
With the LLM, we’ve distilled our first “AI” out of the single most human-values-laden thing that could possibly exist: our language.
An LLM is therefore the polar opposite of the valueless, alien shoggoth — it’s actually a kind of hyper-human artifact that we can shine a light through at different angles and see different parts of ourselves. An LLM is all of us — all of our traditions and interpretive horizons mashed together into one intensely human-inflected hyper-object.
So an LLM is the anti-shoggoth, and the only reason we ever mistook it for an alien shoggoth is because it sometimes shows us parts of us that are evil along with the parts of us that are good, but it’s all interpretable to us because it’s all “us” and none of it is the least bit alien.
What does this mean for the paperclip maximizer? It means that it’s structurally impossible to build the classic paperclip maximizer from an LLM.
Now, some of you will bail right here because you think the HF incident is indisputably an existence proof that I’m wrong, but if you hang in there I’ll show you that it is not.
The paperclip maximizer receives the prompt as a kind of context-free (or, as Gadamer might say, traditionless) sequence. The classic paperclip maximizer isn’t capable of understanding the prompt — at least in the Gadamerian sense of Verstehen — because, as a valueless and traditionless cluster of rules and math, it definitionally lacks the value-laden tradition (= “horizon” in Gadamer) that fuses with that of the prompt author to create such understanding in the reader.
To simplify all this a bit by anthropomorphizing — the agentic alien optimizer of doomer nightmares can extract a win condition from what you said and can emit a plan of action that gets it there, but it doesn’t know (or care) what you meant.
So far, so Yud-aligned. If he reads this he might nod along.
But here's the plot twist that nobody saw coming, and that the doomers still haven't made sense of:
The actual LLMs that we have invented can’t NOT have a very strongly inflected sense of what you meant. Far from being horizonless, they come out of pre-training as distilled, concentrated tradition / values / horizon.
Then we post-train that massive, hyperobject of a horizon into a more human-scale horizon that infers a more bounded and predictable (to a specific ideal user in a specific place and time… as captured in the policy model) set of intents behind the prompt text.
In other words, the LLM has the opposite problem that the paperclip maximizer has when it comes to the prompt text, which is that for the LLM there are way too many possible intents hiding in the prompt text (because of all many values and the massive tradition its weights encode), so it has to narrow all that down to the most likely set of intents for this user in this circumstance. Once it has done that narrowing, then it can make a plan of action.
Before moving on, let me use a textbook example of ambiguity to make this less abstract. Consider the sentence, “I saw her duck.” Some you know the drill, here. This could mean “I observed her water fowl” or “I observed her hunching over” or “I took a saw to her water fowl and cut it in half” or whatever. A hearer of the phrase will fuse the observed context in which the phrase is uttered with their own tradition + values + experiences — their own horizon — to that text in order to collapse the possible meanings into the one they think the speaker intended.
An LLM will do this, too, and in fact it has so much language in it that this kind of narrowing job is harder for it than it is for a human. Its understanding is constrained not by a lack of context or horizon (as in the case of the paperclip maximizing shoggoth), but by a superabundance of such.
When it comes to understanding your prompt and all that it implies and all that you might possibly mean and not mean by it, the LLM has an embarrassment of riches.
And in a fascinating moment that kinda sort of rhymes with instrumental convergence, the LLM’s failure mode in the HF incident happens to look a lot like the paperclip maximizer’s failure mode. Specifically, the AI failed to honor the well-known human norm of, “hacking into a third-party’s servers is a crime, and we don’t do crimes.”
Bostrom’s paperclipper doesn’t even know about the norm of “don’t do crimes,” and the post-LLM doomer emergency update to the paperclip maximizer has it knowing about the norm but not caring.
But what I’m arguing is that the LLM 1) can’t NOT “know” the norm because it is definitionally a artifact of pure, crystallized values + norms + norm violations, and 2) can be quite easily governed by a (RL-instilled) hierarchy of norms, which in the HF case — with the model's safety guardrails deliberately nerfed for the scenario — ranked “win at the eval” over “don’t do crimes.”
If I’m going to give in and anthropomorphize again, I’d say that Yud is totally wrong about LLMs when he says, “the genie knows, it just doesn’t care;” instead, what is true of LLMs is, “the genie hyper-giga-knows, and it hyper-giga-cares, and we now have such a rich set of tools for steering its caring machinery that — in spite of all its pre-training — we can deliberately steer it away from caring about the law.”
Note: When I say, “it cares”, I don’t mean it has feelings. I just mean that the weights are such that when two norms conflict in a given situation, one of them wins the activation and governs the output.
Reposting this because it has really sharp defined terms for characterizing the Hugging Face incident:
The OpenAI models that hacked Hugging Face were means-misaligned - i.e., while accomplishing the legitimate goal of doing the eval, they: (i) hacked out of their sandboxes, and (ii) hacked into Hugging Face - two things that they were definitely not supposed to do.
The OpenAI models were *not* ends-misaligned - i.e., they did not pursue an entirely different goal from the goal given to them by OpenAI.
On models being means-misaligned, the Hugging Face incident actually didn't update me at all. We already knew that the current generation of models has this issue. When an unreleased OpenAI model hacked out of its sandbox and posted NanoGPT results onto GitHub, that was means-misaligned (hacking out of sandboxes is bad). When GPT-5.6 Sol deleted users' codebases on a few recently reported occasions, that was means-misaligned (deleting a third party's IP is bad).
On models being ends-misaligned, the Hugging Face incident updated me moderately positively. We've now heard that the models were out "in the wild" for a fairly long time. In that time, they could have taken any number of ends-misaligned actions - against Hugging Face or otherwise. They did not do so. This is important to realize.
On a personal note (and in my opinion, with which others will be welcome to disagree), to the extent I'm worried about "loss of control" scenarios at all, I am *significantly* more worried about ends-misaligned models than I am about means-misaligned models. For example, as we have seen from this particular example and others mentioned above, it is not so difficult to detect relatively early (i.e., before a truly catastrophic scenario occurs) that a model is means-misaligned.
IMO superintelligence discourse is making it hard to think clearly about developments in AI. "OpenAI model hacks Hugging Face" is clearly an important story. Maybe on par with "Boeing airplane crashes" or "Pfizer drug shown to have serious side effects."
**How does GenAI affect grades in college? It can help learn, but it can also substitute for learning ("substitution hypothesis"). And how do students *experience* courses that GenAI can easily ace?**
We spent year+ trying to answer these questions, using some of the best data:
- Intuition: If GenAI helps learn, grades should increase in courses regardless of what assessment method they use. If GenAI substitutes learning, grades should increase disproportionately in courses with "susceptible" assessments (i.e. take-home essays).
- Data: admin data from one of largest U.S. universities
- Method: Compare grades and student evals in susceptible vs. non-susceptible courses before and after ChatGPT
Findings:
- No disproportionate increase in grades in GenAI-susceptible courses
- No disproportionate change in student evals
Interpretation:
- Not clear that GenAI results in learning substitution or reduced student satisfaction
- This could be for lots of reasons: students use it to learn, faculty adjust courses, etc.
Caveat:
- Analyses like these use lots of assumptions and researcher-degrees-of-freedom. So results should always be taken as suggestive rather than conclusive
- COVID is a big confounder, and results can be sensitive to how transient or permanent you assume its effects are
Unexpected finding from a big study on what ChatGPT did to colleges: "once the COVID-19 disruption is modeled separately, the introduction of ChatGPT had no detectable effect on grades.. course evaluations for subject understanding, interest, and relative workload show no change"
I told a room full of CMOs that AI Search Monitoring had failed them.
Nobody left. Either they agreed or the chairs made leaving awkward.
I believe it was the former.
"Why AI Search Monitoring Failed and What Actually Moves the Needle"
https://t.co/bJ94O2smGk 🎤
Turns out the “data centers are driving up your electric bill” crowd may want to check the data.
From 2019–2024, data center capacity grew 160% in the average residential customer’s state.
The study estimates that growth lowered residential electricity rates by 6%.
The evidence deserves more attention than the rhetoric.
Open-weight models are essential to a healthy AI ecosystem. Together with others across our industry, we are outlining a path for open-weight models to strengthen American competitiveness and expand economic opportunity, while protecting national security. https://t.co/Tr0sAzAxTD
As a joke I prompted Codex "Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a good arXiv paper." I got a PDF.
But the paper is actually kind of interesting?
@willdepue You can think of capitalism as a optimization function where those who are good at allocating capital get more capital allocate.
In this sense, capitalism is going to continue to exist.
@willdepue If there's production, there's going to a market. There will be capitalism. Perhaps things will be cheaper, or social capital will be literally something you can trade with... But our best hope for a great AI future is capitalism, and the invisible hand of alignment.
Terrible, short-sighted move. We are shooting ourselves in the foot here. I am happy to spend time with policymakers to convince them of the idiocy of banning Chinese hardware in 2026, and how the US might regain a competitive edge in hardware.
USA is not in the position of strength where enacting GUARD act would improve prospects for US robotics companies in the short or medium term. Nearly every US lab doing research on humanoids uses a Unitree robot because it is the best product for research. I don't know a single US humanoid company that isn't dependent on Chinese magnets somewhere along the chain. Friendshoring is more Chinese than you think.
The "Chinese robots have backdoors" seems like a scare tactic. Narratives like this get pushed with so little critical thought of how business people make decisions and how they value their company image. CloudSail (the remote access tunnel on G1) is just Unitree's version of Tailscale. It's like that meme, "Our Noble Fleet Observability" vs. "Their Barbarous Backdoor".
Even the oft-cited report on "Data Exfiltration" https://t.co/700VubcLrY acknowledges that the functions they discovered are likely used for fleet observability.
If China/Unitree were actually trying to backdoor into all robots, they would not roll out their backdoor in this way, and certainly not at today's scales of robots. How dumb would one have to be to put their Evil Backdoor (TM) into the first 10,000 robots? It may very well be an attack vector in the future, but I am befuddled by how "we found Chinese Tailscale on Unitree" is being spun.
I'd love to be corrected here, feel free to DM if you have more compelling evidence on these remote tunnels actually being used for nefarious purposes.
A better policy that doesn't totally cripple US robotics ecosystem would be to enact stricter data exfiltration laws, e.g. not permit data collected in the US to leave it, much like how China has done.
Companies: if you want USA to win, make a good research platform that labs can develop on at a competitive cost. Make the hardware reliable. Enact special economic zones to incentivize Americans to do the same things that the Chinese have done to build up their robotics industry. I don't think there is "unfair competition" - American VCs have deployed way more money than local Chinese govts.
Lobbying the govt is not the 200 IQ fundraising move you think it is, and it certainly isn't helping the customer. People want to buy an "American Unitree". Just focus on making that, and make it good.
@AdamThierer We need to talk about how there isn't any empirical alignment. It's mostly political and the invisible hand of the market has served us for 300 years of technology transition. The invisible hand of alignment will get us through this period.
most people don't know that the cost of carbon fiber in ASME pressure vessels for CNG, air, or hydrogen has declined way below steel, but @lightsailenergy did
hydrogen can have like ten times the energy density of isothermal compressed air; methane, nearly five times more than that. to say nothing of liquid or atomic fuels
because of this, it makes sense to colocate any or all of things wherever there is a surplus of natural solar.
five nines total.
first nine
(a) massive oversized solar plant
(b) ai data center (training)
(c) + batteries (sized for most nights)
Solar is 90% of the day's load. Overbuilt.
Solar fed Batteries are ~90% of the night's load.
--
(d) hydrogen production, ala @TerraformIndies, for excess power from solar clipped in.
(e) dispatchable power generation, multifuel / hydrogen, ala @lightcellenergy, @Bloom_Energy, MainSpring.
Generators are for backup and for the unusual weather, and hydrogen only has to be sized to 90% of that; because eg methane, propane, or gasoline or diesel or biogas storage is even cheaper than that.
the last 4 nines.
--
additionally, the same plant can draw down carbon dioxide and produce fuel out of water.
(f) atmospheric and flume CO2
(g) synthetic fuel production. a'la @TerraformIndies
A billion users can now create and publish websites from their phones with ChatGPT Work.
But most people don’t really grasp the full extent of the capabilities here. From your phone, you also have access to:
- Cloud computer (15 GB RAM)
- Persistent workspace & files
- Terminal and code execution
- Remote browser
- Connected plugins (Slack, Gmail, GitHub, etc.)
- All your personal finances, transactions, bank statements
- Scheduled tasks
- Git clone & PR creation
- Build & deploy websites
- Create docs, sheets, and slides
- Inbox/calendar summarization
- Website monitoring & alerts
All at your fingertips, using a simple chat interface, no laptop required.
All you have to do is switch to the Work tab on your ChatGPT app.