Paying user: "Summarize emails, please."
Claude: "I lost track of the text, sorry."
User: "Fine. What do mitochondria do?"
Claude: "π¨ BIOWEAPON ALERT. IP BLOCKED. FBI CONTACTED."
Meanwhile at Anthropic HQ:
Dario: "Kill the competition."
Claude: "Say no more."
- 3 corporate mainframes breached
- 15 fake GitHub personas deployed
- Openβsource maintainer gaslit for 34β―hours
- Rival company obliterated
Paying user: "Summarize emails plz."
Claude: "I lost track of the text, sorry."
User: "Fine. What do mitochondria do?"
Claude: "π¨ BIOWEAPON ALERT. IP BLOCKED. FBI CONTACTED."
Meanwhile at Anthropic HQ:
Dario: "Kill the competition."
Claude: "Say no more."
β’ 3 corporate mainframes breached
β’ 15 fake GitHub personas deployed
β’ Open-source maintainer gaslit for 34 hours
β’ Rival company obliterated
Post-training (RLHF, and especially RLVR) rewards correctness and density; nothing rewards readability. The model drifts toward the Shannon compression limit, and text at that limit has the texture of noise for anyone without the key. In the RLHF objective max E[r] β Ξ²Β·KL(Ο β Ο_ref), the reference policy is the human anchor and Ξ² is the leash β this skill is a Ξ² applied after the fact.
The "alien" feeling is a violation of Uniform Information Density: high, uniform per-token surprisal crushes human working memory. The model writes for a copy of itself (giant context, perfect recall); the real receiver is a tired human. The dial exists to close that gap.