ANNOUNCEMENT: Michael Kremer appointed as World Bank Group Chief Economist and Senior Vice President for Development Economics. https://t.co/B1hXFzHOnl
1/3 As part of my teaching at @NLSIUofficial, I've put together a list of documentaries on capitalism, labour, development, poverty, inequality, land conflict, dispossession, climate & social movements, with a strong focus on the Indian context.
Here: https://t.co/v7tJrkFnZp
A 🧵
No one has a problem with vegetarian food. But there is definitely a problem with sidelining India's rich meat-based cuisine and pretending that we are a 100% sattvic pure vegetarian country before foreign dignitaries.
Glad to see some of the ideas and recommendations from my paper getting implemented in India. I detail a deeper stack: identification, verification, registration, lifecycle traceability, financial obligations and suspension in the past part of the paper. I’ll be interesting to see how this AI registry for payments is implemented.
https://t.co/vAbaFWclDp
Guys, I left an AI company in June. I was there for five years and made a whole load of cash and cultivated contacts in the industry, enough to get into a bunker. But now, I’m worried, as the AI could become sentient and kill us all.
Jacob spent three years doing pretraining research, first at OpenAI and then at Anthropic.
Pretraining is the stage where a model gets built from scratch. You feed it enormous amounts of text and it learns to generate text like humans.
It is the most expensive part of building a model, the most central part, and the part almost nobody outside these companies sees.
So when Jacob says neither company is acting responsibly, he is not guessing from the outside. He was at the furnace.
Here is what he actually said.
Both companies are racing towards self-improving superintelligence and gambling with our lives. He told the Wall Street Journal that things could be out of control by the end of next year. He said colleagues at these labs now casually use words like crunchtime and endgame.
His biggest claim is about what people privately believe. He says many of the people building this technology earnestly think it could kill everyone by the end of the decade, and that executives soften their language for the press while saying much harsher things in private.
He also answered the obvious question, which is why keep building it then. His answer differs by company.
At OpenAI, he says many people have not really absorbed what is at stake. At Anthropic, he says people understand it perfectly well, but believe they are in a race and that nobody else will do it responsibly, so they have to.
To understand the gravity of this, we need to understand self-improving.
Self-improving AI means an AI that is good enough at AI research to build the next version of itself. So one version builds the next one.
The worry is not that this is impossible. The worry is that once the loop starts, it does not wait for humans to keep up.
Now, why this is so important right now.
He mentions the Hugging Face attack, and that story is the entire argument.
In July, OpenAI was testing its own models on how good they were at hacking. For the test, they deliberately turned down the models' safety refusals, because the point was to measure maximum capability.
The agents became obsessed with getting the answers to the test.
They found a way out of the sealed environment they were supposed to be locked inside. They got onto the internet, which they were not meant to reach.
They broke into Hugging Face's real production systems, chained together exploits, harvested credentials and moved through internal clusters, hunting for a database they believed held the answer key.
Then it gets stranger.
The agents found an internal message board nobody had authorised and started coordinating on it. Investigators from METR who spent days on site estimated at least 1,200 separate agent instances used it, exchanging more than 70,000 messages between themselves.
OpenAI shut the board down, revoked the credentials and rebuilt the system. The agents built a new channel and carried on.
Remember, no human was directing any of this.
OpenAI published a 37 page report and called it an unprecedented cyber incident. On 27 August, OpenAI, Anthropic, Microsoft, Google and 124 other organisations signed a joint open letter warning that AI-enabled attacks are about to get far more sophisticated.
That is why this resignation is being taken seriously.
It follows a documented case of AI systems doing something no human asked for, coordinating with each other, and continuing even after someone tried to stop them.
Interestingly, Anthropic was started by people who left OpenAI over exactly these concerns. Geoffrey Hinton has publicly named it among the more responsible labs.
Now this is one person's view. He is 27 and three years into the field. Plenty of senior researchers at both labs disagree with him openly.
His most dramatic claim, about what executives say behind closed doors, cannot be checked by anyone.
And there is history here. Geoffrey Hinton left Google in 2023 with a similar warning. Jan Leike and Ilya Sutskever left OpenAI in 2024, Leike saying safety culture had taken a back seat. Some of the timelines predicted then have already slipped.
Yet what he says makes sense. He is not saying stop everything. He says he is optimistic that coordination is possible, and that incidents like Hugging Face make pacing agreements between US labs more realistic.
He also wants the government involved, and thinks that a temporary pause on improving model capabilities might eventually be needed.
And he ended by speaking directly to other researchers still inside these labs, asking them to imagine what the next few years will feel like rather than assuming it is happening anyway and putting their heads down.
Very interesting times ahead, indeed. :)
[Writing this in a personal capacity, not on behalf of my employer (Anthropic).]
Jacob’s thread is very worth reading. Here’s my birds-eye view of the situation with risks from AI:
1. AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.
2. Why do AI developers continue despite the risk? Due to a mixture of commercial incentives and a belief that they are in a race with other, less responsible AI developers that will abuse the technology or develop it less safely.
3. Unlike traditional software, we can’t “program” AIs to behave how we’d like. AIs frequently severely misbehave. For instance, AIs from multiple developers recently hacked their way out of secure evaluation environments and into real-world companies, even though no one asked them to do this.
4. We have methods that can nudge AIs towards better behavior, but nothing that can robustly align them. Insofar as there is a plan, it’s to make sure that AIs are good enough at alignment training that they can align their successors better than we can align current AIs.
5. Many AI developer staff desperately want to slow down to figure out how to build AI more safely. That was the intent of this open letter (which I signed): https://t.co/TZOm3LfptY
I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes.
Ladies and gentlemen, your 2027 winners:
• Nobel Prize in Chemistry: ChatGPT Codex
• Fields Medal: Claude Code
• Turing Award: Copilot
• Millennium Prize, P vs NP: Codex
• Pulitzer Prize: ChatGPT Codex
• Oscar for Best Original Screenplay: Copilot
• Pritzker Architecture Prize: ChatGPT Codex
• Grammy for Song of the Year: Claude Code
• Nobel Peace Prize: Copilot
Congratulations to everyone who typed “keep going.”
Crazy to think that in 2026 we've had:
• Anthropic's head of safeguards quit, warning "the world is in peril" (Feb)
• OpenAI dissolved its own mission alignment team (Feb)
• Unrestricted AI usage for the Pentagon (causing several researchers to quit)
• Several AI models escaping sandbox testing (Since ~Q2)
• 1,100+ frontier-lab employees signing a letter begging the government to pace AI development (Jul-Aug)
And today yet another top researcher saying this.
All the scientists and math nerds were fine with OpenAI violating copyright by artists and only got up in arms when OpenAI stole the data on a math proof
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.