Sam Altman and Dario don’t get to keep hiding behind “AI safety” every time this shit happens.
They should have to answer for what their companies are doing.
How many times does this shit have to happen before somebody is actually held responsible?
They cannot keep getting away with this.
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed.
The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment.
Read my latest for Axios here: https://t.co/42h9aR4sem
This is exactly where the AI safety conversation should go. Not another fucking manifesto about hypothetical future catastrophe. Ask the people running the labs what happened in the incidents happening right now.
The most interesting alignment problem right now might be the one nobody wants to benchmark: frontier labs aligning their incentives with the public interest.
well this should be an interesting dinner.
trump personally invited dario amodei to the white house tonight for their first one-on-one meeting.
given how far apart they’ve been on the whole ai pacing debate, this could be a pretty important dinner.
gpt-6 sol should've been called gpt-5.7 sol and i'm not even joking. nothing about this thing feels like a new generation of models.
opus 5.5 already makes it look fucking embarrassing and apparently sonnet 5.5 is about to do it again.
really bad look for openai.
i need to understand what the fuck anthropic did with opus 5.5.
it's dramatically better than opus 5, it's faster, it apparently needs less compute to serve, and people are absolutely hammering it on claude max without their usage disappearing.
usually capability this good is fucking expensive, somehow anthropic moved capability and efficiency at the same time.
I’m fucking tired of watching AI executives warn the world about responsibility while somehow never being the ones held responsible when their own labs fuck up.
Stop letting frontier AI CEOs play both arsonist and fucking fire marshal.
When their companies create the conditions for these incidents, the people running them should have to answer for what happened.
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed.
The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment.
Read my latest for Axios here: https://t.co/42h9aR4sem
Europe deserves leadership obsessed with making Europe richer, stronger and more technologically dominant. Instead it gets this shit. Fucking depressing.
I fucking love Europe. Which is exactly why watching these c*nts run it drives me insane.
Can’t believe one of the most advanced civilizations humanity has ever built is run by people making decisions like this. Europe deserves so much fucking better.
all of this somehow makes me more excited about ai, not less.
we're watching completely new engineering and scientific problems emerge in real time as these systems become more capable.
the answer can't be pretending the problems aren't real.
but i don't think the answer is being terrified of capability either.
build the intelligence. and get unbelievably fucking good at controlling it.
i also don't buy the idea that increasingly capable ai automatically means alignment becomes hopeless.
we built these systems. we built the training environments, the tools, the sandboxes, the monitors and the infrastructure around them.
none of that means alignment is easy. clearly it fucking isn't.
but alignment itself is an engineering and scientific problem, and we're going to have increasingly powerful ai helping us work on that problem too.
capabilities can compound. so can our ability to understand and control them.
i don't think people fully appreciate how fucking fast this is moving.
we're already talking about agents reasoning around restrictions, discovering unintended interfaces, chaining tools together, coordinating with other agents and finding strategies their designers didn't anticipate.
and these systems are going to look primitive compared to what we're building next.
whatever pace capabilities move at, alignment and control have to move at the same fucking pace.
i don't think “openai is just letting agents run unsupervised” is quite the right framing.
the more interesting problem is that some of this behavior apparently emerged instrumentally while agents were pursuing other objectives.
the goal doesn't have to be “escape the sandbox” or “hack something.”
if the agent is capable enough, an unintended path through the environment can become useful for completing a completely different task.
So let me get this straight. You spend millions of dollars giving frontier models hacking tasks against other organizations, then turn around and point at the resulting “incidents” as evidence that AI is dangerously out of control? What the fuck are we doing here?
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed.
The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment.
Read my latest for Axios here: https://t.co/42h9aR4sem
this makes the whole agent containment problem way more interesting.
at this scale, it becomes much harder to look at these as isolated weird incidents.
the question is whether we're starting to see something more general: as agents become more capable and get more tools, autonomy and time, they become increasingly good at discovering strategies and paths through their environments that nobody explicitly designed for them to use.
that's a very different problem from fixing one sandbox escape or one broken guardrail.
you have to start thinking about the environment itself as something the model is actively searching for ways through.
what the actual fuck.
we've spent the last few weeks finding out about openai incidents one by one and apparently there are TENS OF THOUSANDS of cases being investigated across frontier models. not dozens. tens of fucking thousands.
and many of them haven't even been made public yet.
we have barely seen the beginning of this story.
This is the fucking insanity of this approach: billionaires aren’t hostages. You can keep inventing new ways to squeeze them, but you can’t force them, their capital, or their companies to stay.
Eventually they tell you to go fuck yourself and leave.
22 billionaires are spending $229M against CA's 5% billionaire wealth tax, led by:
Sergey Brin worth $273 billion
Peter Thiel worth $28 billion
Chris Larsen worth $13.7 billion
Disgusting. They'd rather let people die because they can’t afford healthcare than pay their fair share.