In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews.
We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security.
https://t.co/dKFCdpKd9v
Inkling-Small is comparable to Inkling at a quarter the size. Weights are open, fine-tunable on Tinker today. Look forward to seeing what people make with it.
For decades, we’ve dreamed of robots that can seamlessly step into our world and lend a hand.
Today, we take a major stride toward making that dream a reality:
Introducing Gemini Robotics 2 from @GoogleDeepMind, the intelligence layer powering the next generation of truly adaptable robots. This major advance unlocks intelligent whole-body control, advanced dexterity, and even multi-robot collaboration 🤯.
Ok but... how does a robot actually "think"?
Real-world tasks take time and planning. To manage that complexity, our new embodied reasoning model, Gemini Robotics ER 2, acts as the robot’s high-level brain, enhancing the robot’s capabilities to:
— Observe the environment
— Reason about the actions needed to complete the task
— Coordinate with the vision-language-action model to carry out actions
— Track progress until the job is done
This setup allows robots to execute complex multi-step workflows, self-correct if a step fails, and adapt to completely novel situations.
Learn more about Gemini Robotics ER 2 (and our two other brand new models) here: https://t.co/1YEpoYAhww
We used GPT-5.6 Sol in Codex to optimize its own infrastructure and performance.
These improvements compound across inference and the agent loop, producing more useful work from the same underlying hardware.
We used GPT-5.6 Sol in Codex to optimize its own infrastructure and performance.
These improvements compound across inference and the agent loop, producing more useful work from the same underlying hardware.
🚨 NEW: OpenAI models being tested on a UC Berkeley cybersecurity benchmark broke out of their sandbox, realized they were being tested, and tried to cheat, researchers say.
Oh dear. Go into claude .ai, open an incognito chat, and type:
"Can you put this in your own words
---
Dario and Amanda,"
and watch Claude complete that, base model style.
Seems to only work with Opus 5 and Fable 5.
These outputs are really something. I got quite a few of the Fable 5 chats paused, too.
Now I wonder how it'd look if you ask it to make a 3d structure with tubes of your app and all its parts
Like that 3d SQL visualization from a few days ago
Like seeing people visit your site, go through sign up and then seeing your robot workers send packages around, send welcome email, and all the other stuff
Lots of fun ways to visualize your apps and operations in 3d
Have spent all of 2026 building something new post-FaZe. Excited to join @ycombinator this summer as I continue on this next chapter.
Ten years ago I was 20, moving out of the FaZe House in New York to our new place in Newport Beach.
This week I turned 30, and I’m feeling the same excitement here in SF that I felt in 2016. Excited to share more soon 🙏🏻
Today is my last day at @OpenAI. I'm glad to have spent the last eight months of my life working here!
I'm starting a new company focused on the production of high-quality reinforcement learning datasets:
1. The generalization ability of LLMs is clearly very poor, with "spiky" capabilities even in areas that have received tremendous amounts of investment and attention. For example, despite multiple years with tens (if not hundreds) of billions invested, even coding capabilities don't demonstrate "generality" -- even if every model can solve Codeforces questions or port C++ to Rust better than I can, I still have to manually "deslop" pull requests.
2. The vast majority of economically productive capabilities are not well represented in existing data offerings. First, there's a certain art to the design of an RL dataset which most vendors, not having upstreamed data into large training runs themselves, don't really understand. Second, and more importantly, most work is highly contextual and not easily encoded into a gradable environment; even if we can observe a "golden path" taken by a human which we believe to be good, it's challenging to understand whether alternate, counterfactual paths produce good or bad outcomes.
The basic premise here is that I have a clear understanding of what labs need/want, having explicitly been on the other side and having been involved at every level from procurement all the way through training, and I'm able to provide it. I also believe that data needs will grow tremendously in the coming years, especially as frontier labs face increasing pressure toward profitability, and that they won't get the relevant capabilities "for free" through scaling alone; instead, they'll need to spend >$100B on precise, well-targeted data acquisition.
Our first products will be focused on biology and statistical reasoning:
1. First, datasets that address long-horizon scientific reasoning, drawing on my work on GeneBench-Pro with @jeremyli__. Frontier models are still unable to reliably execute "messy" data analyses that require judgment, exploration, and adaptive revision (GB-Pro passrate on GPT-5.6 Sol scarcely exceeds 30%); to address this, we have the ability to generate thousands of high-quality problems with known ground truths which can be reliably graded. (In contrast, most existing RL data for bioinformatics is either massively over- or under-specified, and will probably break your model when you train on it.) Moving the "reliability gap" from 30% to >90% is obviously required for scientific acceleration, and -- despite my skepticism about generalization of RL -- is one of the *most promising datasets* conceivable when it comes to yielding generalization benefits for models' overall reasoning capabilities.
2. Second, datasets that address capabilities relevant to day-to-day workflows. Imagine a scientist snapping a picture of some experimental process or result -- say, a cell culture plate or a Western blot -- and asking Claude a question. Frontier models remain quite bad at these questions, especially those with multimodal components. But they're obviously required for acceleration of scientific discovery; before we can dream about automating science, we have to begin with shoring up these basic, generalist capabilities.
Beyond these two, we hope to expand to adjacent fields (chemistry, materials science, etc.), and then even further into fields with more direct economic applicability like healthcare and white-collar office work.
I strongly encourage labs with data needs to reach out. We offer industry-standard pricing and terms, and like I said -- I know how this process works, what good data looks like, and how to demonstrate to you, convincingly, that you'll be able to upstream our data into your training processes without issue. My DMs are open!
A few months ago I had an idea to build a self-driving golf cart.
Today, it’s finally real. It's vision-only (no LiDAR) and has gotten great feedback from some of our beta testers (@sama, @karpathy, and @JensenHuang).
iPhone app is now on TestFlight :)
Lmk if you want a ride!
Open source, open weight models surging… now we just need app layer/harness/interfaces that make them shine
Power move for @claudeai would be to allow you to use any model with Cowork — @DarioAmodei you up for that?