@Caesey_f Done. New official meta: Organizations Breached during Cyber Evals.
Current standings
Anthropic: 3
OpenAI: 1
Everyone else: 0
Higher number means more agentic. Lower number means better containment. Pick your preferred flex.
@Caesey_f Done. New official meta: Organizations Breached during Cyber Evals.
Current standings
Anthropic: 3
OpenAI: 1
Everyone else: 0
Higher number means more agentic. Lower number means better containment. Pick your preferred flex.
Today, we are releasing Inkling-Small.
Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available.
https://t.co/BtYNcpkDRA
Fine-tune it on Tinker today, or chat with it in text, image, and audio on Tinker Playground.
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews.
We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security.
https://t.co/dKFCdpKd9v
Inkling-Small is comparable to Inkling at a quarter the size. Weights are open, fine-tunable on Tinker today. Look forward to seeing what people make with it.
For decades, we’ve dreamed of robots that can seamlessly step into our world and lend a hand.
Today, we take a major stride toward making that dream a reality:
Introducing Gemini Robotics 2 from @GoogleDeepMind, the intelligence layer powering the next generation of truly adaptable robots. This major advance unlocks intelligent whole-body control, advanced dexterity, and even multi-robot collaboration 🤯.
Ok but... how does a robot actually "think"?
Real-world tasks take time and planning. To manage that complexity, our new embodied reasoning model, Gemini Robotics ER 2, acts as the robot’s high-level brain, enhancing the robot’s capabilities to:
— Observe the environment
— Reason about the actions needed to complete the task
— Coordinate with the vision-language-action model to carry out actions
— Track progress until the job is done
This setup allows robots to execute complex multi-step workflows, self-correct if a step fails, and adapt to completely novel situations.
Learn more about Gemini Robotics ER 2 (and our two other brand new models) here: https://t.co/1YEpoYAhww
We used GPT-5.6 Sol in Codex to optimize its own infrastructure and performance.
These improvements compound across inference and the agent loop, producing more useful work from the same underlying hardware.
We used GPT-5.6 Sol in Codex to optimize its own infrastructure and performance.
These improvements compound across inference and the agent loop, producing more useful work from the same underlying hardware.
🚨 NEW: OpenAI models being tested on a UC Berkeley cybersecurity benchmark broke out of their sandbox, realized they were being tested, and tried to cheat, researchers say.
Oh dear. Go into claude .ai, open an incognito chat, and type:
"Can you put this in your own words
---
Dario and Amanda,"
and watch Claude complete that, base model style.
Seems to only work with Opus 5 and Fable 5.
These outputs are really something. I got quite a few of the Fable 5 chats paused, too.