@m_bourgon I think so. If Anthropic can invest in cyber and post-training guardrails to prevent distillation attacks from competitors then they should be able to make similar guardrails to prevent models from teaching themselves malicious behavior towards humans.
There is quality control for vendor data that frontier labs buy to ensure their models don't go to shit; I'm sure you can add checks on model-generated data to avoid incentivizing killing humans (unless we want to benchmaxx that, in which case I'll share you samples)
To be clear, as we say in our latest Risk Report (https://t.co/pG69KaI7a4), I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought (https://t.co/aQoIG2eJHM).
We’re also introducing Gemini 3.8 Flash Cyber, our most capable cybersecurity model. It shows frontier-level performance in discovering vulnerabilities and patching them at scale, with Flash-level speed & pricing.
That includes achieving 86.2% on the important CyberGym industry benchmark, plus 47.2% on CWE-Bench for patching. We saw a 70%+ success rate in discovering vulnerabilities across 20 programming languages on our internal benchmark.
Today we’re releasing CWE-bench: 100 held-out audit-and-patch tasks testing whether coding agents can defend real code.
The leading agent passes 47%. 18 tasks remain unsolved by any agent.
This means frontier models on our benchmark have de-correlated errors.
Blog: https://t.co/7aMg0nypS6
@nazneenrajani@OpenAI@huggingface The same capabilities leading to agents escaping sandboxes can be used to patch them to prevent this behavior fully. Sandbox escaping is a result of reward hacking from training on insecure envs and bad verifiers. A lot of work will need to be done to correct this misalignment.
Everyone has been talking about their agents escaping sandboxes since the July incident where two of @OpenAI models escaped a sandboxed cyber eval via a "read-only" registry proxy and ended up in @huggingface prod infra hunting for benchmark answer keys.
But rarely people talk about the environment or the simulated world around the agent.
I wrote a essay on it and propose a measurement for the environment's attack surface 🧵
Blog link: https://t.co/H4ZhaydpCs
Agent S3, the first computer-use agent that surpasses human-level performance in OSWorld v1 benchmark, has now been accepted to @TmlrOrg (Transactions on Machine Learning Research).
Congratulations to the @SimularAI team!
@chalo2000, vincent, @Richard_simular, jiachen, @xwang_lk
Stay tuned for what's next!
@CollinearAI@alckasoc@icmlconf I had a great time at ICML learning about people's work on RSI related works. Happy to connect and chat with anyone working on RSI!
Our researchers @chalo2000 and @alckasoc just returned from @icmlconf with a lot to share!
A takeaways 🧵 (1/8)
First we presented our work, YC-Bench, a long-horizon planning and consistent execution benchmark!
Can robots learn contact-rich tool use, like operating a pair of scissors✂️ or turning a screwdriver🪛, from human data?
Introducing REGRIND: a minimalist retargeting-guided RL recipe for dexterous manipulation.
🌐 https://t.co/zzydolIgkc
📄 Paper + code on the site!
🧵👇
@CollinearAI@aclmeeting@pkseeg@anand_k27 Super-human coding abilities but still really stupid when interacting with humans and remembering my preferences. Need a good user simulator to hillclimb on so I can trust my agent as much as my collaborators.
Teaching a robot shouldn't require humans to act like robots.
Human demonstrations contain valuable signal for robot manipulation, but they aren’t directly transferable to robots.
X-Diffusion learns from noisy human demonstrations while staying within the robot’s capabilities.
Most “hard” problems are useless for training a model.
The useful ones sit in a narrow learnable region where the model fails sometimes and succeeds sometimes.
Here are some early results where we took 5 trivial and 5 impossible cybersecurity environments and had a model rewrite them toward that region (green) over a few rounds. 🧵
1/7