Excited to announce Gemini 3.8 Flash and 3.8 Flash Cyber @google https://t.co/5QKEEqE1IT
Our cyber model has frontier capabilities being competitive on finding and patching benchmarks with much larger models, at a cheaper cost.
@GoogleDeepMind
Wiz Red Agent performs penetration testing from the outside (“black box” testing) whereas CodeMender scans codebases from the inside (“white box”). Using both products offers defense in depth.
To get ahead of the increasing cybersecurity attacks using AI, we are scanning critical US infrastructure for public good using Wiz’s Red Agent and 3.8 Gemini Flash Cyber.
We’re also introducing Gemini 3.8 Flash Cyber, our most capable cybersecurity model. It shows frontier-level performance in discovering vulnerabilities and patching them at scale, with Flash-level speed & pricing.
That includes achieving 86.2% on the important CyberGym industry benchmark, plus 47.2% on CWE-Bench for patching. We saw a 70%+ success rate in discovering vulnerabilities across 20 programming languages on our internal benchmark.
Excited to announce Gemini 3.5 Flash Cyber, our lightweight Gemini cyber security model fine-tuned to find, validate, and patch vulnerabilities quickly and efficiently. Gemini Cyber offers a cost-efficient and highly capable alternative to large, costly cybersecurity models. Code security with frontier AI should be affordable. https://t.co/iKpeHcxtbX @GoogleDeepMind
Agents are finding more vulnerabilities than ever. But it turns out there are gaps in existing vulnerability discovery. Over the past 90 days vs. a year ago, web vulnerabilities (XSS/SQLi/CSRF) are down 66% and memory safety exploitability is down 3.5x.
We built the Agentic Vulnerability Coverage Map to track it all, updated daily. Introducing the Berkeley Vulnerability Initiative: https://t.co/qiZ4eThb0n. ⤵️
At Google I/O, Demis announced our code security agent CodeMender that not only automatically finds vulnerabilities, but also patches them. Finding is not sufficient to secure code. Developers are drowning in vulnerability reports, and patching is difficult. They need agents to help them patch at agentic speeds.
Well before Mythos in 2025, we introduced CodeMender, an AI agent that patches vulnerabilities. Through a series of research innovations, we got to a point that CodeMender can patch complex software.
@GoogleDeepMind
At Google DeepMind, we are invested in security & privacy for generative AI. We're growing our S&P team with 20+ roles open, spanning code security, contextual security & prompt injection defense, security post-training, model threat defense, and others. https://t.co/JQLeqzVaop
@IronRedSandHive We have designed our agent to write code changes in the style of the codebases. The agent is also including extensive validation evidence in the fixes, to explain why this fix addresses a root cause issue and it does not have unintended side effects on functionality.
I am proud to share the announcement about our CodeMender project at @GoogleDeepMind, an agent that can automatically fix a range of code security vulnerabilities. From only a modest-compute run, our agent submitted 72 high-quality fixes to vulnerable code in popular codebases, and maintainers accepted and upstreamed them.
https://t.co/TApULhUSzB
At Google DeepMind, we are excited to grow our Security & Privacy Research Team! If you are an exceptional researcher or engineer in security & privacy or machine learning, apply to one of the following:
- RS, Tech Lead, AI Based Cyberattack Defense
- RS, AI for Secure Code
- SWE, AI for Secure Code
- Sr Staff SWE, AI for Secure Code
- RS, Cybermonitoring
- 5 more to be posted soon: Sr Staff RS, Security for AI; RE, Post-Training for Security and Privacy; RE, Reinforcement Learning for Security; RE, Gemini Security; TPM, Security & Privacy Research.
https://t.co/CK0e2dmE3D
🚀 Introducing DeepSWE 🤖: our fully open-sourced, SOTA software engineering agent trained purely with RL on top of Qwen3-32B. DeepSWE achieves 59% on SWEBench-Verified with test-time scaling (and 42.2% Pass@1), topping the SWEBench leaderboard for open-weight models.
💪DeepSWE is trained with rLLM, our modular RL post-training framework for agents. rLLM makes it easy to build, train, and deploy RL-tuned agents on real-world workloads — from software engineering to web navigation and beyond.
🤗As always, we’re open-sourcing everything: not just the model, but the training code (rLLM), dataset (R2EGym), and training recipe for full reproducibility.
🔥Train DeepSWE yourself. Extend it. Build your own local agents. No secrets, no barriers.
DeepSWE and rLLM mark our major shift: from training language reasoners to building language agents that can truly learn from experience.
We believe the future of AI lies in experience-driven learning — and we’re here to democratize it.
Welcome to the era of experience. 🌍
Links below: (1/n)