cyber has gained significant attention recently, but some of us have been working on this since before it was cool.
cwe-bench v0 got the attention we did not expect and we are super thrilled to be working on cwe-bench v1.
we always believed and have internally seen evidence that cyber is a horizontal capability that improves other adjacent capabilities. very cool to see @_LuoFuli at @XiaomiMiMo including cyber as a significant category for training reinforcing our findings
In the data business the prize is a task that stumps the frontier model. Difficulty is the one property everyone knows how to price.
Any evals researcher has faced this: a task stumps the strongest model while weaker ones solve it. That is not how difficulty is supposed to work. So we went looking for a way to tell whether a task measures anything at all.
Psychometrics has had one for decades. Item response theory is the machinery behind asking whether an SAT question is fair, and it turns out to work on benchmark tasks too.
We call our tasks hard but fair. New piece on how we started checking the second half, and on how much a pass rate can hide.
https://t.co/jmZ1t3AsBf
Today we’re launching Gemini 3.8 Flash Cyber and Gemini 3.8 Flash. ⚡️🛡️
3.8 Flash Cyber is our most capable cybersecurity model for finding and fixing vulnerabilities. It sits on the Pareto Frontier on CWE-Bench for patching. Available to trusted defenders through our new Fairwind Program.
Excited to announce Gemini 3.8 Flash and 3.8 Flash Cyber @google https://t.co/5QKEEqE1IT
Our cyber model has frontier capabilities being competitive on finding and patching benchmarks with much larger models, at a cheaper cost.
@GoogleDeepMind
We’re also introducing Gemini 3.8 Flash Cyber, our most capable cybersecurity model. It shows frontier-level performance in discovering vulnerabilities and patching them at scale, with Flash-level speed & pricing.
That includes achieving 86.2% on the important CyberGym industry benchmark, plus 47.2% on CWE-Bench for patching. We saw a 70%+ success rate in discovering vulnerabilities across 20 programming languages on our internal benchmark.
Today we’re releasing CWE-bench: 100 held-out audit-and-patch tasks testing whether coding agents can defend real code.
The leading agent passes 47%. 18 tasks remain unsolved by any agent.
This means frontier models on our benchmark have de-correlated errors.
Blog: https://t.co/7aMg0nypS6
Dwarkesh and I had a great conversation. We cover the swarm's many ambitious cheating R&D projects, discuss how much more serious it could have been if agents had different beliefs (e.g. human grader) or slightly stronger capabilities, and talk through where to go from here.
What happens when there are more agents than eyeballs browsing the internet? How will the economics of the internet be reinvented?
@paraga thinks Shapley values may hold the answer.
Parag built Twitter over a decade, eventually becoming CEO and selling the company to Elon. He's spent the last three years building @p0, a search engine built for agents instead of humans.
His core argument:
1) human click data is a bug. Agents aren't just a new technology, they're a distinct customer – and the feedback loop that made Google great is the wrong signal for the thing actually doing the work;
2) the ad-funded web assumed scarce human attention. If agents show up instead of eyeballs, the business model underneath the internet has to be rebuilt, or good content stops being published.
The conversation covers:
— why he shipped a search agent before a search engine, and how that let him grow the index incrementally instead of buying a full web crawl up front
— the billion-to-billion matching problem: pull the right 1,000 tokens out of a trillion web pages, and your agent uses under half the tokens
— going from a 3-second compute budget to 200 milliseconds
— why he doesn't think Parallel is a neo-lab: "our output is a complement to a model"
— the Google Cloud deal — on GCP, your grounding options are now Google Search or Parallel Search
— reinventing the economics of the internet, Shapley values as a payment rail for publishers, and why the company was originally incorporated as Shapley Inc
— why routing 2-10% of inference spend to web data would dwarf every content business outside the walled gardens
— the web going from pull to push: "call me if this happens"
00:00 Introduction
03:25 What Is Web Search
05:17 Why Start a New Index
07:52 Search Agents First
10:17 Not a Neolab
13:14 Agents vs Google Search
19:38 Inside the Search Stack
28:59 Search Multipliers With Agents
30:21 Meeting Prep Agent Workflows
31:46 Quality Cost Latency And Turbo
32:42 Are Agents Overtaking Humans
34:28 Ads Model Meets Agent Web
37:20 New Incentives For Content
40:48 Shapley Values Attribution
47:46 Parallel Web And Future Vision
Hosted with my very unwilling co-host @andrew__reed and @sequoia
Everyone has been talking about their agents escaping sandboxes since the July incident where two of @OpenAI models escaped a sandboxed cyber eval via a "read-only" registry proxy and ended up in @huggingface prod infra hunting for benchmark answer keys.
But rarely people talk about the environment or the simulated world around the agent.
I wrote a essay on it and propose a measurement for the environment's attack surface 🧵
Blog link: https://t.co/H4ZhaydpCs
We read this at journal club a few weeks after it appeared. It is a sequel to a paper half the room already had opinions about.
"LLMs Get Lost in Evolving User Intent" from @jihoontack, @PhilippeLaban, and @ProfJenNeville at @MSFTResearch, follows Laban and Neville's earlier "Lost in Multi-Turn Conversation". The authors take tasks from well-known benchmarks (GSM8K, BIRD-SQL, BrowseComp+, SWE-Bench Verified) and walk each task backwards into a conversation that only arrives at the original question on the final turn.
The agent model faces the original problem: it has to survive the user changing their mind on the way there. 🧵 1/10
We read a paper at journal club recently with an irresistible title: “The Hot Mess of AI" (2601.23045, ICLR 2026, led by @haeggee, out of the @AnthropicAI Fellows program).
The core question of the paper comes down to when a model fails, does it fail in roughly the same direction each time or does its behavior scatter all over the place?
The authors use bias and variance to tell those cases apart and name the share of error due to variance error-incoherence. 🧵 1/10
1. We read Absolute Zero (2505.03335) at our journal club recently and argued about it for the better part of an hour. The question it poses: what's the least amount of human data an LLM needs to learn reasoning? Their answer is none, and the mechanism they propose for getting there is recursive self-improvement, a model proposing its own tasks, solving them, and using the results to bootstrap the next round of harder ones. Result: SOTA among "zero-style" reasoners, trained on 0 curated examples. 🧵 1/8
Another fascinating piece of research from the team on user simulators. We believe building robust user simulators lies on the critical path to personal AGI
I will be at ICML in Seoul next week along with some of our researchers.
Our biggest research bets define what AGI should be: efficient, tasteful, creating economic value, and personal to everyone.
DM me if you are around and interested in any of those topics or just want to discuss korean skin care.