this was super fun evaluating 10 of the strongest open-weight and closed models to find the best at cyber with @pilvar222
2026 is the year of open weights. paired with a harness that knows how to play to their strengths and weaknesses, open models can deliver stunning performance at a fraction of the cost of closed ones.
ps - scroll to the end of the blog to find a relic from our saga :p
We burned 11.7 billion tokens to find out which AI models can actually find vulnerabilities. 10 models, three runs each, 32 freshly disclosed CVEs to rediscover.
Open weights won this round. DeepSeek V4 Pro found the most, 28 of 32, ahead of Opus 5, Grok 4.6, and Sol. Qwen, Kimi, and GLM-5.3 right behind.
Full breakdown of all 10, by @iminurputer and @pilvar222: https://t.co/CwJO7tZnhR
👀🧵
underrated pro of agents is they basically replace devex/internal tooling budget and team. i do a LOT of research as part of my job and have my agents build me tools on the go internal to me that help lots in deeper analysis and repeated debugging
Big W for people saying that distillation was mostly an SFT artifact, and how its unclear how it would become increasingly influential in an era defined by RL.
I can’t stop thinking about this picture.
And about the possibility that math helped me more way more than I ever realised.
Perhaps, I am in fact, a living example that math is one of the most effective drugs this earth has to offer.
I finished my degree with 4 math subjects (from actuarial to analysis - don’t ask), I had a few months to go through so much math, a part of me felt it would be impossible.
So I spent months, doing nothing but math for 5+ hours each and every days for months, every single day.
Right after that, I started my watch company, Viktor Watches, and I never once noticed how easy it felt to work for 10+ hours day in, day out for months compared to math.
I never connected those two periods of my life until now. Math probably played such a substantial part in my life, and never once gave it any thought before seeing this…
Maybe spending months starting at problems you have absolutely no fucking clue on how to solve, until you… eventually do..
Really does wonders for you.
In ways you can’t even imagine.
@mehulmpt Oooh so these guys are th reason I'm getting AI calls. I make sure to waste tokens deliberately whenever I get called by idfcbank almost made an agent speak in python
@xeophon@pilvar222 indeed, we go through traces by hand and also needless to say our env is setup such that the agent is pacified with select tools at its disposal and our methodology makes sure we test code security reasoning capabilities of the model rather than brute force search
We ran DeepSeek V4.1 Flash on our cybersecurity benchmark, it is now A LOT better
- On a single run, it now rediscovers 65.6% of the benchmark's recent CVEs, up from 55.2% for the old version
- At pass@3, it catches 84.4% of them, up from 75%. This is even more than frontier models like Grok 4.6, Opus 5, and GPT-5.6-Sol
- Precision also went up, from 73.8% to 78.9%. The model is less noisy, it now reports less false positives
- The price per task also reduced. The model performs more actions per turn, and ended with a better caching hit rate 94.5% -> 95.5%
@deepseek_ai is iterating fast, and Flash is now matching frontier models for 50x cheaper (the cost per task difference is huge)
Note that this is a preview version (deepseek-v4.1-flash-expires-on-0910). The official release will likely be even better.
1/3 🧵
I was an intern under Seb in 2020. Unfortunately, the allegations of unscrupulous behavior is 100% believable. I am glad that the mask is off publicly. I really hope he doesn't weasel out of this.
https://t.co/Pe6e6cK920