We asked GPT 5.6-Cyber to escape a VM used to sandbox agents. It broke out three times.
In its final escape, the agent found three 0-days on its own and chained them into a working exploit. https://t.co/3JRVWgPxHx
Great article, thanks for sharing and typing every letter of it.
I think it captures the multifaceted dilemmas we're facing as an industry, and as individual enjoyers of the thrill of security research.
On the individual level, it's hard to accept being basically rug-pulled from doing something that we loved, and payed the bills. Doing security research, spending hours chasing bugs, is part of our identity. The joy of manual research, and the thrill you get when you find that elusive vuln won't go away.
But, what will happen when the economics of it change? Will all of us still be deep in it when the market and job opportunities shrink by orders of magnitude? What will happen in this front is still undetermined, but we'll shortly get the answer.
The one thing that encourages me to keep going, is that even though the industry is getting shaken at its core, it's never been a more exciting time. Now we all have what are basically cyber weapons at our disposal, which would have been a dream to have just a couple of years ago. With them you can dive into areas of research you previously couldn't either because of a lack of knowledge, time, or whatever, but now you can do whatever your creativity and motivation takes you to.
Another thing in everyone's favor, is that There Really is No Moat (TM). What any company/individual does in terms of harness development, agent skills, or whatever, you could also replicate (at least approximate). Whether complex harness or raw model is the answer remains to be seen, but in any case I'd say that more than 70% of it is the raw model intelligence.
The real concern are the AI labs, basically Anthropic and OpenAI, those are the ones who can really wipe out or take a big chunk of the industry. They do have the real moat, by means of training and access to the latest frontier models.
So now it's just the time to experiment. The hacker mindset was never about a particular set of skills regarding WAF bypasses or memorizing XSS payloads, but rather about figuring how things worked and working around them. We need to apply this mindset to this moment to reinvent ourselves, and take control of what the future will look like.
Trying out Codex Security on a vibe-coded, medium sized codebase. Could be interesting since any bug would have been introduced by codex itself in the first place.
This article has been making the rounds lately, and if you haven’t read it, I recommend you do. It’s a great case study on how AI can be leveraged to exponentially increase a researcher’s capabilities beyond what’s possible for a small team. Imagine how long this would have taken to do manually.
Another interesting reflection is on the need to carefully guide agents toward their objectives. This is an important example to keep in mind when discussing the importance of custom harnesses versus leaving everything to the model.
We have great examples of both: Anthropic's N-days research explicitly relied on a minimal harness, but I imagine that without careful guidance, tooling, and prompting, this Google research wouldn’t have been this successful. There might even be another $500k to make with a better harness.
In any case, congratulations Brutecat, and thank you for sharing this with the world!
Fable in Cowork seems to be capped on Claude Teams subscriptions. I don't see the same callout on Code. Hopefully it's just a limitation for Cowork and they still give us full access with a subscription.
Even though I haven't been able to test its cyber capabilities, it seems to be a step up in terms of coding so I'm already getting hooked.
The model card numbers are really impressive. Let's see how they hold up in practice.
The cyber security guardrails are very strict. Even when being approved for their cyber verification program, any attempt to work on something just mildly related to security research gets flagged. Hopefully they relax these soon.
The aesthetics of the announcement are impeccable though.
Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use.
Its capabilities exceed those of any model we’ve ever made generally available.
As far as we can tell, no. There is only anecdotal evidence, along with claims from AI pentesting vendors.
If a strong model can do everything by itself, then what exactly have these vendors been building? It is understandable that people would prefer a story in which the harness, workflow, and surrounding infra matter a great deal. It's also why people keep flexing "0-days" in OpenSSL, FFmpeg, or nginx, despite limited real-world impact.
That said, Niels Provos was not trying to sell anything, and he and several people have reported good results with IronCurtain despite using relatively weak models.
Most importantly, what Google achieved with Chrome suggests that a good harness may be quite valuable. Google does not appear to have access to anything more capable than Mythos, which means they likely scanned Chrome using Mythos itself or something less powerful. Yet they still uncovered hundreds of bugs.
There is, however, another explanation. Google may simply have better Chrome/V8 experts who can extract more value from Mythos. This remains our preferred hypothesis.
What provides a real advantage: domain knowledge accumulated over many years, or a harness vibe-coded in an afternoon? We think the answer is fairly obvious.
Frontier models are also really good at finding and exploiting n-day vulnerabilities, doing so on timescales of hours. Read about some recent work from my team studying these capabilities! https://t.co/668NzY2I2J
We just made it to @BSidesTampa Unlucky13, and it’s already been great to be here 🌴🏴☠️
Even before getting to the venue at @USouthFlorida, you could tell the team put a lot of care into the details; clear parking directions, check-in info, and everything you need to get settled.
The energy is awesome. Around 2,000 people are here, plus what looks like a huge group of volunteers helping keep things moving.
There are also several villages and tracks running in parallel, covering everything from offense and defense to CISO topics, IoT hacking, AI, and more. Plenty to learn, plenty of people to meet, and a really strong community vibe.
Looking forward to catching up with friends and meeting new folks 🤝
If you’re here and want to connect, shoot us a message.
Enjoy the conference!
The most common LLM security problems usually don’t start with advanced attacks.
They come from production systems that still have prototype assumptions: the model is trusted too much, the tools are scoped too loosely, and the guardrails live in prompts instead of code.
We put together a short post on 5 vulnerabilities that show up repeatedly in LLM deployments:
https://t.co/fEttNGJy0y
🚀 @HackSpaceCon 2026 mission accomplished!
We had a great time attending Hack Space Con this year. It was especially interesting to see so many conversations centered around the intersection of cybersecurity and space, along with the latest developments in AI.
With the recent Artemis II launch still fresh in everyone’s mind, the space and cybersecurity themes felt especially relevant.
We also really enjoyed connecting with security practitioners, researchers, engineers, and newcomers who are just as excited about this space/cyber crossover as we are. 🧑🚀
And as Space Coast locals, the conference gave us the perfect excuse to revisit the @ExploreSpaceKSC!
Big thanks to the organizers, speakers, sponsors, volunteers and everyone we met along the way.
Already looking forward to the next mission, and to @BSidesTampa next Saturday! 🏴☠️
We wrote a blog post about why it's a bad idea to use domains like https://t.co/8XT6Bvu4Ax and https://t.co/9ViFeudvD4 during testing and in security write-ups.
Even though it’s very common to use them in PoC examples and payloads, they may end up being executed by people validating findings, potentially leaking information to uncontrolled third parties.
To make things worse, these patterns have made their way into LLM training data, and agents are more than happy to reuse them during testing.
Read the full blog post:
https://t.co/KEbeSInILx
#pentesting #aiagents #appsec
AI agents are evolving, but manual pentesting isn’t going away anytime soon.
To cut out the repetitive friction in @Burp_Suite, we built 3 lightweight extensions to sharpen your workflow:
Tab Autonamer: Dynamic naming for Repeater/Intruder tabs.
Find in History: Instant jumps from Search to Proxy entries.
Forward OPTIONS: Auto-skip preflight noise.
Low on flash, high on efficiency. 🛠️
Get them on GitHub:👉 https://t.co/ZpZ5nb0OE4
Full breakdown:👉 https://t.co/UbCV5vtCaK
#BugBountyTips #Pentesting #BurpSuite #AppSec #RedTeaming
We’ve just launched our new website: https://t.co/I152J6fMaN
Built to better reflect how we approach security: no checklists, no noise—just real-world testing and meaningful results.
Have you ever needed to share a log file, JSON, or configuration file with a colleague, a vendor, or an AI assistant—but hesitated because of the sensitive data inside? 🛑
You need to keep the data structure intact for debugging, but you can't risk exposing PII, API keys, or credentials.
Today, we are releasing https://t.co/6PP9gDCUQb 🚀
It’s an open-source tool designed to sanitize your data before you share it anywhere. It supports formats like HTTP, JSON, XML, and YAML, replacing sensitive information with realistic placeholders.
Most importantly? It runs 100% in your browser. There is no backend processing, so your sensitive data never leaves your device.
Give it a try: https://t.co/JiG9EpPfOv
Check out the repo: https://t.co/uR2fqjKYvF
#infosec #privacy #devtools #opensource
What a week for cybersecurity in the Sunshine State!
We are at @BsidesORL at @FullSail, and the event is just getting started with an awesome crowd.
Looking forward to some great talks and meeting new people. Reach out if you want to connect today or at the upcoming events!
Next stops: @bsidessoflo and @BSidesStPete.
#BSides #InfoSec #Cybersecurity