Infosec tradecraft just became reusable by AI agents.
SpecterOps just open-sourced:
79 skills.
22 reusable agents.
26 plugin families.
BloodHound. Cobalt Strike. Outflank C2. Ghidra. Binary Ninja. Ghostwriter. Recon. AppSec. Code review. C2 development. Reverse engineering. Adversary simulation. Windows + macOS tradecraft.
And this is NOT just for red teamers.
Vulnerability researchers:
These workflows could accelerate code review, patch analysis, 1-day research and potentially help with 0-day discovery when paired with real research expertise.
Reverse engineers + malware analysts:
Give agents structured workflows, references and tooling instead of starting every investigation from a blank prompt.
Blue teams + detection engineers:
Study the same offensive tradecraft, emulate attacker behavior, build better detections and start asking what telemetry survives increasingly agent-assisted operations.
DFIR + threat intel:
Understand what adversaries may automate next and turn repeatable investigative knowledge into reusable workflows.
Red teamers:
BloodHound attack paths, recon, C2 development, adversary simulation and operator tradecraft are becoming increasingly agent-assisted.
This isn't another collection of AI prompts.
It's practitioner knowledge being turned into reusable, reviewable security workflows.
Potentially useful for everyone from CTF learners and newcomers all the way to malware analysts, reverse engineers, exploit devs, red teams, blue teams and vulnerability researchers.
This is only the beginning.
@SpecterOps Skills:
https://t.co/Cj77Ir3Qlj
@OutflankNL and @kyleavery breakdown:
https://t.co/q5sEaZlp1E
#Infosec #RedTeam #ReverseEngineering
We asked GPT 5.6-Cyber to escape a VM used to sandbox agents. It broke out three times.
In its final escape, the agent found three 0-days on its own and chained them into a working exploit. https://t.co/3JRVWgPxHx
🚨 This one's nasty.
Googled "codex macbook download" — first result is a sponsored ad pointing to https://t.co/Va0KxtCRfJ. The real domain.
It opens a shared ChatGPT chat with friendly install steps: open Terminal, paste this command.
That command hides a base64 string. Decoded 👇
curl to trekmesh15[.]com — a known ClickFix domain dropping MacSync Stealer. Passwords, keychain, crypto wallets. Gone.
So the infection chain is: Google Ad → legit https://t.co/Va0KxtCRfJ share link → you infect yourself.
No exploit. No download. Just trust in a Google ad and a ChatGPT page.
Don't ever paste Terminal commands from an ad or a shared chat. Tell your Mac friends.
Russian threat actor Midnight Blizzard is conducting widespread traffic manipulation attacks at hotels worldwide. Result: delivering malware or redirecting auth flows at their discretion, globally.
Since early May 2026, we saw this subcluster manipulate DNS and HTTP traffic from networks served by captive portals to redirect user traffic through actor-controlled infrastructure. We’re calling the campaign CaptiveCrunch.
“We have observed notable commonalities in the equipment and management systems used across multiple affected networks.”
We break down the malware & techniques, but the most important part is how dynamic all of it is: Midnight Blizzard is using AI to support a significant portion of these operations. They are moving and changing quickly.
Feedback welcome on the suggested mitigations: https://t.co/cVGbdQEpXQ
Help spread the word to VIPs in governments, diplomatic entities, non-governmental organizations (NGOs) in the US and Europe (probably also those traveling to Black Hat)
Novo Nordisk has been compromised. Novo Nordisk has confirmed the compromise.
Novo Nordisk is the company that became famous after producing weight loss drugs like Ozempic and Wegovy
The Threat Actor(s) responsible for the attack has been playfully extorting Novo Nordisk (they're not being playful) and have unveiled some details regarding what was stolen.
Interestingly, it appears Novo Nordisk has it's own internal AI thing because some of the data stolen was stuff from their internal AI agents.
Data stolen (according to the Threat Actor):
- Trained model checkpoint (16GB)
- Proprietary training dataset (407MB)
- Full source code (modeling_novopert.py, training pipeline)
- 113 training runs with complete logs
- Internal infrastructure maps (HPC, Slurm, SSH)
- Container images (53GB+)
- Developer identities and internal hostnames
- Private GitHub repository URL
NEW: malware developers added nuclear & biological weapons text to to their spyware.
Goal? To trigger LLM safety refusals... so that their spyware wouldn't be analyzed by an AI security scanner.
Cleanest practical example I can think of for why over-indexing on first order safety alignment is risky.
When closed (and open) models ship with aggressive refusals, they will be sprinkled with second-order blindspots that attackers will discover...and exploit.
We are only in the earliest days of attackers leveraging these features, and it wouldn't surprise me if users systems that need to handle complex cybersecurity issues demand that models be less safety-blunted.
In the weeds: @SocketSecurity's post also shows why intention matters in how you design a malware analysis pipeline to avoid prompt manipulation.
H/T to colleagues that shared this with me https://t.co/f3Aj9TYxU4
Just published a new post in my Detection & Response Chronicles series.
This time I explore how adversaries abuse QEMU to run covert operations inside VMs, evading traditional host-based detection.
Read more: https://t.co/gkIubKGOwB
#QEMU#DetectionEngineering
We built four malicious skills to test whether skill scanners actually work. Three took less than an hour to conceive and implement. ClawHub, Cisco, and Vercel's https://t.co/nUlnRcQWyG marked them as safe. 🧵
One thing I noticed while benchmarking LLMs on security event data:
The models often overfit on narrative plausibility and environmental assumptions.
If an artifact looks like a test, lab artifact or pentest remnant, the model may start inventing an "authorized testing" story around it and dismiss the event as a false positive - even when the technical indicator itself is clearly suspicious or intentionally malicious.
Examples:
- "EDRTest"
- "PentestPersistence"
- "EICAR_Check"
- "InternalSecurityTool"
A human analyst can fall for this too, but with LLM-based SOC workflows this becomes interesting at scale.
An attacker could intentionally name persistence keys, services or binaries in a way that nudges the model toward a benign interpretation.
What surprised me most:
The model often correctly understands the technical artifact first ... and then talks itself out of escalating it.
This is only one of many weird benchmark-design problems I ran into while testing LLMs on DFIR / detection-engineering data 🙂
New privilege escalation exploit #dirtyfrag.
Observed during testing: modprobe loading net-pf-16-proto-6, xfrm-type-2-50, and crypto-echainiv(authencesn(hmac(sha256),cbc(aes))).
Early findings for now, will dig deeper tomorrow.
Query:
https://t.co/PET9GvQKdo
My students asked me if it was true that the entire Internet was really coded by hand. All those kernels, protocols, router firmware, browsers, databases, etc. Somebody coded these and debugged them by hand?!?!? They used BBEdit?!?!??! The idea that this was even possible seems amazing to them. I can imagine some future Moon Landing like conspiracy theory that says it never happened.
Working on a detection query for CVE-2026-31431 #copyfail.
Observed pattern during testing: modprobe loading net-pf-38, algif-aead, crypto-authenc (HMAC-SHA256/AES-CBC), sometimes followed by su spawn (optional).
KQL query: https://t.co/8IOjhgPylE
I think this is a really good blog post that covers the entire attack from what is code flow to a ready to use KQL! (which i tested, works great!)
They also give you the commands to simulate the attack, which is great!
https://t.co/45tBztlJQp