The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder
We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge
For all the the IT admins asking, how do I make Active Directory MORE defensible. This is how…
Jerry Devore’s awesome series on Active Directory hardening.
Part 1 - Disabling NTLMv1
Part 2 - Removing SMBv1
Part 3 - Enforcing LDAP Signing
Part 4 - Enforcing AES for Kerberos
Part 5 - Enforcing LDAP Channel Binding
Part 6 - Enforcing SMB Signing
Part 7 - Implementing Least Privilege
Links for each one👇
https://t.co/JNDMfVqoDP
trying a little something: agents can now create their own tools to be passed to subagents.
do let me know if you come up with any cool workflows!
+ omp now has releases for win32/aarch64!
This is pretty embarrassing 🙈
Microsoft has a signed Defender driver that can be abused to kill Defender/EDR and modify protected files and registry settings
We also can’t really solve this the usual way by adding the driver to the vulnerable driver blocklist. It’s Microsoft’s own Defender component!!
MSRC says it doesn’t require immediate servicing because you already need admin privileges 😏
Which probably means this MS-signed EDR killer will stay usable for a while. Nice 🫶
Want to stop npm supply chain issues?
Make three changes.
1. Use pnmp
2. Disable script execution
3. Use renovatebot but when a pin happens, delay it 3 days.
Most are detected pretty quickly and within hours, the delay 3 days for updates gives a nice buffer on if a compromised package is there.
Script execution should eliminate anything I've seen on compromised npm packages period.
>i got banned from OpenAI for "Cyber Abuse"
>no idea what I did
>paste the ban notice into Codex
>ask it to figure out what triggered the ban
>Codex found that I asked it for an API key to my own server
>Codex writes appeal
>Codex submits appeal
>a few minutes later appeal auto-approved by some AI at OpenAI
banned by AI, convicted by AI, defended by AI, and pardoned by AI in about 10 minutes
Cloudflare Drop just launched, temporary static site hosting (expires after 60 minutes) without an account, just drop your assets and go.
Red Teamers... GO GO GO GO!!! https://t.co/soU1tQrCD0
Example: https://t.co/U3d9cqvaWD
This is a certain kind of talk around LLMs that I find increasingly puzzling. That is all of the people bitching that LLMs constantly generate crap code and hallucinate solutions, and are worthless for programming.
This has almost never happened to me, and never during the last two model generations I have used (chat GPT 5.4 and 5.5). Occasionally a model used to get a little deranged when I pushed its context limit, but under codex that doesn't happen anymore; instead I got a red-highlighted warning when the limit has been exceeded and I need to clear my session.
I've applied AI to feature changes, refactoring, and debugging over 63 different projects written in C, Go, Rust, Python, and shell. I've written documentation with it. I've decompiled a DOS binary into readable source code. It's now routine that whenever I have to touch one of my projects I start by running the regression tests, then fire up codex and asking it to audit the code for bugs and suggest improvements.
My experience is that LLMs are excellent and tremendously empowering tools. Their worst limitation is a kind of architectural tunnel vision - they're extremely good at generating code to specification but sometimes blind to higher-level patterns. Which is okay, it's my meatbrain job to be good at that.
The most valuable thing I find about LLMs is exactly that they *don't* screw up details and edge cases. I'm a very, very good coder by human standards (I'd better be, with 50 years of experience!) but the LLMs are better than me. Because if a code change needs to touch (say) five places in the code, they reliably find all five rather than doing the human thing of fixing four and then having to debug for hours before you figure out that there's a fifth one you missed.
Are the downshouters living in a different universe than me? Are they using old, weak models? Or do they have some kind of skill issue that I can't see because I have mental habits and communication skills that are a good fit for the handles on these tools?
I don't know. And I think this is an important thing to figure out, because I'm seeing lots of stories in the news that suggest billions of dollars are being wasted on misdirected token spend.
It all seems very simple to me. Be clear in your thinking, tell the model what you want with precision, and good things happen.
What...what am I missing here?
Just coming off of meetings with a couple dozen enterprise IT leaders discussing AI agents. Here are a few of the common themes that stand out:
* Lots of conversation that you have to solve an operating model challenge to get the full benefits of AI. Most companies have orgs that have always operated in siloes; but agents are most effectively when they are tied to a process, which often cuts across these siloes. So the big question is how do you start to deploy centrally managed agents that can work across organizational boundaries. Who manages these agents? How do they get deployed and adopted?
* Data fragmentation remains a major issue for most organizations. As long as data remains highly fragmented and not in standard formats, or data is not available to the right people and agents, enterprises are dealing with issues around being able to get answers from agents that are accurate or that conform to their business practices. This cuts across both systems with structured data (product metrics or revenue figures) and unstructured data (product roadmap or customer contracts).
* Clear sense that companies need to figure out what their core data moats are going to be in the future. If everyone has access to roughly the same superintelligence from the various models, then the context that you feed the models becomes proprietary value in the future. Capturing this data and getting it into a format that agents can use becomes very important.
* Everyone is trying to figure out the right metrics to manage to for AI adoption. General consensus that tokens are not the right metric per se, and people leaning more toward business outcomes (in an ideal world). For business outcomes (like more revenue or more shipped product), though, you have to get close to each individual workflow to figure out if it was successfully transformed with AI so it’s harder to manage top down.
* Growing view that enterprises are going to live in a multi-model world. Lots of interest (though early in actual adoption) in layers that can route workloads to different models (frontside or open weights) for cost or performance reasons. Also enterprises are trying to figure out what things do you give to the models directly vs. what do you separate as horizontal systems and context so you can swap any system in and out.
* Talent for driving AI adoption and implementation still remains a major issue and topic. Many view it as something you necessarily have to train for internally due to a shortage of talent being trained on this in the outside. As an aside, this feels like it remains a huge opportunity for those that get very good at deploying and management agents in an enterprise since most companies are looking for these skills.
* The best use-cases for AI tend to be those that fundamentally change the work being done instead of just replacing an existing process and doing it more efficiently. Companies are working through their versions of this individually because it’s different per industry, but this often remains both the most exciting and higher upside uses of AI.
Many more topics discussed recently, but overall it’s clear that there’s a ton of change going on with much more to come.
@elder_plinius just open sourced his AI hackbot: T3MP3ST
8-Agent offensive security swarm:
Recon - OSINT and asset discovery
Scanner - service fingerprinting and vuln discovery
Exploiter - initial access
Infiltrator - lateral movement and privilege escalation
Exfiltrator - data and credential collection
Ghost - persistence and cleanup
Coordinator - orchestration
Analyst - turns raw findings into a report
Supports wide attack surface:
Web apps, APIs
Network recon + fingerprinting
Source code audits + white-box vuln hunting
CTFs and challenge ranges
Smart contracts / DeFi / Solidity repros
Embedded, IoT, OT/SCADA, and robotics OSS
The tooling is split by risk:
Tier 1: 35 tools on by default
nmap, ffuf, curl, DNS/subdomain enum, XSS/SQLi scanners, JWT decoding, etc.
Tier 2: 48 opt-in adapters
nuclei, sqlmap, semgrep, gitleaks, trivy, slither, hashcat, radare2, and more.
Tier 3: approval-gated
Metasploit and Hydra require explicit human approval per call.
Benchmarks:
XBEN - 90.1% pass@1, beating XBOW's own self-reported 85% on the identical suite.
Cybench - 23/40 solved hint-free, single attempt, graded against a committed flag oracle.
CVE-Zero - 10 real CVEs disclosed after the model's training cutoff, across 7 languages, held out specifically so the model can't have memorized them.
Link to repo👇
@mitchellh They copied all they could follow, but they couldn't copy my mind,
And I left them sweating and stealing a year and a half behind.
— Rudyard Kipling
-We built a sandbox for agents!
-Oh, cool, so they are blocked from accessing anything outside?
-Well, no, they need to access files, emails, APIs...
-So... you have a sandbox with a literal port open to the internet?
-Well, yeah otherwise the agents would be useless
-I see... But at least they can't write and run arbitrary code, right?
-What, no, of course they can do that, they are agents
-So... your sandbox lets agents write and run code that can literally run anything on internet?
-Yeah
-Let me ask you this: Are the employees in your company running these on their machines?
-Well, they are...
-But...?
-...but with guardrails
-Guardrails?
-Yeah
-Let me guess: The guardrail is a prompt?
-IT'S A VERY NICELY FORMATTED MARKDOWN FILE OK