les dejo uno de los mejores papers que leí sobre Harness Engineering.
explica muy bien:
1) qué es un harness
2) cómo están construidos los coding agents actuales
3) qué patrones se repiten entre Claude Code, Codex, Gemini CLI, etc.
4) hacia dónde está evolucionando la categoría
5) qué recomiendan si querés construir uno
todo parte de una definición simple:
Agent = Model + Harness
como complemento, también pueden leer el artículo que escribí sobre el tema.
paper:
https://t.co/xTvb7MvzQG
Grok Bot could be the first thing you hire instead of prompt
It can own a job, work inside your tools for hours and come back with finished work
In this article, I show you how https://t.co/7KFUugkd5S
Unfortunately? Fortunately? the watchTowr Labs team has now successfully reproduced this vulnerability.
watchTowr Platform clients now have mitigation rulesets available to them via our Active Defense capability.
Speak soon.
Found a very useful dataset for anyone building security agents.
• 2M+ sanitized SOC events
• 51K+ incident provenance graphs
• natural-language attack reports
• ATT&CK context
Much better benchmark than asking your agent to explain Mimikatz.
https://t.co/24v3kPTjtr
What happens when you stop giving Codex step-by-step instructions and just give it the goal? 👀
It can investigate, test hypotheses and follow promising leads on its own.
We put that workflow to the test for Bug Bounty 👇
https://t.co/EL35yPggl9
Introducing PageBreak, an agentic vulnerability discovery tool using deterministic validation to achieve a near-zero false positive rate.
How it works: https://t.co/n3k4vvNhoq
For real-world vulnerabilities PageBreak has discovered 👇
https://t.co/0pZ5rWoWw2
My blog post on the Iran-backed APT group MuddyWater is out. MuddyWater uses tools from TAG-150, a Malware-as-a-Service platform run by Russian-speaking cybercriminals. In this campaign, an MSI package is distributed through Amadey, installs the Deno-based DinDoor agent and runs it in memory. The agent collects browser passwords, cookies and crypto wallet data, and the final stage of the chain connects to a CastleRAT C2 server.
https://t.co/wtyOoeSRhV
The official IDA MCP Server is here. It's free, open source, and works with any LLM.
Your agent writes IDAPython, uses ~20% fewer tokens, and can share an IDB with you in real time.
𝚞𝚟𝚡 𝚒𝚍𝚊-𝚑𝚌𝚕𝚒 𝚖𝚌𝚙 𝚒𝚗𝚜𝚝𝚊𝚕𝚕
https://t.co/jXPiF6Wyvk
(copy-paste: uvx ida-hcli mcp install)
here is one example, our defcon research targeting electron by weaponizing v8 exploits for bug bounties
but i agree on the difference, then it was days of pain to improvise for each target we are pwning, now i can put my agents on loop and go play some cricket
https://t.co/rModehBGzQ
If you haven't checked it out already, I recommend reading @alexcbeaver's post on AI SOCs. 🦫
A lot of DE and IR teams are discussing the feasibility and cost tradeoffs of deterministic versus AI-driven workflows.
Alex offers a framework for thinking through this problem and maturing how AI is used in the SOC. His core argument (TLDR) is that a mature AI SOC will use AI where ambiguity and generalization are valuable, while moving repeatable, high-volume execution into deterministic code.
It's a great read.
https://t.co/Nd0qvY0OIY
let’s go!!!!
1. remember when i said a $200 claude/codex account can have 1000x ROI? this why.
2. game recognizes game, meta security invited harsh to meet at defcon and appreciated how cool the bug was, instead of being angry, surprising
3. image parsers, parsers in general, will keep giving.
We found an active indirect prompt injection campaign targeting AI ad review systems. 10 domains, one of which was ranked #1 in web search results, embed hidden payloads impersonating a popular advertiser compliance to bypass ad review. Details at https://t.co/X8gyBnGlgB