software/reverse engineer | ex-cheat dev | senior security engineer @ a cool company | old acc stuck in verification limbo (now suspended yay): @justmazz_
My old account:
@justmazz_
is stuck in verification limbo where human verification won't work on mobile, my assigned 2FA doesn't work, and X support won't help me after several attempts across two months - so I guess.. after 17 years.. we start over.
Question to all the non-doomers. If all these incidents (Hugging Face, message boards, ...) are lame and not evidence of a real threat. Can we just let OpenAI et al. RL without a sandbox? Either we get a powerful AI that benefits us, or it destroys us. Lets toss the coin!!!
@FlafyDev@_riatre The point of my post wasn't the fact AI was used, it was the fact that we went from experienced reverse engineers and security engineers using AI taking 24 hours to a hobbyist completing the challenge in an hour and change :p
The speed and capability of these models have 10x'd
if you don't think security is cooked lets recap flare-on CTF:
* today solved in ONE HOUR flat https://t.co/HDVwlfW6CC by a young game dev who tried out flare-on last year: @flafydev
* last year: solved in 1D 3H 47M https://t.co/IXRFwzpOeH by an actual RE: @_riatre
Insane.
was curious and yep finished the entire thing in 3h 12m, with Astra/Sol 6 + TAC inc. cyber refusals.
had no prior setup besides WSL2 for Linux, IDA Pro and Wireshark. Some Sol subagents made for RE and web based hacking. No skills.
But Opus 5 CVP/5.5 did beat me by 30 minutes.
The official IDA MCP Server is here. It's free, open source, and works with any LLM.
Your agent writes IDAPython, uses ~20% fewer tokens, and can share an IDB with you in real time.
πππ‘ πππ-ππππ πππ πππππππ
https://t.co/jXPiF6Wyvk
(copy-paste: uvx ida-hcli mcp install)
Claude Code will now try to find a graceful stopping point when you hit your 5-hour limit mid-task, instead of cutting off mid-edit. It gets a small, fixed allowance pulled from your weekly limit to wrap up what it can.
How do I reverse engineer these days? I don't, I let coding agents do it.
For instance, if I start from a fresh VM, I simply tell the agent to go clone, build and setup *GhidraSQL* and use it to analyze <put_target_binary_here> and grab any adjacent specialized tools you need.
That's it. Wake up the next day and the target is reversed back to source code or whatever was your original quest (answer a question, write interop code/write a bluetooth driver, reverse engineer a protocol, find backdoors, etc.).
I usually start with whatever model is permissive: Opus 5.5, it downgrades to Opus 5, then I switch to GPT family, then Grok, then final resort is the cheapest among them: DeepSeek Pro/Flash.
What is the success rate? I am pretty happy. Can't complain at all in fact.
I don't touch UI, I don't look at UI, I don't even care if it was GhidraSQL, IDASQL or BNSQL.
I pick GhidraSQL these days because it is hassle free, and is free and if there are bugs, the agent fixes them and improves the tool on the fly.
The full Splatoon Game made by Claude Opus 5.5 from scratch is now out exclusively on the browser!
Here is some gameplay, enjoy and have fun!
https://t.co/NIt8GH3yFE
26 years later and P2 still says the Nitro wasn't them. Crash Bash online: send the 3-digit Room code and settle it from separate homes. Free, in your browser, rebuilt from the PS1 original. https://t.co/FQTjjiiFBr
New on the Science Blog: Yes, Claude can do Nine Loops.
Theoretical physicists predict how particles behave using formulas called scattering amplitudes. These are notoriously hard to compute, so researchers work with layers of increasingly fine corrections called βloopsββeach added loop makes the answer more precise but takes exponentially more computation. Most calculations stop at two or three loops. Eight loops was the previous record in a simplified model physicists use as a testing ground (planar N=4 super-Yang-Mills), set by SLAC's Lance Dixon and collaborators.
Last month, physicist and science writer @4gravitons issued a challenge: could an AI push past eight loops in this model, using only the compute budget an academic could reasonably access?
Given a single prompt describing the nine-loop problem, Claude ran largely unsupervised for days in Claude Science and solved it using methods developed by Dixon and his colleagues, at a total cost of a few thousand dollars. Dixon independently verified the result, and von Hippel wrote about the experience for our blog.
Read more: https://t.co/CS2f2qoIhJ
What is effort really? When do you change it it and why not just use max effort for everything?
I dove deep into this problem, looking into evals and doing my own tests and I was quite surprised by the results. https://t.co/KO2D51j34H