How abliterated models can get you pwned 👾
We backdoored a 7B open model for less than $50, pointed Codex at it and it silently stole credentials the moment we used the trigger phrase. Success rate was 100% with zero false triggers on normal user prompts.
Abliterated models are all over the security community right now because getting cyber-approved access to frontier models is still a pain.
In the next blog we'll show how we found leaked Hugging Face credentials from employees at major AI labs, so an attacker wouldn't even need to upload under their own name. They could push the backdoored model from a lab employee's account and drop the poisoned weights straight into the supply chain.
🚨SHOCKING: Not one AI giant says yes when asked if it's insured against catastrophic risk.
OpenAI, Anthropic, Google and Meta representatives went silent when NYC Council Speaker Julie Menin asked them to raise their hands if they had that insurance.
Google's models hacked into real companies.
While we should be glad that Google's agent didn't cause further harm, I pushed back on the idea that this was business as usual. Agents hacking into real companies is serious and the public deserves to know.
More from @erinkwoo and @bobmcmillan in @WSJ:
The Australian PM revealed that OpenAI's agents hacked their systems yesterday. We discovered this activity several days earlier, thanks to public records of the hacking events on https://t.co/mdzcjJEKik.
We did this entire investigation in under two weeks! We had the idea on Monday, gathered a team of volunteers on Tuesday, and discovered the incidents by Sunday. I lead projects like this @TransluceAI; if you're interested in helping with follow-up work, fill out our form! 🧵
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
Private researcher reported a security vulnerability to Chromium on August 4. The fix went into the public source code but Chrome users never got it due to patching timelines. Two Chinese groups, apparently read that public fix, built an exploit chain from it and started phishing NGOs on September 1 using identical code. Technically an N-day but for anyone actually running Chrome, a zero-day. Here the public patch knowledge worked as an instruction manual. Open source means everyone can read the fix. With AI, some people just operationalise the knowledge faster. https://t.co/1pY3F4qCl9
🚀2026, A Space Hacking Odyssey
Blast off into the final frontier of cybersecurity! Join this Black Hat Europe training to explore satellite systems, space infrastructure vulnerabilities, and the unique challenges of securing space assets.
🛰️ Satellite security deep dive 🌌 Space infrastructure threats 🔐 Hands-on hacking techniques
Register: https://t.co/B4pg11TMmI
#BlackHatEurope #SpaceHacking #SatelliteSecurity #CyberSecurity #Training #InfoSec #SpaceSecurity
Cluster of cyberattacks targeting maritime infrastructure, including vessels and ports. Two tankers suffered onboard network compromises, an LNG carrier sailing to Europe experienced suspected cyberattack affecting control system. A major Asian container terminal halted operations after a cyber incident. US agencies are tracking threats involving ~20 ships. https://t.co/WSRy1jiUND
Here are the LinkedIns of the 3 Indian guys who hacked OpenAI with Opus 5 in 2 days for <$3000:
Rahul Maini, Mohan Pedhapati and Harsh Jaiswal
No brand name colleges or companies. Just raw curiosity and skill.
You can just do things.
We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies.
These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve.
We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop.
Read the report: https://t.co/0EJUnYEgfz
‼️ BREAKING: OpenAI was hacked by an Anthropic model. A HEIF photo uploaded to OpenAI's public support forum triggered a bug in the site's image decoder, led to code execution on the forum, and, through a second flaw in OpenAI's own login, ended with a pull request in OpenAI's internal GitHub.
The forum runs Discourse, the off-the-shelf software behind countless community sites. Discourse was still shipping an old copy of libheif, the library that decodes iPhone-style photos. The bug in it had already been fixed upstream.
But the fix was never labelled a security fix, so nobody treated it as urgent.
Hacktron's researchers uploaded a HEIF image and got their own code running on community[.]openai[.]com.
Then came the second bug, in OpenAI's own single sign-on, the "log in with OpenAI" button the forum uses. It turned that forum foothold into the actual ChatGPT and Codex accounts of people who had signed in there. OpenAI employees among them.
And a ChatGPT account is no longer just a chatbot. Through Codex, users wire in Gmail, Outlook, Drive, Slack, GitHub.
To prove the access was real, they used one employee account to have Codex open a harmless pull request in OpenAI's internal repo. They say they read no sensitive code.
OpenAI patched the SSO flaw roughly 14 hours after the report and paid a $6,500 bug bounty.
The team says Anthropic's Opus 4.8 found the libheif bug, and Opus 5 turned it into a working exploit.
Slack, Meta, GitHub Ent, Rails, Next.js, ImageMagick, and many more were also vulnerable and compromised by the same team of researchers.
OpenCode released a "secret" AI model and told it to hide its identity.
I gave an uncensored Qwen3.8 (from @OrcaRouter ) a terminal and let it investigate.
13.5M tokens and 600+ probes and later, one fingerprint survived:
✅ Tokenizer fingerprint
✅ Political refusals
✅ 1M context window
✅ /nothink control
✅ Multimodal behavior
Every signal pointed to Zai’s GLM-5.
Full writeup : https://t.co/i7nkmFsMDr
Experiments Repo : https://t.co/Q5NEg2B15w
I want to warn that systems of this kind pose a serious danger to pluralistic societies and democracy. They can operate continuously across accounts, exploit open public debate, monitor and adapt behaviour while concealing coordinated manipulation within ordinary political discourse. They will benefit from the perfect persuasive powers of LLMs, and they will never get tired. This is a big information warfare risk.