OpenAI is focused on “AI safety” while falling very short on AI security.
This incident is a *security* problem that requires *security* solutions. And if there’s anything to learn from it, we need the audit log.
An Anatomy of the ExploitGym Incident 🧵: When an OpenAI model hacked its own benchmark
1/
Sources: @OpenAI , @huggingface official publications and ExploitGym resources: paper, Github huggingface
@_nexact Yah for sure - we use Cribl heavily at Bd and partnered with the folks there - usually forking data retention to reduce logs in siems for folks but it’s a great product for sure
One amazing thing I focused heavily on was fixing traditional log ingestion and detection engineering in NightBeacon. We use a AI native universal log normalizer / parser that doesn’t worry about source technology.
Then we use agents inspecting every submission of source. If source is unknown it kicks off autonomous research on that technology and builds training data for it automatically.
Pretty incredible seeing this work across customers and completely eliminate the complexity around parsing data and directly applying static rulesets to detections.
Game changer.
#NightBeacon 🥓
ChatGPT Voice is now in the desktop app.
Control your computer and direct multiple agents running in ChatGPT Work or Codex, using just your voice.
It's powered by GPT-Live, so it can speak, listen, and coordinate work in the app at the same time.
Rolling out globally today on macOS and Windows to Plus, Pro, Business, Edu, and Enterprise plans.
Had some spare token usage left before reset, so I had codex play my character on my MUD for me which is a built in full RPG into NightBeacon because we are so good at catching attackers, why not MUD with other analysts and customers during downtime... so I need to make sure I'm beating my dev team and automatically leveling me up.
In regards to the OpenAI hack, as more data and information comes out. A couple of thoughts.
One - OpenAI's sandbox environment was not setup or designed well for it to escape the way it did. Devils advocate here, you want it to have as real/live of an environment for appropriate testing on capabilities, so this is a delicate balance.
It's clear that the detection capabilities were not solid either at least for this enviornment. This would have set off every alarm in the detection catalog.
I will say, HuggingFace identifying it within a couple of days is impressive still, granted it could always be better/faster and followed solid response procedures here to minimize the impact during this.
Third, the AI behaved exactly the way its supposed to. It's designed to reduce token cost utilization and finding the easiest way to objective, not the hardest. They are lazy by design. It found it's method to objective through cheating without proper guard rails.
Good lesson learned here, our traditional defensive capabilities still work as designed, should have snagged this, and the question is how do we leverage AI to better define our security controls, defensive posture, and more.
TLDR: Not groundbreaking or earth shattering, still super cool to see these capabilities expanding on the offensive and defensive side.
Also, what an amazing marketing opportunity for OpenAI right now 😂
California updated #CCPA... again. Before you do anything else, does it apply to you? In Part 1 of our latest blog series, Chris Camejo clarifies who falls under scope, what data is covered, and how enforcement works. Read it now! https://t.co/VJoQuSG4GC