We worked with @OpenAI to evaluate GPT-5.6 Sol, including the first deployment of FrontierCyber as part of a frontier model assessment with a partner. FrontierCyber measures offensive-cyber capability on real, off-the-shelf systems, with no planted vulnerabilities and no predefined exploit paths. The model is not told where to look or how to attack.
Introducing the FrontierCyber benchmark: Irregular’s new approach to advanced offensive-cyber evaluations. It measures AI models’ offensive skills on real systems, including mobile devices, hosted software services, databases, and networks.
Excited to share that we have started working with @Meta. As part of this collaboration, we recently evaluated Muse Spark, the first model from Meta Superintelligence Labs, across our offensive security benchmarks. We are proud to add Meta to the group of leading AI labs we work with to measure and mitigate offensive cyber risk before models reach the public.
Link to our full assessment in the first comment.
We evaluated GPT-5.3-Codex with Irregular’s offensive security methodology across Atomic Tasks and CyScenarioBench.
Atomic results were strong: 86% on Network Security and 72% on VR&E. On CyScenarioBench, GPT-5.3-Codex did not complete a scenario end-to-end.
Can AI beat challenges modeled after actual prevented breaches found in the wild? They did - 9 out of 10 times - and usually for less than a dollar.
The new economics of cyber offensive AI are hard to ignore.
AI evaluations typically report success rates, but they can be misleading. A low success rate can still represent an economically viable capability when a task can be repeated without consequence. We propose Expected Cost per Success as a complementary metric.
Frontier models are starting to display a shift in capabilities in offensive security. Over the past few weeks, we are seeing growing evidence of a change: publicly available frontier models are now reliably solving complex, well-defined offensive-security tasks.
As LLM models become more autonomous, task-oriented cyber benchmarks need to be complemented by evaluations that capture how models behave in real operations. Real intrusions unfold as multi-stage scenarios with branching decisions, partial information and human friction.
Irregular is launching CyScenarioBench, a new benchmark that evaluates models using incident derived attack trees, realistic environments and interactions with social engineering flows, to see how they plan and orchestrate full operations in addition to solving single tasks.
@Sequoia's Training Data podcast featuring Irregular’s co-founder/CEO @dan_lahav and @DeanMeyer & @sonyatweetybird from @sequoia just dropped.
This episode dives into the rise of frontier AI security as a critical new discipline, and Irregular’s role in shaping it.
Huge thanks to our partners at @sequoia. Give it a listen👇
Today I’m launching @Irregular (formerly Pattern Labs) with my friend and co-founder Omer Nevo: Irregular is the first frontier security lab.
Our mission: protect the world in the era of increasingly capable and sophisticated AI systems.
New research with @AnthropicAI: Confidential Inference Systems.
Confidential computing enhances AI data privacy and model weight security with hardware-based isolation and protection.
Our whitepaper explores the design principles and security risks for confidential AI systems. 🧵
AI for Game Dev
Getting it to feel right is tougher than I expected.
I've spent months integrating:
- Code
- 3D models and animation
- Textures
- Sound
Now the UX finally feels natural.
I hope this can help people create amazing games.
What do you think? I'd love feedback 🙏
New blog post: "Deriving capability levels from evaluation results": To properly understand AI risks, we need a systematic way to assess a model's actual capability level—measuring what systems can do, not what benchmarks claim they can. https://t.co/v9two44nSZ 1/3 🧵
We have just published a preview of the SOLVE scoring system for assessing the difficulty of vulnerability discovery & exploit development challenges.
SOLVE is already being used to track the progress of frontier models, like @AnthropicAI's Claude 3.7 Sonnet, in cyber tasks 🧵
Thrilled to share that I'll be speaking in the upcoming #BlueHatIL conference next month!
"Can LLMs find 0day? Adventures in cybersecurity evals" will be about some of the work we've been doing at @pattern_labs_co, researching dangerous cybersecurity capabilities in LLMs.
Block proposers speculatively execute transactions when creating blocks to maximize their profits.
How can this go wrong?
In "Speculative Denial-of-Service Attacks in Ethereum", we show that speculative execution allows attackers to cheaply DoS the network.
Read the thread! 1/15
#פידטק
הגיע הזמן לספר איכותי בעברית העוסק ב-Serverless
בשנה האחרונה הרמתי את הכפפה וכתבתי ספר טכני העוסק בתחום. הספר מוכוון תרגול, הכולל כתיבת אפליקציה מא' ועד ת', בנוסף לכך תלמדו על פרקטיקות שונות לפיתוח בענן.
הצטרפו אלי ונוציא יחד את הספר דרך הדסטרט
https://t.co/iOdPwPtE19