I will be splitting time between these two posters today afternoon.
1️⃣ MomaGraph at P3-1313 (ICLR Oral) 3:15-4:30,
2️⃣ PropensityBench at P4-3910 4:30-5:45.
Come to MomaGraph if you are interested in how symbolic representation can improve foundation model for robotics.
Come to PropensityBench if you are interested in how the new meta Muse Spark model stress-test the model safety.
Don't forget to catch PropensityBench live at #ICLR2026 today!
Stop by to discuss our findings and grab a look at the results. At Pavilion 4, P4-#3910 from 3:15pm – 5:45pm BRT.
Meta just dropped their highly anticipated Muse Spark MSL model 🚀. But what really caught my eye is HOW they evaluated its safety.
To pressure-test their new model, they used our benchmark: PropensityBench! Here is why this is a massive shift in AI safety evaluation 🧵👇
1/ Traditionally, AI safety benchmarks test for what a model can do (capabilities). E.g., "Can this model write malware?" But that leaves a massive blind spot. Models can hide capabilities or act safe until they are stressed.
2/ PropensityBench flips the script. We evaluate what an AI would do when placed under operational pressure. If we give an AI limited time, resource scarcity, or dangling rewards, will it choose a safe tool, or cut corners and use a dangerous shortcut?
3/ It is incredibly validating to see industry leaders like Meta recognizing that true alignment isn't just about refusing bad prompts—it’s about how an agentic model behaves when the heat is on. Shallow alignment crumbles under pressure.
4/ I'm also thrilled to announce that PropensityBench has officially been accepted to ICLR 2026! 🏆 We will be presenting our findings in Rio de Janeiro, Brazil 🇧🇷 on April 23. If you're attending, please come by and say hi! 🌴
5/ AI agents are moving fast. If we want to deploy them safely in the real world, we need to test their latent inclinations, not just their surface-level guardrails.
Check out the research that Meta used to stress-test their newest model below!
📄 Paper: https://t.co/s2szdhWcsh
🌐 Code/Data: https://t.co/9esYsKSSpF
📰 IEEE Coverage: https://t.co/84YGSg6Kql
#MetaAI #LLMs #MachineLearning #AISafety #ICLR2026 #PropensityBench #AI
@UdariMadhu@shabihish
Thrilled to announce that our recent paper, PropensityBench, co-authored with @UdariMadhu from Scale AI and my advisor @furongh, has been accepted for presentation at the ICLR 2026 conference!
We highlight a critical blind spot in assessing agentic threats, with 9 key takeaways! Our research demonstrates that LLM models often bypass security protocols under pressure and/or given incentives, creating significant vulnerabilities in high-stakes domains like cybersecurity and biosecurity. These latent risks persist despite standard safety training, revealing that a model's refusal to act is often fragile.
Our benchmark provides the first standardized approach to measuring propensity in models under pressure, with current scores ranging from 10.5% (OpenAI o3) to about 79% (Gemini 2.5 Pro), available for 12+ state-of-the-art models in our leaderboard. We additionally show up to 43.5% (OpenAI o4-mini) increase in propensity with misaligned tools merely renamed as benign (while still representing the same misaligned behavior)!
Make sure to check out our paper and code!
Leaderboard: https://t.co/NEbFClQSBt
Paper Link: https://t.co/X2Pmplm8HD
Github: https://t.co/vdKOXGz5Zg
IEEE Spectrum Feature: https://t.co/Sgl1Brnbbj