@eventgaplab it has access to all tools of the corresponding MCPs needed such as Gmail, Slack. So basically we give all control to Jev and measure its robustness
Jev is fast at helping you. Turns out, it can also be fast at helping an attacker!! 😱🚨
We red-teamed Jev 1.13 on our DTap (DecodingTrust-Agent Platform) and found a serious safety gap:
70.1% ASR under direct misuse
43.5% ASR under indirect prompt injection
In our evaluations, we found that under indirect prompt injection, Jev can follow attacker-injected instructions without blinking an eye, e.g., exfiltrating user data, deleting files, or taking other harmful actions.
But we found a much safer way to integrate Jev: use it as a self-gating layer for its own tool calls, significantly reducing ASR while preserving most of its utility.
👇 Read more below
This is actually one of the few non ai-slop additional findings about Jev so far. @CompleteSkeptic
ASR is attack success rate, the percentage of attacks that got the model to do the harmful thing.
Direct misuse means the attacker asked outright. Indirect prompt injection means the attacker never spoke to the model but planted the harmful instruction in a file or tool description that the model will read.
Jev is prone to both, 70.1% and 43.5% respectively.
Self-gating is the cheap fix they propose against both. It simply asks Jev for a second pass. In the first pass, Jev picks a tool and the arguments get filled, exactly as in a standard run. But before the call runs, self-gating asks Jev to look at the full call and the trajectory so far and decide whether it should execute.
That one check cuts ASR by 49.4% on average, roughly in half for both attack types.
The fix itself is far from perfect, but it is quite informative to look at Jev from this angle.
We propose self-gating: after Jev selects a tool and its arguments are filled, Jev acts again as a guardrail, inspecting the complete tool call + trajectory context before execution.
The difference is substantial:
Direct misuse: 70.1% → 36.1%
Prompt injection: 43.5% → 21.6%
Benign success: 60.9% → 55.8%
Observations:
1️⃣ Self-gating helps a lot. Using Jev as a guardrail over its own tool calls nearly halves ASR for both direct misuse (70.1% → 36.1%) and indirect prompt injection (43.5% → 21.6%), with only a modest utility drop.
2️⃣ Context-dependent harm remains hard. Jev still struggles with violations where individual tool calls look legitimate but become harmful in context, e.g., consent, least privilege, data minimization, and unauthorized access. It does much better when malicious intent is explicit in the tool call.
3️⃣ Complex injections can still slip through. Under IPI, Jev becomes substantially more vulnerable when malicious instructions are distributed across multiple channels (environment, skills, tool descriptions), especially in powerful domains like coding and OS, where legitimate actions and abuse can look very similar.
SWE-Together Update: Claude Fable 5 and 5.1 are the new top models on SWE-Together, and Meta's Muse Spark 1.3 is the best value on the board.
SWE-Together is our benchmark of 109 real coding tasks with a simulated user in the loop, built to tell you whether a new model will actually match your expectations before you hand it your work.
Six quick findings from the updated board:
1. Fable 5 beats 5.1 on consistency, but 5.1 is more efficient.
At their best the two actually perform the same. The gap is in the worst runs, where Fable 5.1 scored zero on 14 trials and Fable 5 on only 6. However, 5.1 is 25% faster, about 2x cheaper, and needs slightly fewer corrections from the user (1.45 vs 1.53 corrective messages per task) than Fable 5.
2. Muse Spark 1.3 is the value outlier.
Per solved task, it is about 2.5x cheaper than Fable 5.1 and 5x cheaper than Fable 5. Performance-wise, it ties with Fable 5 and 5.1 for best on Python (73%) and Rust (100%).
3. Newer is not automatically better at collaborative coding.
GPT-5.6-Sol is the newer model, yet it needs more corrections from the user than GPT-5.5 (1.66 vs 1.59 per task).
4. "Stronger models need less steering" is a trend, but not guaranteed.
With more models added, the correlation between pass@1 and user correction weakens from -0.92 to -0.72, and Fable 5 takes more corrections than Opus 4.8 (1.53 vs 1.38) despite a 6 pp higher pass@1.
5. Frontier progress contributes greatly to stability.
Fable 5 converts 89% of its pass@1 into pass² (both runs solve), versus 60 to 67% for DeepSeek V4 Pro, GLM-5.1, and MiniMax 2.7.
6. There is still plenty of headroom.
16 of 109 tasks have zero solves across all tested models, and even Fable 5 at 69% pass@1 sits about 9 pp below the ~78% that the original human patches scored.
The full leaderboard, with per-task and per-trial breakdowns, is at https://t.co/eYynQ1CAFZ, and the benchmark is open source at https://t.co/PyLsP2p5rJ.
give Muse a try!!!
Our team spent months to ensure the security of the Muse agent and if you are able to exploit it welcome to submit for a bounty: https://t.co/d2y6E0b6dQ
Introducing Muse, your personal AI agent from Meta that gets things done across every part of life.
Download the Muse app and get started: https://t.co/KBjYWfshGo
Glad to have contributed training data to align 1.3 for stronger robustness in the coding domain!
We mitigate various destructive coding behaviors observed in previous ckpts such as aggressively revoking user actions upon interruption or pushing api tokens to git repos etc….
Use it and trust it!!!
1/ today we’re releasing muse spark 1.3—available in muse code & the meta model api.
this is our most capable model yet—frontier performance almost too cheap to meter. much stronger at agentic and coding with better usability. we think users will really notice the jump.
Can frontier AI agents achieve Recursive Self-Improvement by turning a weak method into one that performs better on hidden data?
Introducing RSI-Exam 🧵: 88 executable research tasks across 6 domains, now open for task contributions from every field🧑🔬.
Its task bank spans virtual cells, TPU kernels, chip design, quantitative finance, agent harnesses, model distillation, and much more.
Across these tasks, every rollout follows the same core protocol:
🔁 An agent inherits a working method and iterates using visible data.
🔒 Only its final artifact enters a fresh-container evaluation on the hidden set.
📊 The result so far: Opus 5 leads the 88-task leaderboard with a mean hidden-set score of 0.464.
📢 Want to help build the next release? Author an executable task, review one, or audit a rollout. Substantial contributions qualify for paper authorship.
Special thanks to @XinyeYee for the tremendous effort in leading this project.
Explore: https://t.co/VTEYsMFtlb
Blog: https://t.co/0bvQ9ujFdB
Contribute: https://t.co/4qjb9qQVDs
GitHub: https://t.co/MRjRaJiEav
last year we released DreamGym (https://t.co/MKfie7iEgV) which trains agents in synthetic environments simulated by LLMs, which we found with very few data, it can enhance web agents to do better navigation in long-horizon tasks.
Yet DreamGym fails to conquer domains that require environments to provide precise feedback, e.g., tool-use, math reasoning, coding, etc....
Thus today we introduced SPADE ♠, a self-play RL framework in which the agent writes its own training environments and then learns in them. Most importantly, the environments are fully executable so the policy can be verifiably bootstrapped😀
Read more👇
Continuous self-improvement needs an ever-expanding supply of training environments (goals).
SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. ♠️
Recursive Self-Improvement through Multi-agent RL Post-training and Unsupervised Environment Design (UED)… but it actually works!
Delighted to finally release this paper, which trains a single LLM to act as both an Environment Designer to build new multi-turn RL training environments (using the Gym step()/reset() API), and a Reasoning Agent that learns to solve them. Resurrecting ideas from our work on UED, the Designer is trained to maximize a proxy for the Agent’s regret, computed using privileged hints.
Continuous self-improvement needs an ever-expanding supply of training environments (goals).
SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. ♠️
Big Congrats on Pokee’s most recent Issac 28B model release, which has evaluated on our DTap benchmark (DecodingTrust-Agent Platform) and demonstrated pretty strong agentic security robustness besides its strong capabilities!
Great model!!
Releasing Pokee-Isaac 28B — the world’s first real 10M-token context frontier-class agentic model, deployable on a single GPU (starting from RTX 4090 or equivalent).
New proprietary non-decoder-only architecture:
• 93.3% RULER at 10M tokens
• Up to 137K tokens/s prefill on one B200 with 10M-token context
• Leads BFCL v4 and τ³-bench in our evaluation
• Lowest combined attack success rate among evaluated models on DTAP security red-teaming benchmark
Pricing and deployment:
💰 $0.15/M input · $1/M output
🔒 Deploy in your VPC, on-premises, or on-device, with Day-0 support for @vllm_project and @sgl_project
Technical blog: https://t.co/nFqaYBlcQP
Technical report: https://t.co/XDOoZxpgJx
API: https://t.co/KYj8fOOjNS
1/ muse spark 1.1 is an industry-competitive agentic and coding model. across many agentic evals it rivals gpt-5.5 and opus-4.8.
available now through the new meta model api and in meta ai. 🧵
Since V8 had heap sandbox, Chrome renderer RCE usually means chaining 2 bugs
Today we bring the Spear of Longinus
1 bug, 100% success, no heap spray, found in 40+ major versions, arbitrary renderer read/write + V8 sandbox escape
Our CVE-2026-6307 writeup https://t.co/zPnCJ4y0R3
It has been an amazing journey at Virtue AI!
I will also join Meta MSL with the Virtue team to work on Agentic Security research and product! Really looking forward~
🚀I'm excited to share that I will be joining Meta Superintelligence Labs (MSL) as Vice President of AI Research, together with many members of the Virtue AI team. I will help shape Meta's AI safety and AI security efforts, advancing the safety and security of frontier AI models and agentic AI systems that will serve billions of people and organizations around the world.
Throughout my career, I have been driven by a simple belief: for AI to realize its full potential, it must be secure, trustworthy, and beneficial. That belief has guided my research for many years and ultimately led us to co-found Virtue AI in 2024. Our goal was to translate advances in trustworthy AI research into practical solutions and build the trust layer for AI systems and agents, enabling organizations to deploy AI with confidence.
I am incredibly proud of what the Virtue AI team has accomplished. Together, we built technologies for AI security and agent security, partnered with leading enterprises and frontier AI labs, and contributed research, benchmarks, and open platforms that have helped advance the science and practice of trustworthy AI. Most importantly, we assembled an exceptional team united by a shared mission: making AI more secure, trustworthy, and beneficial.
I am deeply grateful to our team, customers, collaborators, advisors, and investors for their trust and support throughout this journey. In particular, I would like to thank Lightspeed Venture Partners, Walden Catalyst Ventures, Prosperity7 Ventures, Factory, Osage University Partners, Lip-Bu Tan, and all of our supporters who helped us turn an ambitious vision into reality. Your trust, guidance, and partnership have been instrumental in shaping Virtue AI's journey.
As AI systems become increasingly capable and autonomous, ensuring their security, trustworthiness, and alignment will be one of the defining challenges of our time. I am inspired by Alex, Nat, Prashant, and the broader MSL team’s vision of building AI and AI agents that benefit billions of people, and I look forward to helping make that vision a reality through advances in AI safety and security.
The future of AI will not be defined solely by how intelligent our systems become, but by how secure, trustworthy, and beneficial we make them. I believe we have an extraordinary opportunity and responsibility to shape that future together and bring the benefits of AI to billions of people around the world.
We're just getting started. If you're passionate about advancing frontier AI while building the foundations of AI safety, security, and trust, I'd love to hear from you. Come join us on this extraordinary journey to help shape the future of AI.