Thanks to @AiEleuther for bringing us together through SOAR, and many thanks to @baseten for sponsoring the open-weight inference compute that made these experiments possible!
1/
New paper: Agent Memory Is a Surface for Endogenous Authorization Laundering.
Persistent memory can make LLM agents act without authorization -- even without an attacker.
Paper: https://t.co/aB53SEF2d6
Code: https://t.co/ls1GCjU5lR
6/
Once false authority is stored, executors almost always trust it: 98.6% of matched trials lead to unauthorized action.
Replace only the memory with the exact authorization state, and that falls to 0%; that localizes the bottleneck to memory.
We've pushed a version update to the Terminal-Bench dataset and leaderboard.
Terminal-Bench 4.0 calibrates task resources (time, cpu, memory), implements task fixes, and removes saturated tasks.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work.
Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort.
Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%
We compared linear attention architectures (DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2) and explored cross-layer routing.
Folks from @Alibaba_Qwen and Moonshot AI starring the repo 👀
Paper: https://t.co/JxPGUDhHCa
Code: https://t.co/u1SVdeN1a6