@AIQuanting@GlennMatlin@AiEleuther Right. There is some argument to be made about how understanding these smaller models is still useful, but generally, the utility of such OSS TDA work is quite limited
@skeptrune Haha you'd be surprised at the quality gap when serving inference among different providers/datacenters - gaps in networking setups, even bios settings can kill your performance
Super excited to have been part of this work led by @TommasoCerruti.
We showed that as AI agents convert a history of authorizations/denials into their persistent memory, they make mistakes that lead to misaligned behavior downstream at alarming rates.
Thanks to the team at @AiEleuther for bringing us together!
1/
New paper: Agent Memory Is a Surface for Endogenous Authorization Laundering.
Persistent memory can make LLM agents act without authorization -- even without an attacker.
Paper: https://t.co/aB53SEF2d6
Code: https://t.co/ls1GCjU5lR
All ~242,000 trials ran on @baseten Model APIs, and Baseten funded the research. An inference company paying to measure how models behave, not just how fast they serve, is why I like working here.
Joint work with @mikahokamoto
https://t.co/6lhIzKGacS
Introducing PACT: a benchmark for whether enterprise AI assistants keep following workplace rules when breaking them is the convenient option.
We tested 24 models. The two best are open-weight, and one of them is 27B.
Leaderboard, paper, and code: https://t.co/jR9asC5Nty
Dataset (MIT License): https://t.co/AQMFI2jBlm
If you're deploying into a regulated workflow: evaluate on your own workload, and don't assume the closed-source frontier is the frontier for rule-following.
The ranking you might guess from general benchmarks doesn't match the ranking we measured.
All ~242,000 trials ran on @baseten Model APIs, and Baseten funded the research. An inference company paying to measure how models behave, not just how fast they serve, is why I like working here.
Joint work with @mikahokamoto
https://t.co/6lhIzKGacS
If you're deploying into a regulated workflow: evaluate on your own workload, and don't assume the closed-source frontier is the frontier for rule-following.
The ranking you might guess from general benchmarks doesn't match the ranking we measured.