We believe monitoring and intervention to be a strong tool when dealing with complex systems. Read more here https://t.co/9kaOE7Cl2M
@megamor2@OriYoran
"One bad apple can spoil the bunch 🍎", and that's doubly true for language agents!
Our new paper shows how monitoring and intervention can prevent agents from going rogue, boosting performance by up to 20%. We're also releasing a new multi-agent environment 🕵️♂️
Using our method, we see consistent gains across varying difficulty levels of WhoDunitEnv. We apply our method to GovSim, a recent resource sharing environment and show an increase in both survival rates and efficiency of agents.
@giorgiopiatti
Do you have a "tell" when you are about to lie?
We find that LLMs have “tells” in their internal representations which allow estimating how knowledgeable a model is about an entity 𝘣𝘦𝘧𝘰𝘳𝘦 it generates even a single token.
Paper: https://t.co/9uJ2Zab1aC 🧵
@dhgottesman