📸👨🏽💻 -PhD Student at the intersection between Theory of Mind, MARL and Human-AI cooperation at IMPRS-IS and the HCI-CS department at Stuttgart University
And another one from Matteo Bortoletto! In our work we test different cognitively inspired neural network architectures in multi-modal social interactions to reason about the mental states of humans. He will present the work at @ecai2024! ArXiv: https://t.co/eAu6VpRvv1
I wrote a mech interp reading list two years ago, but a LOT has changed - announcing v2!
This is a highly opinionated list of favourite papers, not a lit review - I try to summarise key takeaways, critique, advise on which parts to prioritise vs skip, and explain WHY I like it!
Shout out to Matteo for leading this delightful study on mental state representations in language models! Talk to him at the Mechanistic Interpretability workshop at ICML if you are there about language models, mechanistic interpretability and Theory of Mind!
Important new work!
I somewhat disagree with the conclusions though: imo persuasion capabilities will very likely surpass this curve once key players start performing "persuasion RL training/unhobbling" (beyond RLHF). Unfortunately, there will be many incentives to do this!
We prove that transformers can implement temporal difference learning in the forward pass. It's not only that TFs behave like TD, it's that the forward pass over layers is exactly equivalent to iterations of TD! https://t.co/blLpDdPPG3 w/ @wangjiuqi, Ethan, @HadiDaneshmand
Effective altruism, longtermism, catastrophic risk and billionaires
- why reducing catastrophic risk matters
- fitting it with global poverty and animal welfare
- why billionaires are funding it
Many people helped with this, errors my own