The safety funding ecosystem is scaling up: our grant took six weeks from first conversation to confirmation. More capital is entering the field fast via the OpenAI Foundation and the Anthropic IPO. It's time for everyone in AI safety to be more ambitious.
We're excited to announce that Resolution has a $160M grant from Coefficient Giving: $108M unconditional, with a further $52M conditional on hiring and compute needs. We'll use it to grow teams across our research portfolio and invest heavily in research automation. 🧵
📢 Do common RL training incentives significantly hurt chain-of-thought
(CoT) monitorability? No.
In our new paper, we investigate how different training incentives influence monitorability. We find that monitorability is easier to degrade than to improve.
🧵 (1/10)
Every frontier AI system should be grounded in a core commitment: to protect human joy and endeavour. Today, we launch @LawZero_, a nonprofit dedicated to advancing safe-by-design AI. https://t.co/6VJecvaXYT
Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?
@Yoshua_Bengio, Michael Cohen (@Michael05156007), Damiano Fornasiere, Joumana Ghosn, Pietro Greiner, Matt MacDermott (@MattMacDermott1), Sören Mindermann (@sorenmind), Adam Oberman (@oberman_adam), Jesse Richardson (@PoliticalKiwi), Oliver Richardson
Marc-Antoine Rondeau, Pierre-Luc St-Charles, David Williams-King
@Mila_Quebec
Work done with @james_D_fox, Francesco Belardinelli, and @tom4everitt. See more work that uses causal models to help design safe AI systems at https://t.co/cjuSHFqGI4
🔦 @NeurIPS2024 spotlight paper we’re presenting today. Making AI systems more agentic is a hot research topic. But powerful agents bring worries about misalignment and loss of control. Can we measure how agentic an AI system is? 🧵
As we move towards more powerful AI, it becomes urgent to better understand the risks in a mathematically rigorous and quantifiable way and use that knowledge to mitigate them. More in my latest blog entry where I describe our recent paper on that topic.
https://t.co/emiQxTvWrd
How should we understand A.I. agents?
This blog by @tom4everitt provides one of the clearest and most complete accounts I've seen yet. Well worth checking out – alongside the wider causality research agenda:
https://t.co/TTxuMwxfZq
Do LMs deceive and manipulate? Can they be held responsible for their actions? Do they have agency? These questions all depend on the concept of *intention*. In our @AAMASconf paper, we provide a notion of intention which, we argue, can be applied to LMs. https://t.co/4C7SN6Ur9o
Excited to share our new paper https://t.co/JijwCc66w1 (Oral, ICLR 2024, w/ @tom4everitt, @GoogleDeepMind). In it we answer the question, do agents need to learn causal world models? https://t.co/JijwCc66w1. 🧵
@wooldridgemike@lrhammond@MattMacDermott1 Multi-agent influence diagrams (MAIDs) are a popular graphical representation of games and have recently been used for AI Safety problems (https://t.co/YiytPx8lRv)
Our paper extends the underlying theory to cover imperfect recall settings.
https://t.co/LWHSi68Kyt