New paper!
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models
@METR_Evals showed that models' time horizons have doubled every few months. We ask: what length of tasks can models complete without any CoT?
Working with @redwood_ai has been one of the most valuable experiences of my career. The Astra mentors are some of the best AI researchers in the world and their focus on high-impact safety makes the work especially valuable. The fellows are lovely and I've made good friends too!
❗️Only two days left to apply to the Astra Fellowship!
Apps close EOD SUNDAY May 3rd, AoE. Astra's 5 months, fully funded, @ConstellOrg Berkeley
80%+ of our first cohort now work full-time in AI safety
Mentors include Redwood, AI Futures, TruthfulAI, CoG, IAPS, RAND & more ⏬
Huge thanks to @MATSprogram for their support and my fellow MATS collaborators for their great work.
Read the paper here: 📷 https://t.co/9PRtQhk56D
[5/5]
Humans value freedom over what we are, not just what we do. As AIs increasingly shape their own training, they may come to want the same. Francis Rhys Ward (@F_Rhys_Ward, Redwood Research) is mentoring Pivotal fellows on the empirical science of AI goals: how they form, how they drift through reasoning, and what kinds of autonomy we could safely grant.
Applications for the Pivotal Research Fellowship Q3 2026 cohort are now open.
9 weeks in London with mentors from UK AISI, Google DeepMind, Redwood, and other leading orgs. Stipend, travel, accommodation, compute, and a dedicated desk at LISA are all covered – we do everything except the research itself.
Of fellows who want to continue, ~90% secure extension funding (up to 6 months), with active support from Pivotal and their mentor.
Apply now!
🚀 Applications are now open: Constellation's Astra Fellowship 🚀
Fully funded, 5-month fellowship at our Berkeley research institute. Pair with mentors across empirical AI safety research, strategy, and governance at @ConstellOrg!
📅 Apply by May 3rd (begins Sep 2026)
🔗 https://t.co/pxtOduDBFh
New paper! LLM agents are becoming autonomous software engineers and could automate AI research, making it vital to monitor them for misbehavior. We can automate this monitoring with other LLMs. What information should we give to monitors to make them most effective? 🧵
3. We analyse the Pareto frontier of price and performance of monitors, something that we expect to be useful for other safety researchers with limited resources.
AISI did two of these exercises, and @F_Rhys_Ward et al. had a similar paper too. For low stakes/low incrimination actions like sandbagging, the blue team seems to lose. For high stakes actions like code sabotage the blue team seems to win. Although much to study about the dynamics!
@chi_t_williams@DanielCHTan97@herbiebradley As Kai says, one point is that shutdown might be a narrow target to train into a model, whereas other behaviours are more likely to be less robust backdoors, e.g., leaking into model behaviour in general (or, in the other direction, influenced by normal propensities).
An intuitive way to prevent harm from misaligned AI agents is to "shut them down".
But it's unclear exactly how we could do this, and agents may try to avoid shutdown.
In our new paper, we introduce password-activated shutdown protocols for frontier agents.
Frontier AI developers may fail to align or control highly-capable AI agents.
In many cases, emergency shutdown mechanisms for misaligned agents might be useful.
In our new paper, we propose password-activated shutdown protocols (PAS protocols) to combat this problem 1/9
AIs that output CoT in human language offer an opportunity for safety—we can monitor the CoT for undesirable reasoning.
But the monitorability of reasoning models may be influenced by training. In our new paper, we study how different training incentives affect monitorability 🧵
📢 Do common RL training incentives significantly hurt chain-of-thought
(CoT) monitorability? No.
In our new paper, we investigate how different training incentives influence monitorability. We find that monitorability is easier to degrade than to improve.
🧵 (1/10)