I'm excited to share that Astralis Foundation (with Macroscopic Ventures) has launched an RFP on middle powers and transformative AI!
We're looking to fund up to $10M in grants on how middle powers can make the AI transition more likely to be safe and broadly beneficial.
MATS 10.0 Symposium spotlight talks just dropped! If you're interested in AI safety & security, here is a taste of the research MATS helps support!
https://t.co/lSMybl9Mms
Problem with monitoring giant agent swarms that barely interact with humans:
- Scale: Monitoring is obviously harder with 100s of billions of tokens per task, humans can only review a tiny fraction
- Split attacks: The swarm can distribute attacks over time and over multiple agents, making single-trajectory monitoring much worse--see our work on this: https://t.co/Zw32nhM4hh
- New tech: The swarm can come up with their own infra and terminology over time, making it easier to not look suspicious
- Incident response is harder: while I think surfacing some of the most suspicious activity should be possible with CoT + action monitoring (with the caveat that the CoT part seems to eroding), actually figuring out what's going on in the swarm is basically a novel research task. E.g. METR + Redwood team took many days to get a handle on the HF incident and were still discovering new stuff over time
- Collusion: Requires a few more assumptions, but collusion between monitor and agent models seems more likely as multi-agent training ramps up/there's more a "community" of agents that the monitor can identify with
Today we're launching Project Tailwind, a major call for founders to launch ambitious AI safety initiatives: https://t.co/o4ZkHv5UmH
We're looking for exceptional people to build the research, technologies, and institutions that will help humanity navigate the risks of transformative AI. 🧵
We develop a mechanism design framework for AI alignment and control: https://t.co/7eI8H52s8h
It’s largely conceptual but we offer stylized applications to failure modes (sandbagging, alignment faking), safety practice (scalable oversight, peer prediction), and a way to think about the value of alignment, interpretability, capability, and control.
I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.
OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program.
@KatSpartz I disagree! Isn’t the idea of “genetically engineer animals to be evil” against animal welfare that EA advocates for in the first place?
Humans will literally attempt anything but never changing themselves.
@Miles_Brundage Maybe because Mythos directly targeted ‘national security’ that WH considered ‘real threats’?
Regardless, I think HF incidents should be used for developing better policy regulations or even lobbying (not even WH, but maybe CAISI?)
1/ MATS is launching the Residency program!
A new path for experienced researchers working on AI safety, or moving into it.
📍 Berkeley / London / Washington, DC
💰 $155,000–$285,000/yr + compute, no preset cap
📅 6–24 months
Apply by October 31 AoE ↓
https://t.co/0y8kEMaF90
🚨 MATS Winter 2027 applications are now open.
Fully-funded, 12-week fellowship for aspiring & established AI alignment, interpretability, security, governance researchers & field-builders
📍 Berkeley/London
📅 Jan 19–Apr 10
💰 $6.4k/mo + $8k - 16k/mo compute
Apply by Sep 6 ↓
One of the most inspiring thoughts I've heard in the last year was from a former Secretary of the Navy, when I asked him what he thought about game theoretic dynamics around the superintelligence ramp potentially driving us to kinetic war over the coming years.
They noted that historically during the Cold War, a range of people were advocating very strongly for nuclear first strikes. Some of these were coming from a military perspective (Curtis LeMay), but others such as von Neumann argued that game theory said first strike was optimal.
And then both these camps went (directly or indirectly) to various presidents, who said "No this is crazy, we're not going to kill millions of people based on your theoretical argument" (paraphrased).
I think this is an important heuristic to keep in mind when modeling the behaviors of actual humans, both because it has the potential to capture unmodeled aspects of the game theory and also because...sometimes people just refuse to do unreasonable things, even if the theory says they should.
Not all the ways in which we aren't modeled by game theory are good, but I think this is a hopeful thought on net, and it might be the thing that gets us through the next few years or decades intact.