1/ New paper on AI control with @olivermakins@ShamsZohreh@MaryPhuong10 🧵
A misaligned AI deployed as many monitored copies can't do much alone. But what if the copies coordinate, splitting one attack so each piece slips past its own monitor?
When we began working on this in January, all of this was very speculative, and the hypothesis cooperation could happen unmonitored was heavily questioned. Fast forward to August...
When we began working on this in January, all of this was very speculative, and the hypothesis cooperation could happen unmonitored was heavily questioned. Fast forward to August...
1/ New paper on AI control with @olivermakins@ShamsZohreh@MaryPhuong10 🧵
A misaligned AI deployed as many monitored copies can't do much alone. But what if the copies coordinate, splitting one attack so each piece slips past its own monitor?
Quote from OpenAI on the livestream.
'Already, Sol has been transforming our research program. As one example, GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna.'
8/ A monitor reading the whole squashed attack catches it easily. But a real defender can't know which commits belong together. FakeLab (9 services, 86 tasks, 4 attack types) is shared with safety researchers on request. Paper: https://t.co/AwHqvXYscy
1/ New paper on AI control with @olivermakins@ShamsZohreh@MaryPhuong10 🧵
A misaligned AI deployed as many monitored copies can't do much alone. But what if the copies coordinate, splitting one attack so each piece slips past its own monitor?
7/ One hopeful result: the effect varies by attack type. Stealing model weights resists splitting because its core action, handing out raw weights, is intrinsically alarming and no cover story hides it. Every attack keeps ~1 hard-to-hide commit. That's the monitor's foothold.
@Ric_RTP Speculation. If true, this means the natural outcome of AI takeoff is essentially the resource curse. (Full reasoning: because it's so capital intensive it ends up concentrated so it can be controlled by a small number of actors, therefore curse)
Another quite successful prediction by @DKokotajlo : Fable is intentionally nerfed for frontier ML research. This is within ~3 months of Daniel's prediction of Q1 2026 (made in 2023).
Although I don't think Mythos is automating ML research to the same extent as his prediction.
I don't really agree that nobody tells you. This screenshot is from a very famous 2014 blog post. We already knew everything we needed back then, it's just that few had eyes to see