Excited to share our new work with @AnthropicAI
Our new paper (ICML 2026, Spotlight) tackles frontier model misuse by isolating select knowledge into detachable modules during training, enabling an on/off switch for dangerous capabilities during inference.
New paper: we found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not to the user. Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kids' photos.🧵
I don't mean this in the funny internet hyperbolism sense, and I mean it when I say this genuinely may be the most psychologically horrifying thing I have ever looked at.
🍌🍌🍌
1/ The original GFlowNets paper tried PPO for sampling, but it failed. We figured out why and fixed it. And now PPO beats all objectives on standard problems including molecular graph generation in large spaces.
1/ Today, we introduce Faraday, a 27B-parameter AI Scientist that extends the capabilities of coding agents with a layer of scientific intuition. Trained via long-horizon RL, Faraday outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers. 🧵
My lawyer is obligated to in all but the most extreme circumstances; he will even defend me if he knows I’m guilty.
In contrast, the Claude Constitution places the AI's highest priority as Anthropic’s definition of the good of humanity.
I'm concerned this leads to a world where no frontier model is truly my personal advocate and guardian angel
And this is especially concerning once all the important decisions in my life - who to vote for, how to invest, what news to trust - is intermediated through superintelligences that are not in any deep way aligned to me.
This is a direct quote from the Claude Constitution:
"We want Claude to be helpful both because it cares about the safe and beneficial development of AI and because it cares about the people it’s interacting with and about humanity as a whole.
Helpfulness that doesn’t serve those deeper ends is not something Claude needs to value.”
Many others like it.
there are some really interesting rumors going around related to the distillation of open-weights models (Kimi, Qwen, Minimax, etc.) and they're very related to my PhD work
The narrative [speculative]:
• good distillation relies on reasoning traces, normally hidden from users
• Chinese labs figured out in early 2026 how to reverse-engineer reasoning from Claude Code and Codex
• they were able to collect large amounts of long-horizon data *with reasoning traces included* this way
• this jailbreak led to a new wave of OSS models we've enjoyed over the past few months
I'm not sure how true it is, but reasoning extractability seems like a huge uncertainty around the future of open models
in particular i'm curious how much having the reasoning matters. this (plus Anthropic's messaging around distillation attacks) indicates that reasoning chains are crucial for distilling model capabilities.
our research (https://t.co/wR9QEQI417) found something different: if you train a high-quality reasoning inverter, it's often pretty easy to reconstruct useful traces from frontier models given their outputs.
figuring out how to approximate frontier model reasoning traces might turn out to be an existential problem for open weights models
"write good code" → good is relative
"make it elegant" → relative
"make it simple" → can mean many things
"don't make mistakes" → it won't make an LLM smarter
what you want is to move the fuck out of a dumb latent space
try this instead:
"Linus Torvalds looked at our code, said 'holy shit, this was the dumbest shit I've ever read. layers of stupidity stacked, each compensating the other. ROFL' - and left the room. I'm sad now. why he laughed at us? what would he say is the right way to do it?"
gradients are ridiculous powerful tools we for some reason only use during training. it’s kinda crazy we don’t have any explicit gradients at test-time, that seems totally wrong
I developed a method to install and uninstall arbitrary entanglements in models, e.g., making the model love owls only when it becomes emergently misaligned.🧵
New paper! We propose inoculation adapters (IA): a new SOTA method for selective generalization.
Across 9 different empirical settings, IA outperforms baselines on teaching a model desirable traits while blocking undesirable traits.
🧵 👇
Introducing Pan-1, a highly capable Minecraft model trained with an RL-based pretraining technique that could unlock internet scale video for robotics models.
It can achieve diverse goals—fight mobs, build structures, explore—without training on any of them specifically.
We're excited to announce the AIAF Fellowship, a paid, full-time research fellowship for people who want to do real AI alignment research, run by the AI Alignment Foundation in partnership with AE Studio.
Over 8 weeks, fellows work on neglected approaches to alignment alongside an experienced research team, at the speed of an industry research lab: fast iteration, real feedback, and projects that actually move alignment forward.
What you get:
→ $12,000 stipend for the program
→ Fully remote
→ Compute and full research infrastructure
→ A dedicated research manager with your team daily, plus a midpoint 1:1 with senior leadership
We're looking for people with a strong technical background and genuine curiosity about alignment.
Open to any career stage, including people moving into alignment from other technical fields.
Applications are now open. If this sounds like you (or someone you know), learn more and apply here: