Created an open-source implementation of this with an interactive terminal REPL, currently supporting Qwen3.5 and Gemma4.
Details at https://t.co/eymussPIip
New Anthropic research: A global workspace in language models.
Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with.
We found a strikingly similar divide inside Claude.
Across our experiments:
💫 +44 pp average SR over the original policy
💫 8 simulation and real-world tasks
💫 2 policy backbones
💫 <500 edit steps per target mode
💫 0 runtime overhead
Huge thanks to our amazing advisors! @ManlingLi_@RuohanZhang76@zhiwen_fan_ 🧵 (4/4)
You can clone behavior into a robot policy. Can you un-clone it?
🤖 Introducing Behavior Uncloning: our recent work explores how to remove undesired behavior modes from a trained robot policy.
📄 https://t.co/oTq60PXoha
🌐 https://t.co/tCrGZ4dDeC
🧵 (1/4)
During uncloning, the mode redirection signal provided by a lightweight classifier is distilled into policy weights, and a retain loss preserves task competence.
Then, the policy is deployed with the original inference path, and the classifier is discarded. 🧵 (3/4)