I'm excited to share my latest work, which focuses on how to give causal transformers the ability to model arbitrary conditionals. This work was done with a bunch of excellent folks @EricElmoznino@g_lajoie_@le0gagn0n@sarthmit@tejaskasetty . https://t.co/paKZUnvZfZ
Excited to share my research from the Anthropic Fellows program with @TrentonBricken!
We built a "diff tool" for AI models and found features that act as behavioral switches: turn one dial down in DeepSeek/Qwen and they start talking about Tiananmen Square. Turn one up in Llama and it outputs American propaganda. We even found one that gets GPT-OSS-20B to start outputting copyrighted text.
Eager to see what else can be found with this technique!
Thanks to @AnthropicAI and @sleight_henry for the support!
🚨 New Paper 🚨
TL;DR we derive scaling laws of lr, momentum, and batch size for modern first-order optimizers through the lens of recent convergence bounds for LMO, a framework that includes normalized SGD, signSGD (approximating Adam), and Muon
https://t.co/hcdDMOvKYH
Remember all the self-distillation papers that came out last week. Well, we also propose it 😅, but…
But alongside something better 😎 π-Distill
We show that with this method, you can distill closed-source frontier models even tho their traces are hidden 🔒.
Both our methods can reach and even surpass the performance of the industry-standard SFT + RL with access to reasoning traces 🤯.
🔬And we spent ~100,000 hours GPU hours on a comprehensive analysis, not because the method is finicky, but because we wanted to understand why it works so well.
🧵
1/10
How can we predict multiple plausible targets from a single context in joint-embedding self-supervised learning (SSL)?
Check out our paper titled “Self-Supervised Learning from Structural Invariance” accepted at #ICLR2026! Previously Best Paper Award at @unireps 2025.
https://t.co/mN5e1huPO9
We introduce AdaSSL, which models the target uncertainty and relaxes the standard assumption that the positive pair share the same semantic features.
Derived from first principles, we realize @ylecun’s JEPA with a learned latent variable for jointly learning better representations and world models, extending SSL’s utility to a broader range of data types.
1/🧵
LOTs of discourse lately about the correctness of the KL-regularization term used in RLVR fine-tuning of LLMs.
Which estimator to use? Whether to add it to the reward or loss? What’s even the difference? 🤔
In our new preprint, we evaluate these choices empirically. 🧵
1/n
🚨 New paper!
“Understanding Adam Requires Better Rotation-Dependent Assumptions.”
Come check out our poster at @NeurIPSConf, or DM me if you would like to chat!
📅 Wednesday, December 3
🕐 4:30 PM PST
📍Exhibit Hall C,D,E #908
🚨Reasoning LLMs are e̵f̵f̵e̵c̵t̵i̵v̵e̵ ̵y̵e̵t̵ inefficient!
Large language models (LLMs) now solve multi-step problems by emitting extended chains of thought. During the process, they often re-derive the same intermediate steps across problems, inflating token usage and latency.
Metacognitive Reuse: turn recurring LLM reasoning into concise, reusable “behaviors”. The model learns named skills from its own chains-of-thought and reuses them to think faster & cheaper.
Arxiv 🔗 - https://t.co/zA1gB4eYTG
Very excited to release a new blog post that formalizes what it means for data to be compositional, and shows how compositionality can exist at multiple scales. Early days, but I think there may be significant implications for AI. Check it out! https://t.co/bxHUgljfK2
(1/n)🚨You can train a model solving DFT for any geometry almost without training data!🚨
Introducing Self-Refining Training for Amortized Density Functional Theory — a variational framework for learning a DFT solver that predicts the ground-state solutions for different geometries and generates the training data itself!
🧵
This work is the result of an amazing collaboration with @CristianGabell1@hatemhelal@dom_beaini@k_neklyudov
📜 Paper: arXiv:2506.01225
💻 Code: https://t.co/f6XjqV7KVa
Excited that our paper "Addressing Concept Mislabeling in Concept Bottleneck Models Through Preference Optimization" was accepted to ICML 2025!
We show how Preference Optimization can reduce the impact of noisy concept labels in CBMs.
🧵/9
New preprint! 🧠🤖
How do we build neural decoders that are:
⚡️ fast enough for real-time use
🎯 accurate across diverse tasks
🌍 generalizable to new sessions, subjects, and species?
We present POSSM, a hybrid SSM architecture that optimizes for all three of these axes!
🧵1/7
📢UNIQUE Student Symposium 2024: Limited places!
If you are interested in Neuro-AI, join us on May 8-10 @ CERVO Brain research Center (Quebec City)! Bus and hotel available - see conditions 🧠
👉Register now: https://t.co/hWmLYCmDAE
⏳Poster submission: https://t.co/OghqavfMkD
I am grateful to have a helping hand in @pierrelux, @riashatislam and @pierthodo when I first started my research career as an MSc student in 2018. The advice they gave me on reading research papers, organizing my thoughts, and thinking like a researcher is still with me. (1/N)
Bing Liu wrote the book on Lifelong Machine Learning. That's not a metaphor, folks. It's the closest thing we have to a sacred text! 📖 Join us at #CoLLAs2023 and get schooled by the master himself! 🧠 Register here: https://t.co/ueLgjFzrTR"
🌟UNIQUE Student Symposium: EXTENSION DEADLINE🌟
USS is a 2-day conference for students on Neuro-AI🧠🤖. Join us on June 5-6 at campus MIL!
⌛️Deadline to register has been extended to May 30 ! Register now: https://t.co/3zdYhrhwFJ
👉More info: https://t.co/Fqj4mpWpPq