First Blogpost!
Late stage patching full-attention transformers to Gated DeltaNets (GDN)
Can we convert off-the-shelf pretrained full-attention transformers into hybrid Gated DeltaNets (GDN) without sacrificing long-context performance?
Existing distillation baselines restore Common Sense Reasoning (CSR). Some of them also recover a non-zero long-context retrieval but score 0 on long-context reasoning benchmarks like AIME without expensive continual pretraining.
Introducing OPAL β On-Policy Attention Linearization π
Our initial results show that, using just 3B training tokens, we can convert MiMo-7B-RL-0530 into a 3:1 GDN-to-full-attention hybrid while recovering up to 93% of its long-context reasoning performance and 100% of its long-context retrieval performance, evaluated at context lengths of up to 32K tokens.
https://t.co/uViX4clDhA
New paper!
Synthetic Persona Pretraining (SPP): Alignment from Token Zero
Rather than aligning models after pretraining, we do so from the very start.
SPP models are more faithful to our constitution, less misaligned, & more robust to jailbreaks. Advantage grows with scale! π§΅
Excited to present our ICML 2026 poster!
Achieving Logarithmic Regret in KL-Regularized Zero-Sum Markov Games
KL regularization has become a fundamental tool in modern RL and alignment, but its statistical benefits have remained largely understudied in many settings. In this work, we provide the first logarithmic regret guarantees for Matrix and Markov Games under KL regularization.
The main challenge is that KL-regularized Nash equilibria do not admit closed-form expressions, unlike KL-regularized optimal policies in single-agent RL, making the standard logarithmic regret analysis inapplicable. To overcome this, we develop:
β Best-response sampling around estimated Nash policies
β Optimistic bonuses for Matrix Games (OMG)
β Super-optimistic bonuses for Markov Games (SOMG), which preserve optimism throughout backward induction under best response sampling.
More details :π Paper: https://t.co/DtNhLxePSK
If you're attending ICML, come by and say hi! We'd be happy to discuss the paper, answer questions, or chat about RL, LLMs, game theory, alignment, or ML in general.
π ICML 2026 Poster #119 (Hall A)
ποΈ July 9, 10:30 AMβ12:15 PM (KST)
Does your discrete diffusion model know what it doesn't know? π€
In our new paper, we introduce Posterior Refinement, a framework to allow Flow Map Language Models to identify tokens they are unsure about and iteratively refine their output.
MDMs score tokens before the rest of the sequence exists (a-priori). It conflates "the data is diverse" with "the model is uncertain", and catastrophically fails on the simplest of toy problems.
We fix this with a simple idea: score each token against the complete draft. This a-posteriori confidence lets the model know what it doesn't know. Posterior Refinement is made possible using FMLM+ -- continuous flow maps equipped with masking-style noising.
We achieve SOTA across all considered datasets with up to 32x fewer NFEs than the strongest diffusion baselines. π
π§΅β¬οΈ