[5/N] SoTA Results π
Posterior Refinement delivers large efficiency gains across all considered benchmarks. On Sudoku, FMLM+ reaches 97.9β/β92.0β/β71.2 (Easy/Med./Hard) with just 4 NFEs β surpassing every baseline evaluated at 128 NFEs. On GSM8K, it reaches 19.0% accuracy at 32 NFEs, a 32Γ speedup over the strongest non-autoregressive baselines at matched accuracy. On TinyStories and OpenWebText, FMLM+ matches or surpasses every diffusion baseline with up to 8Γ fewer evaluations.
Does your discrete diffusion model know what it doesn't know? π€
In our new paper, we introduce Posterior Refinement, a framework to allow Flow Map Language Models to identify tokens they are unsure about and iteratively refine their output.
MDMs score tokens before the rest of the sequence exists (a-priori). It conflates "the data is diverse" with "the model is uncertain", and catastrophically fails on the simplest of toy problems.
We fix this with a simple idea: score each token against the complete draft. This a-posteriori confidence lets the model know what it doesn't know. Posterior Refinement is made possible using FMLM+ -- continuous flow maps equipped with masking-style noising.
We achieve SOTA across all considered datasets with up to 32x fewer NFEs than the strongest diffusion baselines. π
π§΅β¬οΈ
[4/N] Path To Scaling
We show that the masked diffusion training objective is exactly the boundary case s = t = 0 of the FMLM+ training objective. This correspondence makes the growing pool of pretrained MDMs directly usable to accelerate FMLM+ training, via either distillation (using the MDM as a teacher) or direct warm-starting from its weights.
Grateful to have contributed on this amazing project, under @max_simchowitz and @servo97's guidance.
Working on OGPO gave me a deeper appreciation for the many moving parts in modern offline-to-online RL systems, for which OGPO comes up with an intuitive & elegant solution!
Interaction with the real world is the major bottleneck in robot learning. So what would robot RL look like if we didnβt need to limit compute per interaction? Our latest work, Off-Policy Generative Policy Optimization (OGPO, accepted to ICML26) embarks on answering this question (spoiler alert: when done correctly, it helps massively!).
π§΅(1/N)