Seriously, what is the goal of today’s visual generative models?
Are pretty videos/images and low FIDs enough – or should we also demand something closer to human-like creativity? Our paper tries to answer this question 🧵
No player has been involved in more shots during a game at the 2026 FIFA World Cup than Mohamed Salah has versus New Zealand today (10 - 5 shots, 5 chances created).
We’re training models wrong and it’s due to chatGPT. Even the modern coding agents used daily still use message-based exchanges: They send messages to users, to themselves (CoT) and to tools, and receive messages in turn.
This bottlenecks even very intelligent agents to a single stream. The models cannot read while writing, cannot act while thinking and cannot think while processing information.
In our new paper, see below, we discuss LLMs with parallel streams. We show that multi-stream LLMs can …
🔵Be created by instruction-tuning for the stream format
🔵Simplify user and tool use UX removing many pain points with agents and chat models (such as having to interrupt the model to get a word in)
🔵Multi-Stream LLMs are fast, they can predict+read tokens in all streams in parallel in each forward pass, improving latency
🔵 LLMs with multiple streams have an easier time encoding a separation of concerns, improving security
🔵 LLMs with many internal streams provide a legible form of parallel/cont. reasoning. Even if the main CoT stream is accidentally pressured or too focused on a particular task to voice concerns, other internal streams can subvocalize concerns that would otherwise not be verbalized.
Does this sound related to a recent thinky post :) - Yes, but I don’t feel so bad about being outshipped with such a cool report on their side by 23 hours. I’ll link a 2nd thread below with a more direct comparison. I actually think both are complementary in interesting ways.
Along with Categorical Flow Maps and Flow Map Language Models, we now have three separate papers heralding the triumphant return of continuous methods for language diffusion😶🌫️
Can you tell I'm excited?🫨
https://t.co/SKS4OFtSG8
https://t.co/kJ3cuFsggd
https://t.co/aXXU4bUSMT
After supervising 20+ papers, I have highly opinionated views on writing great ML papers. When I entered the field I found this all frustratingly opaque
So I wrote a guide on turning research into high-quality papers with scientific integrity! Hopefully still useful for NeurIPS
Who's at #ICLR2026 next week?
I'm not (😢), which means someone else is going to have to organise a diffusion circle! I'm counting on you!
Let me know, happy to amplify the announcement📢
Students trained on teacher-generated data don’t just learn the task, they can inherit hidden teacher biases, even from seemingly harmless data. Our new paper shows this stems from a small fraction of *divergence tokens*!
1/n
Flow-LLM Blogpost :D https://t.co/0HiyNPJHsk
In the last few weeks, a bunch of work on flows for language came out 🌊
That is exciting, because it makes truly parallel text generation feel real: generation where models can keep refining the whole response during inference, instead of committing token by token.
I wrote an intuitive and animated introduction to the area — why autoregression has a structural ceiling, why discrete diffusion only partly escapes it, and why flows may be the first genuinely parallel alternative.
Here's an overview of the key parts of the blog - and let's chat at #ICLR2026 :)
You don't imagine the future by mentally rendering a movie. You trace how things move -- abstractly, sparsely, step by step.
We built a model that does exactly this. It predicts motion, not pixels -- and it's 3,000× faster than video world models.
Myriad, accepted at @CVPR 2026
🚨2 PhD positions with me @AmlabUva on learning causally grounded concepts 🚨
Are you interested in improving the #interpretability, #robustness and #safety of AI by integrating #causal reasoning? Join us in beautiful Amsterdam 🇳🇱🌷🚲
Deadline: 20 April
https://t.co/fN8UVRLyXF
The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything else. Knowing how to evaluate models in a way that correlates with actual usefulness is a core competency for AI labs, and any new lab is unlikely to be successful without first figuring that out.
🤯 big update to our flow map language models paper!
we believe this is the future of non-autoregressive text generation.
read about it in the blog: https://t.co/DfBXrYmJc8
full details in the paper: https://t.co/coiNXj4ucC
we introduce a new class of continuous flow-based language models and distill them into their corresponding flow map for one-step text generation.
we beat all discrete diffusion baselines at ~8x speed!
v2 gives a complete theory of the flow map over discrete data, with three equivalent ways to learn it (semigroup, lagrangian, eulerian). it turns out you can train these with cross-entropy objectives that look very similar to standard discrete diffusion — but without the factorization error that kills discrete methods at few steps.
beyond improving results across the board, we showcase properties that are unique to continuous flows. in particular, inference-time steering and guidance become straightforward. autoguidance brings generative perplexity down to 51.6 on LM1B, while discrete baselines completely collapse at the same guidance scale.
we also show reward-guided generation for steering topic, sentiment, grammaticality, and safety at inference time — and it works even at 1-2 steps with our flow map model. simple, well-understood techniques from continuous flows just work incredibly well in practice for language.
we’re extremely excited about the future of this class of models.
stay tuned for results on scaling, reasoning, and reinforcement learning-based fine-tuning. 🚀
Gaussians are not empty inside!
This is common mis-"common-misconception"
When you look at the chi squared distribution it feels like gaussians are essentially shell with radius d, which is the visualization here.
But its not the case. The inside has strictly larger density.
Its just the nature of high dimensional space (where most of the oranges are at the peel @tszzl) thats pumping more space at the outer crust, not more density!
Btw ||x|| ~ d^1/2 ± O(d^1/4) (as chi squared dist has mean d and std sqrt(2d))
Egyptian programmer Badr El-Khamisy launched a digital initiative to honor every Palestinian who has been killed in Gaza
So far, over 60,000 names have been documented, each represented as a point of light on the screen. Clicking a point reveals their name, age, and birthday.