First blog! Part 1 of writing speed-of-light GEMM kernels on Blackwell. This part covers basic GPU architecture, CuTeDSL fundamentals, loop tiling/TMA/TMEM, and swizzling.
https://t.co/MiDlRXzZoE
This is Google’s new diffusion LLM, DiffusionGemma’s denoising canvas over time.
Diffusion LLMs can generate tokens in flexible order. But in practice, do they just become autoregressive anyway?
1/ 🧵
But for both, once a token is accepted, it’s unlikely to be re-noised.
Google claims a benefit of DiffusionGemma’s uniform state diffusion over masked diffusion (e.g., LLaDA, Dream) is error correction via re-noising - how often does this happen in practice?
3/4
Trying a a magic squares puzzle prompt:
The prose tokens have the same pattern, but not the magic square tokens!
Perhaps a better model of diffusion LLM denoising is “easiest token first”, which for prose induces casual order.
2/4
Long-tail scenarios remain a major challenge for autonomous driving. Unusual events—like accidents or construction zones—are underrepresented in driving data, yet require semantic and commonsense reasoning grounded in control.
We propose SteerVLA, a framework that uses VLM reasoning to steer a driving policy via grounded, fine-grained language instructions.
Paper: https://t.co/xtvYQeQWH1
Website: https://t.co/7KjBgHE8BX
@erisaonX If we list all the births in the world in chronological order, it will behave the same as fair coin flips, there is no special meaning in “starting” or “stopping”
@erisaonX As long as all births are independent, any stopping function of the family’s choosing will result in 50-50. The stopping function across different families can also be different.