This is a very good teaser but the details behind the scenes are even cooler!
For some intuition on CSM, Cerebras made everyone understand that if you can pool SRAM together on a big wafer you are going to go very fast. Etched can pool the SRAM without needing to keep the dies attached. This means they can push from 8x to 128x to 1000x+ chips all working together in ultra low-latency scale up domains, and then at the same time they can use LVI to pack the die with tons of flops for throughput. These are beautiful ideas and it’s great to see them coming to life!
@SouthernValue95 It literally doesn't matter? CXMT cannot satisfy China demand this decade. So there's 0nreason besides the small price arbitrage. It's like bannin Iranian oil. India and China buy it anyways, but price diff is less
🚨 Why does Self-Play RL for LLMs keep collapsing? Most fixes focus on the reward signal. In our new paper "Survive or Collapse", we show that's the wrong lever. The true binding constraint is actually Data Gating: deciding which generated tasks enter the training pool. 🧵 1/n
I will be presenting our recent work on writing efficient GDN kernels on B200 GPUs at MLSys 2026 today (11 am–1 pm PDT)!
FlashInfer ran a kernel competition for B200 GPUs. Our team (@thepushkarp + me) won 🥇 1st place on the Gated Delta Net track (more details here: https://t.co/QnVVWWE6cr)
Do join if you are around.
#MLSys2026
We built a kernel abstraction to rewrite the entire transformer stack as GEMM + Epilogue kernels!
Neural net architectures such as transformers consist entirely of matrix multiplications and elementwise nonlinearities such as RMSNorm, log sum exp, and gated activations. Fusing these elementwise nonlinearities into GEMMs in both the forward and backward passes allows us to make training and prefill as compute-bound as possible!
Our kernel abstraction CODA is implemented in CuTeDSL, and by abstracting away the fixed prologue and main loop of the GEMM kernel, we expose an epilogue function where LLMs like Claude can easily implement elementwise nonlinearities in fusions approaching speed-of-light!
Your drifting model is secretly a fixed point for the Wasserstein gradient flow on...
...the KL?
...an approximation to the Sinkhorn?
...Is it even a Wasserstein gradient flow at all?
https://t.co/QJLh86Hi0d
@liwenliang@agalashov@JamesTThorn@ValentinDeBort1@ArnaudDoucet1
My NeurIPS 2025 paper looked into the numerical errors present in synthetic data used for neural surrogates... and uncovered a counter-intuitive finding. Since the paper can be hard to digest, here is a lightweight introduction: https://t.co/M3sw0Yj8bM
Really elegant idea to use symmetry as a guiding principle to derive several popular optimizers, incuding Muon. Looking forward to reading this more deeply!
Why do you need to talk about CUDA streams and CUDA events for a blog post on Continuous Batching?
I have recently started reading more about LLM inference optimization. Upon asking around, I was quickly greeted by the term "Continuos Batching" (CB).
> batch your input prompts
> some generations finish ahead of others
> swap out completed generations with new prompts
The CPU needs to prepare the batch and swap out completed generations while also scheduling it to the GPU. This is a classic case of computation and communication parity.
If the communication is blocking, the GPUs are kept idle and it hurts the performance of CB. How do we work on this?
@remi_or_ in his article takes us through the CPU and GPU decoupling methods in CB.
> What are CUDA streams
> Why need CUDA events
> Forcing synchronizations to mitigate race conditions
> Bringing it all together
PS: It is one of the best blog posts I have ever read, and I request you all to give it the love it deserves.
The most underrated math theorem but Google secretly used it to change the world.
Perron-Frobenius Theorem (positive matrix version):
Let A be an n×n matrix with every entry a_{ij} > 0. Then there exists a unique positive real number λ > 0 (the Perron root) and a unique (up to scaling) positive vector x > 0 such that:
A x = λ x
Moreover:
λ = ρ(A) (spectral radius of A)
λ > |μ| for every other eigenvalue μ
λ is simple (algebraic multiplicity = 1)
Your search results aren't just a list; they are the coordinates of a high-dimensional vector pointing toward the most authoritative nodes on the web.
Today in continuous diffusion language models, we have:
- Spherical flows https://t.co/Pc0DhnKU7M
- Hyperspherical flows https://t.co/FmvxuepBnG
Another case of convergent evolution! Two different takes on the same core idea, published within days of each other.
In my next blog post, I hope to share my experience of reading the CPU and GPU trace of a Large Language model from scratch.
While working on it, I wanted to share my traces with other colleagues of mine and was quite amazed that there was no good way to do it.
I have devised a simple way to share traces now:
- Build a HF bucket
- Sync the HF bucket with your local traces
- Use `https://t.co/p6LSyO4jDm<bucket url>` link to share
If you need a GIST, let me know. 🤗