Little Enby Researcher 🏳️⚧️🏳️🌈 Tech Lead at Amazon for Agent Safety research; Open Source contributor to democratize post-training喵. | 持证含糖 l SRS 09/2026
Got discharged from the hospital today and got the best possible welcome-home news: our paper is accepted to #NeurIPS2026! 🎉
Reading, Not Thinking: Why do MLLMs solve a problem as text, then fail when the same text becomes pixels?
Turns out, often they can read — they just stop thinking.
We study 7 MLLMs × 7 benchmarks × 5 input modes and find visual inputs can trigger 5–19× shorter responses and much less reasoning.
A simple self-distillation approach boosts image-mode GSM8K from 30.71% → 92.72%, with gains transferring to unseen benchmarks.
https://t.co/SHJL5KHjj9
Please read/share! ❤️
Looped transformer meets agent swarm.
Our new @Neurips2026 paper shows how to loop compute across multiple agents.
Different LLMs recursively share latent thoughts w/ each other. Better performance with up to 75% fewer tokens.
Great work led by @Jiaru_Zou w/ awesome collaborators.
New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s!
This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it.
Specifically:
-(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time.
-Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used.
-Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2.
-Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step.
-Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params.
-Hand-rolled flash attention for 64 dim heads.
There are several additions that add accuracy too:
-(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15.
-(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application.
-A couple additional dynamic skip connections in the network.
The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large,
only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead.
https://t.co/Ycrzy6JFC3
As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: https://t.co/YEfM1VpOTi
[1/n] Introducing Diffusion Reward Models (DRM)
Reward Models are at the core of LLM post-training. But almost all of them make the same simplification:
they collapse human preference into a single score.
The problem is that human feedback is often not a point — different people can hold distinct, even conflicting, judgments about the same response.
So we asked: Can a Reward Model learn the distribution of human preference?
Here’s what we found 👇
📄 Paper: https://t.co/e0EMNa3I64
💻 Code: https://t.co/KVosB5D5PI
🤗 Models: https://t.co/wwdDpOtgbL
200 high-quality, long-horizon, open-ended RL environments are here!
FrontierSmith is a NeurIPS 2026 Spotlight (top 3.7%)! We’re releasing 200 problems for studying RL on challenging open-ended tasks:
https://t.co/90FjTjAjnv
And stay tuned: FrontierCS team will soon share something very interesting about training in these environments 👀
RSI-Jev v3.0 is out, and this is the first version where RL actually helped.
Our AutoScientists loop ran 59 RL experiments on the 2B decision model. Most failed.
The pattern was pretty clear: when every example already gives the model the right answer, plain supervised training is very hard to beat.
What finally worked was giving RL a signal the labels didn’t contain: whether the model ranked the right memory first.
That made it find the right memory first 60% more often on a long-memory benchmark, with no regression elsewhere.
The lesson we’re taking from this: RL helps when the reward teaches something the labels can’t.
But honestly, the part we care about just as much is why the other 58 experiments failed, so we wrote all of that up too.
Report: https://t.co/YjaMKguYgC
Code, weights, experiments: https://t.co/KHKS4XtWzv
Students sometimes ask me if it still makes sense, in this accelerating age, to pursue a PhD in AI.
Perhaps counterintuitively, I think it's a great time to do so.
I wrote up some thoughts on this here: https://t.co/RRqGR0qi1R
"Contrastive World Models"
If world model has to reconstruct every pixel, it'll pretty much waste most of its capacity modeling irrelevant background noise.
So this paper removes Google DeepMind's Dreamer pixel decoder and instead trains the latent state to identify features of the correct future observation.
It matches Dreamer on clean environments, but performs much better with moving distractors and natural-video backgrounds.
https://t.co/S5b0QWmwuH