Agentic RL CPU sandboxes spend >75% of the time idle waiting for GPUs. Wall-clock vs. active vendor billing hides a massive 20x cost gap. 1,000 always-on environments run $60k/mo on vendors like E2B, but cost $0 on your cluster's spare CPU cores. https://t.co/yfXzdZTmZ7
We are of course talking about Ajinomoto Build up Film (‘ABF’). This is the dielectric insulator film that goes into the organic package substrate of most modern processors today. How did it come from Japanese seasonings heavyweight Ajinomoto?
Umami was first discovered in 1908 by Kikunae Ikeda while studying kombu broth, umami’s secret lies in glutamic acid. This the very compound that Japanese seasoning company Ajinomoto would later crystallize as MSG (monosodium glutamate).
on TMEM: 256 KB (128 lane * 512 column * 32 bit) per cooperative thread array.
densegemm accum is in TMEM on sm100.
(same with nvfp4/mxfp8 scale factors for A/B tensors)
hopefully this thread inspired you to learn a little bit more about sm100 programming / memory!
feel free to check out this repo for all runnable scripts! https://t.co/fwVQoCCUYa
'bank conflict' is a term thrown around commonly. here is an explanation of how this actually works on sm100 and pre-sm100.
our theoretical model is only off by 0.009%
and we show how swizzled could make ldmatrix 5x faster on sm100, and densegemm inner loop can be 87% worse in the worst bank conflict case.
1/n
on sm100, thankfully most of this is flipped
tcgen mma is issued by one thread (elected by the warpgroup). 32 lanes don't spam your SMEM at once anymore.
also sm103 has bit 52 (secret tip)- for K=48B operand
23/n