Turns out you can make LLM inference fully deterministic across devices, with no loss to quality or speed.
This weekend at the @SpaceXAI hackathon I got Qwen3-0.6B to produce identical hashed logits from a 512-token generation across 2 GPUs and 3 CPUs: an A100, an H100, an Apple M5 Max, an AMD EPYC, and an Intel Xeon.
The main reason why inference isn't deterministic is because floating-point addition is not associative. (a + b) + c ≠ a + (b + c), since every add rounds. Accumulation order changes with the hardware used and kernel selected, so the same prompt can give you different outputs even at temperature 0.
Integers ARE associative. So why doesn't integer quantization already fix this? Because, while weights and activations get quantized, the non-linear ops (softmax, normalization, SiLU) dequantize back to float and requantize afterwards. Each of those steps hands you back to floating point rounding.
True integer-only inference does exist, but it's historically been motivated by edge hardware without FPUs, which doesn't make much sense for LLMs. One 2024 paper (I-LLM) did it on LLaMA from that angle and didn't get much attention. Nobody seems to have looked at it from the determinism side.
I wrote my own implementation, simplifying the approach from the paper, so that every operation between the input ids and the int32 logits is exact integer arithmetic. To test it I chain-hashed the logits at every step and ran that across the devices and configurations below. Every integer run gave the same hash: 64430dd985f8. Every fp16 run gave a different one, all diverging on the very first token.
WikiText2 perplexity came out to 20.72 vs 20.95 for fp16 (slightly better than the float baseline), and CUDA-graphed integer decode hits 106 tok/s at batch 1 on an A100, 3.6x the fp16 eager baseline.
Github repo is listed in the comments. Plan to do a writeup over this eventually!
nvidia is worth close to india’s economy of $3.9 Trillion. let that sink in. nvda 30k employees compared to 1.5 B people. the ratio is insane.
this is jensen’s world baby.
Another fun thread about long division! We were discussing the arithmetic of complex numbers in precalculus, and I asked students how we would divide them. One student half-jokingly suggested we try long division. So we did!
(1/8)
🆕 Blog post 📣
I wanted to run JavaScript in WebAssembly, so I decided to turn JavaScript into a compiled language by building a JS-to-C++ transpiler.
How does it work? And is it a good idea? Well, that’s what the blog post is about!
👀👇
https://t.co/xjSi2S8dxv
Only 184 kB of neural nets. A new demo that will ship with the 'toolkit for baked dynamic neural global illumination in the browser' - catchy, no?
#threejs#webgl#gi#machinelearning
We need a MrMLBeast YouTube channel.
"I Trained a 10 Trillion Parameter Model to Memorize Wikipedia"
"I made ResNet50 Converge on ImageNet with a 0.000001 Learning Rate"