GLM-5.2 on 8× DGX Spark — v18 update 🚀
Jump from 1,200 → 1,329 t/s prefill
Peak decode hits 66 t/s (was 54)
Repo with both v16 and v18 builds:
https://t.co/uFSmpdqw3R
Forum post with benchmark:
https://t.co/mMeqh3Bgcj
Kimi K3 Coder Reap on 8xGB10 - 32t/s avg gen speed in coding task.
This is while waiting for the full 16xgb10 setup to run the full k3. Considering that this one has same no of active experts, it is a good speed for my first, unoptimized, run of KIMI-K3-CODER-REAP-320-MXFP4 with dspark.
I will work on improving these results and let you know.
🚨🚨🚨🚨🚨: most requested since the DeepSeekV4-Flash GA 0731 release yesterday.
Now Abliterated 32/32 100% compatible with DSpark
please if you like my work and want to contribute please can help via X-money, GoFundMe link in profile 🙏 any little bit helps, whether for token credits, coffee, time, power, hardware.
HF repo:
https://t.co/g5xjS1dOb8
@antoniodearmas What is fast enough for you? 50tps+ single stream in coding for glm5.2 is good enough? For me, yes. Multiple requests are also scaling well, till 6-8.