Open-sourcing our DeepSeek-V4-Flash + DSpark setup for 2× DGX Spark (GB10).
Stock container, no custom image. The SM120 topk fix in the wild rounds 256 down to 128 and discards half the draft candidates — we pad up to 512 instead and get +32% decode.
Recipes + scripts:
https://t.co/0w5hNbOTYv
Full frontier LLM — GLM-5.2, unpruned, all 256 experts — running on 4 small NVIDIA DGX Sparks. Locally on a shelf, not in a datacenter. 🖥️
📏 327K context · ⚡ ~25 tok/s · ☁️ multiple options
Open source on github - clone it, edit one line, run it 👇
https://t.co/KwRd6fMacm
Check out how XanuNetworks/North-Mini-Code-1.0-NVFP4 achieved 60.13 tokens/sec on text generation on NVIDIA DGX Spark with vLLM! Shout out to @cohere for making this great new agentic coding model, which now fits in only 17gb NVFP4.
View full benchmark at https://t.co/TmVjnqpQFL