My colleague Wesley Blakemore (https://t.co/XFhu0PzkPn) designed the latest sub-agent experience in Code Puppy (https://t.co/yuK1Nssslj), and I LOVE it.
It's cooking on 16 GitHub issues in parallel and using < 400mb of RAM.
I'm loving it.
@MiaAI_lab@NVIDIAAI You know, I hadn't even considered TP=3 + a ring topology when I ordered by 3rd + 4th spark + the CRS804. I don't regret spending the money though, except occasionally I daydream about piling all of the net worth of my home lab into RTX 6000s.
@MiaAI_lab@NVIDIAAI The model is nice overall. Deepseek is better but m3 has vision. It sometimes overthinks and thatโs one issue with it, but overall I like it a lot.
@MiaAI_lab@NVIDIAAI I got it to 39 single stream with almost no optimization. I currently had to RMA 2 of my sparks so I'm running your INCREDIBLE Deepseek recipe. <3
https://t.co/8kfOMHTEiK - here's my recipe for M3.
Check out how nvidia/MiniMax-M3-NVFP4 achieved 38.99 tokens/sec on text generation on NVIDIA DGX Spark with vLLM!
View full benchmark at https://t.co/GZbBjY3CEK
This recipe uses https://t.co/HUfygAHnBe to perform speculative decoding.
It runs on 4x DGX Spark machines networked together with Tensor Parallelism.
@MiaAI_lab@NVIDIAAI Hey there! You gained a follower today. I'm a 4x Sparker and my preferred model is Minimax M3 NVFP4. I have not gotten GLM 5.2 working well in a stable way.