@Tech2Wild Funny you should ask. I just was too impatient with GLM 5.3 Flash, even on dual Sparks. Hermes seemed to be choking it further. DSV4F is so nice (especially with the updates) and the vision version will be even better once the dust settles. Hello deep seek, my old friend.
@MiaAI_lab I take back everything I said about @MiaAI_lab . Your iteration speed is remarkable. The newest recipe fixes all the bugs I found and I'm back to EXL3 with Hermes on dual sparks. And I love the one line switch between abliterated and stock. Still the Queen!
No disrepect to @MiaAI_lab and @MichaelGannotti but after benchmarking all the recipes I could find for GLM 5.3-Flash with @NousResearch's Hermes on dual Sparks the one that best fits my use case is @sfxnz's open vLLM recipe. I'm not tokenmaxxing I just want a smart reliable local model.
@ShawnTempesta Just to put this all in perspective, when I started in radio -- 50 years ago! -- my first paying gig was $700 a month. I didn't own a car. I didn't own a TV. But man, I loved that job.
What kind of local models could you run on an M5 Ultra with 512GB of unified RAM at 1.2TB/sec? And how much would you be willing to pay for that? It's got to be $20K.
But where are they getting the memory??
Interesting. I asked Fable to design seven of the hardest coding problems it could and gave them to Ox Alpha, Grok 4.6, and GLM 5.3. Ox Alpha and Grok both aced the tests. GLM 5.3 scored 73%.
I don't know what Ox Alpha is but if it ends up being open weights this will be a watershed moment.
I'd love to see it - I run your recipe for DSV4F on my dual sparks and it's a thing of beauty. Thank you!
I'm running Orca's Qwen3.8-27B-Uncensored-Q4_K_M.gguf on my 3090. Quant is Q4_K_M, KV cache is unquantized F16, and effective context is 24,576 tokens (49,152 split across 2 parallel slots).
On a whim I bought a used RTX 3090 last week and upgraded my old gaming machine. Qwen3.8 27B NVFP4 runs beautifully on it. And with the new vision capabilities I've retired the old Qwen VL on the Framework.
3.8 is so good at vision, it even sees me as a young man. Well, younger. 50s-60s. I'm keeping it!
I was questioning my common sense when I traded my 3070 in for a used RTX 3090. And then I knew I'd gone totally crazy when I bought *two* DGX Sparks (spending my kids inheritance).
And yet, today, running Hermes with Qwen 3.8 27B on one machine and DeepSeek V4 Flash 731 on the Sparks, all totally local, I think maybe I wasn't so nuts after all.
We live in amazing times.
Iโm so glad we were at black hat for this. Literally the most important story in tech today. We are at an inflection point. We knew this would happen eventually but itโs here now and itโs both exhilarating and terrifying
https://t.co/J2VXDMrubM