@hypergameme To run a crappy local Chinese model
Like qwen 3.8 flash next at mxfp4
Which surpasses 95% human coders
You don't need 100k
6k is enough in today's price
-- buy 2 amd r9700, 64g ddr5, 512g nvme, done.
@mweinbach tcclaviger/Qwen3.8-Flash-Next-MXFP4-FP8-GPTQ · Hugging Face https://t.co/Yo47c09jmn
This one is the cheapest hw setup you can get:
> 100 tps decode for batch 1,
@talphiel Man, I am from communism China.
I am not worshipping anything
You don't need to be religious to know
That human are limited, transient existence--that only make me be more appreciative of this short life.
Why not AMD?
-- a single RTX 6000 (96GB) can get you about 8 AMD Pro R9700 (32 x 8 = 256GB), with aggregated VRAM bandwidth 640x8 = 5TB/s, vs 1.8 TB/s, 766 x 8 = 6 PFLOPS at FP8, vs 2PFLOPS at FP4.
-- physics is physics, man, people get https://t.co/ONutA0t4u3 to run on just 2 R9700, with > 100 tps batch = 1 decoding, > 3000 tps prefill
--
Got this Qwen 3.8 Flash Next running on my dual AMD R9700 machine
peak 21,000 😨😨😨😨tps prefill speed
stead 52 tps decode speed (no MTP, because I only have 32GB memory 😂)
-- beat OPUS 5 on a long context knowledge retrieving task (Opus gave Qwen 89, Qwen gave Opus 64 😂)
-- buddy, if you have 5K, buy this instead of paying $200 a month!
Added 32GB DDR to my machine ($480, 😰😰😰😫) , now the machine has 64GB DDR5, running at 5600 MT per s
Enabled MTP, vision, 4 concurrent request at full 256k context.
23.5k tps peak prefill speed
92 tps decode speed.
Now the whole machine need about $6000 for a full new build, I think 🤔
-- but I almost have my own data center grade server now 😁🤭🤗
@0xSero $5000 setup (dual amd r9700)
Qwen 3.8 flash next at mxfp4
Got prefill peak at 21k tps, decode 52 tps at 128k context
Beat opus 5 on my private benchmark
-- if only I can afford another 64gb ram, I think the decode speed can easily be x3, but that's another $1200 ? 😂