@MiaAI_lab@MiaAI_lab got 43 but needed to make 65k draft-vocab file for it ( cant find yours in repo ;) ) great job. i can feel the speed difference already
@sudoingX hah hello neighbour Spark enthusiast XD happened to me last week in Rayong with slightly bigger rock. let us know how does 3.8 next and 5.3 flies on 2 sparks, i still didn’t both my second unit...
@ModelScope2022 "Before release of Full family of qwen 4 " -thats a strong hint that qwen 4 family of models will still be released as open source/open weights @QwenDevs 🫰
@bullbear_info@MiaAI_lab cap your Mhz settings to 2200-2000 , put spark sideways and invest 5 $ in usb fan. i use 27b with h3 and trellis2 (3d modeling ) and am under 80C ;)
@MiaAI_lab@MiaAI_lab does non thinking 47 tps is worth more then thinking 29 tps? Every time I see comparisons I go with no, but every time I see posts from anyone showing off the potential speeds I see structural. How is your experiance comparing both Mia ?
@Blackwellboy 1/2 there could be other reason (happened to me before) and it's caused by failed update via dgx dashboard. if you ever have similar drop in clock power and hard unplug doesn’t work you will need to load back previous working version and if that fails do it via bios.
@jun_song@MrTacticalX You know what would be cool? 2bit obliterated Super version of dsv4 flash 0731 that could run on one spark with dwarfstar and @MiaAI_lab recipe ;)
@LocalInference@jun_song that’s so bad, was planning to buy second unit and they up the price. i both mine for 116 000 bath in february, wish i took 2 of them at once xD can we have the ram price crash already?
@MinLiBuilds try 0.7 gpu uti with dspark2 recipe from @mr_r0b0t https://t.co/GrxUgOaFYd
it gives 29 tps in prose already and over 30-40 in coding thinking mode. using 4 agents on xhigh thinking with over 60 tps (need to measure again, had over 60 on 3 agents). its doing great work :)
@Vladimir76897@MinLiBuilds steady 25-29 (dflash2) or 24 for sglang in prose. 1000 tps prefill speed. for coding its faster. thinking 35- 50 no think can benchmax to 75. i use thinking mode. dont understand the reason for using this model without thinking from results shown. hope it helped @Vladimir76897
@MiaAI_lab@shuai_bai_@Alibaba_Qwen It would be funnier if we get comparable size to deepseek v4 flash and got possibility to use dwarfstar to run it like we do dsv4 sladh 2bit on one spark