@needmorevram This thing flies. I’m getting 16/300 decode/prefill on atomic chat q4.25 on full context with vision on llama/5060ti, just ordered another 64 gigs of ram 😭
Just doubled the prompt processing speeds from 539 token/s to 1290 tokens/s.
Qwen3.8-Flash-Next on consumer GPU is now blazing fast.
Next on the list is further optimizations on kernels and swapping which should boost results by 3X-4X.
https://t.co/83diJSdDja
@CryptoHolixxx@DenisMazychev 20-30 years? Absolutely not true. If you didn’t die young, you had good chances to live a long life. But anyway, if you have more options, then not selecting procreation among them is especially dumb.
@kibergnida1337@ItsmeAjayKV не, 3x я наверное преувеличил, но около 30тс на контексте 120к он выдает. Максимальный тоже тянет, но там будут уже твои 15-16. Но это все равно круто, ибо с 27b 120 без MTP - это предел
@ItsmeAjayKV Man, I’m doing context ladder of q4 3.8 flash next on my 5060Ti and so far it is 2x-3x times faster decode than q4 27b, and with much larger context (200k vs 114k already). I don’t think we need 27b anymore, like what’s the point 😂
@needmorevram Running on my 5060Ti much faster than 27B and with much larger context😭 That’s insane. The only thing which is slow is prefill. 200-300 vs 600-800.