@Luisezpeleta@needmorevram I installed Strata last night and have been playing with it today on my 3060 and 3090 GPU (two different systems). I'm seeing remarkable number too all running the Q3_XSS quant of Qwen3.8-Flash-Next.
30+ t/s on a 12GB 3060 using another 60GB of DDR4 system ram.
@itsmissglitch@coldniko You could say I was gobsmacked when I typed in my first prompt and saw the response come back at 30+ t/s. I was in disbelief and expected as you did that this project would amount to the same useless hype that others at let me down with. As long as the 3_xss quant is good...
@itsmissglitch@coldniko Wow, I have an RTX3060 with 12GB VRAM and a Ryzen 9 3900X with 64GB RAM, and with Strata I'm running Qwen3.8-Flash-Next Q3_XSS quant at at 32K context at over 30 t/s!
Another reminder that Q1 GGUFs are too often just headline bait.
“We can run Qwen3.8 27B on an 8 GB GPU.”
Sure.
But if the quantization drops it to the accuracy of a nearly two-year-old 4B model, you’re not really getting what you think you are.
It's not always the case. Q1 for very large models can be good enough. But for small models (let's say below 100B), I've never seen Q1 being useful.