@leerob Thx... It was draining really fast, faster then normal. Grok 4.5 should be computed as First-party model usage now after the acquisition, right?
@superbiche@sudoingX 3090 can do near 260 tps gen with 35B A3B with vLLM, Intel AutoRound 4-bit, fp8 kv-cache without any practical (<1%) accuracy loss
@Psicolicious@UnslothAI FP8 kv cache in vLLM, disable video and tweak image settings (lower resolution supported). Can do 262k context with text-only easily. Lost only 1% (mean) in all benchmarks with this amazing setup.
@danieltvela @The_Only_Signal It really is... but then what would we call Qwen3.6 35B A3B (called Qwen3.6 Flash by Alibaba) running at 160tps on a single rtx 3090 with 262k context in vLLM vs 37 tps 27B then?
@CarlosZarattini Deixa ver se eu entendi sua posição: então o dado que desmonta essa mentira de vez diz que menos de 3% dos beneficiários trabalham de carteira assinada... Puxa, 3% é realmente um desmonte...