> get super stoked
> 328 tps
> wow that's sick
> sign up for account
> run quick benchmark bc I'm not a trusting person
> it's 58 tps the same as everyone else
thanks for playing
@DotDotJames@UnslothAI which inference engine and command line you using? i tried days ago but quality was garbage and slower than my current setting "Gemma-4-26B-A4B-it UD-Q4_K_XL (187 t/s, 530 t/s aggregate in 4 slots parallel)
Kimi 2.7 ranked 2nd after Fable 5 and before GPT-5 xhigh
We have re-run our ErdosBench smoke test on 14 problems with Kimi 2.7, Qwen 3.7 Max, Grok 4.3 and compared it with the top performers from previous runs.
Kimi 2.7 is amazingly good. More below.
@0xSero you sure Diffusion Gemma26B really works? tried it on 5090 and it is slower compared to autoregression for normal prompts (variable input, variable output)
@bendee983 I am same token processing tier with opus 4.6. cache saves 95% context otherwise we would be talking about milions/month.
But inference costs will most likely and hopefully go down in next quarters, this Is the way, let s just wait for it