@DogukanUrker Than you for your hard work. From my testing in my use cases qwen3.6-35B-A3B-UD-Q4_K_XL-MTP with 3060 + offload on cpu is still better choice than any _XS or _S qwen 3.8 27b I have tested so far.
2x faster inference on my hw and produced html toy-tasks seems better on old 35b ๐
@DogukanUrker MTP is with spec-draft-n-max = 2 ?
on Fedora with 3060 12g i can only use:
spec-draft-n-max = 2 with 85k ctx, otherwise llama-server cant allocate more memory, with 32 tps.
With llama branch for DFlash2 UD-IQ2_XXS.gguf ~37 tps (n-max = 4)
@DogukanUrker I have the same setup: 12GB RTX 3060. Is in your case Qwen3.8-27B overthinking? What thinking level/effort have you set up if any? I'm using pi-agent and it seems like setting up an thinking effort is not working there.
Are you using default xhigh?