We will give one banked reset for every day you don't have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can.
First one will land in ~ 3 hours. There is still time to create your account if you don't have one.
Qwen3.8-Flash just got whole lot faster now.
MTP has landed on llama.cpp (WIP) !!
And i got it's benefit on my 3090 🥳
@UnslothAI llama.cpp fork got MTP support, and using their MTP shared-draft file i ran Qwen3.8-Flash-Next-IQ3_XXS quant.
Same flags:
-ncmoe 34
-ngl 999,
- kv at q8_0
- draft: mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf
- n-max 3
Results on my 3090 + 64GB RAM:
Decode avg (Without -> MTP)
- 16k: 33 t/s -> 32 t/s
- 32k: 29 t/s -> 39 t/s
- 64k: 23 t/s -> 38 t/s
Peak decode got massive improvement:
- 8K: 37 -> 43 t/s
- 16K: 35 -> 38 t/s
- 32K: 30 -> 42 t/s
- 64K: 24 -> 48 t/s
VRAM:
21.5GB -> 23.3GB (+1.8GB)
At 64k, average decode jumps from 23 - 38t/s, while peak hits 48 t/s
The interesting part is how much t/s increase i got, i'll be testing with different n-max now and will be coming with more results.
That's pretty wild jump for a 3090 !
A little bit of a prefill tradeoff but completely worth it.
Single 3090 + RAM offload. At this point cannot wait for my 5090 to come back 😪
Same hot-expert implementation from @spiritbuun , will try to port a PR to his repo for the ones interested
https://t.co/DaMiW8brm8
/
好評につき第二弾CP実施中🎊
\
Snapdragon® 8 Elite Gen 5 for Galaxy搭載 「Galaxy Z Flip8」を抽選で1名様にプレゼント🎁
【応募方法】
① @Snapdragon_JPNをフォロー
② この投稿をリポスト
さらに「スナドラGalaxy」とリプすると当選確率が2倍にアップ✨
⏰応募締切:9/13(日)23:59まで