Multi-GPU Support + Tensor Parallelism supported!
Prompt processing gain is around 20%, output gain around 5%-6%. Will be optimized further soon.
ROCm support next ๐
https://t.co/83diJSeb8I
@Ipsi_Cust@itsmissglitch abbliteration can be enabled in "Chats" tab under "Experimental speed projection" toggle. It will uncensor any model you run with the engine
To improve Qwen3.8-Flash-Next engine I need:
- A person with AMD GPU + at least 40GB RAM
- A person with dual GPU setup + at least 32GB RAM
All I need is SSH access for a day on each to pin-test the engine and adapt it for dual gpuโs and amd rocm.
Please share, if youโre not the one! ๐ซก
New one-shot voxel pagoda build โ straight from my very first launch with Strata by @coldniko.
Running the IQ3_S GSQ-RCO quant from ISTA-DASLab on a 5080 + 64GB DDR4 and sustaining ~36 tok/s through Hermes.
Absolutely wild performance from hardware that barely even registers as โseriousโ by homelab standards.
Just doubled the prompt processing speeds from 539 token/s to 1290 tokens/s.
Qwen3.8-Flash-Next on consumer GPU is now blazing fast.
Next on the list is further optimizations on kernels and swapping which should boost results by 3X-4X.
https://t.co/83diJSdDja
Run Qwen3.8-Flash-Next-GSQ-RCO-Coder on 32GB of RAM (64GB RAM+VRAM needed for full 265K context)
44 tps output / 1,300 ppts - Enjoy coding fast on consumer hardware!
https://t.co/Aki3ZOCDKX
https://t.co/83diJSeb8I