- AMD is now +20tps faster
- Unsloth’s Q4_K_XL support
- Dual AMD GPU support
- CUDA 7% faster
- RTX 20 is 12% faster
… and 30 things more
https://t.co/qtnyGpDUW3
Enjoy! 👌
Run Qwen3.8-Flash-next on 8GB+ AMD GPU 🔴
RX 7900 XTX runs the model at a sustained 52-60 output tokens per second + 1250 prompt processing tokens per second.
+ Enable KV Cache - k8v4 for + 10% boost in performance with no quality loss
https://t.co/83diJSdDja
Qwen-Image-2.1 full BF16 running with 2GB VRAM
ncnn/Vulkan for NVIDIA/AMD/Intel/Apple
No CUDA/PyTorch/Python
Portable executable
Text-to-image & Image editing
Up to 10 reference images
Transparent RGBA generation
Dynamic output resolution
Batch generation
https://t.co/mTiok6Qija