If you have an NVIDIA GPU with 16 GB VRAM, Qwen3.8-27B is about to get a lot more interesting.
And no, weβre not talking about some heavily crippled Q2 GGUF.
A new setup is coming that should let 16 GB NVIDIA cards run:
β’ Qwen3.8-27B
β’ Large context windows
β’ MTP speculative decoding
β’ Linux and Windows
That covers cards like the:
RTX 5060 Ti 16GB
RTX 5070 Ti 16GB
RTX 5080 16GB
and older NVIDIA 16GB cards.
The important part is the combination.
A 27B model + large context normally eats into VRAM pretty quickly. Then you add KV cache, multimodal components, and MTP on top.
Getting all of that working on a 16 GB card without dropping down to Q2 is the interesting bit.
For people sitting on 16 GB GPUs, this could be a pretty big upgrade.
Instead of needing 24 GB or 32 GB just to get a comfortable local setup, the 16 GB tier is about to become much more capable.
And it works on both Linux and Windows.
This is exactly the kind of local AI optimization I like seeing.
Donβt buy a bigger GPU.
Make the software use the GPU you already have.
Coming soon.
@MiaAI_lab Hi, Iβm using Fedora Linux and for now Iβm using Qwen3.8 27B with Q4 and KV-cache and context also quantised on my 5080. The performance is acceptable with 70k tokens for context length. Iβm very interested to know what you are doing.