Today we are introducing Escha-W2 quantization.
A 2-bit Qwen3.6-35B-A3B model built for fast, local inference. The complete model is 12.3GB on disk—small, enough to run on a single consumer GPU — while averaging ~100% of FP8 performance across 12 benchmarks, including:
MMLU-Pro: 80.9
MATH-500: 93.8
GPQA-Diamond: 77.8
LiveCodeBench v6: 62.6
BFCL tool use: 88.9
RULER 8K–128K: 89.9
Commonsense-6: 76.1
On a single RTX 4090, the model runs:
225 tok/s single-stream generation
on 12.3GB on-disk model size
Compressing a 35B MoE model this far without collapsing its capabilities required more than a standard quantization pass.
We built an end-to-end compression system combining state-of-the-art low-bit quantization with model-aware fine tuning and recovery to preserve capabilities most vulnerable to low-bit error.
Quantizing the Qwen 3.6 35B model - from the base model to the final deployable checkpoint — took approximately 10 hours to complete.
Escha-W2 runs through a custom Qwen3-MoE runtime, which includes the weight loader, low-bit decoding kernels and serving integration required to execute the Escha format efficiently. The runtime currently supports SGLang and ZML deployment. vLLM and llama.cpp will be supported in the next release.
No retraining from scratch. No specialized accelerator. One consumer GPU.
Model download: https://t.co/hbO3VuoDGU
Runtime download: https://t.co/hrTpO4FvuA
Apache-2.0 model and runtime.
Sitting here watching this BVB vs Augsburg match and an advertisement for @ForwardMSNFC’s match vs Augsburg popped up on the ad board on the sidelines.
Millions of people around the world just saw Forward Madison’s name on their screens. 2020 is freaking weird
Two sets of twins at Notre Dame — Connor and Ryan Green and Connor and Ryan Powers — discuss their college experience and the impact that being a twin has had on it.
https://t.co/6nDqoSorE9