Qwen3.8-27B on a single RTX 4090: 322 tok/s decode.
Same card, same benchmark (Spec-Bench, batch 1, greedy), same session:
vLLM 200, llama.cpp 102.
245 tok/s on the official weights; better 3/4-bit weights do the rest.
Write-up, interactive figures:
https://t.co/YRPcdwwZID
Most late wins were 0.3–1%, smaller than run-to-run noise. Measuring them took forced-text A/Bs (every config decodes the same tokens) and drift-balanced run orders.
An AI coding agent (Claude Code) ran the experiments for four days.
BERT is just a Single Text Diffusion Step! (1/n)
When I first read about language diffusion models, I was surprised to find that their training objective was just a generalization of masked language modeling (MLM), something we’ve been doing since BERT from 2018.
The first thought I had was, “can we finetune a BERT-like model to do text generation?”