The "speedup", is only at large context. they demonstrate lack of loss of performance, not gain of performance.
They developed framework to replace in optimal way subset of attention layers in existing models with linear layers, as this suibset in non-trivial to find, and just slapping linear every say 3 layers instead of attention will degrade performance.
Their paper, “Inverse Scaling in Test-Time Compute,” reveals a surprising phenomenon: in certain tasks, models like Claude and OpenAI's GPT-o series actually perform worse when allowed to "reason" for longer.
Qwen3-Coder is released! This 480B-parameter Mixture-of-Experts model (35B active) natively supports 256K context and scales to 1M context with extrapolation