ox-Alpha is an exceptional example of why Chinese labs are so impressive with tiny budgets.
Looks how fast ZAI adopted innovations by their peers!!!
1.) They adopted DeepSeek-V4's mHC based residual connection and
2.) also, DeepSeek's DSA like Sparse Attention
3.) also, MoonShot's KDA type linear attention
Result:
1.) Strong performance with 4.44 times less KV cache
2.) 3 times less flops
3.) Kick ass model that can be served 10 times cheaply even compared to their own GLM-5.3
Every lab builds on the innovations by other, so they don't have to repeat all the same experiments themselves. This is extremely economically efficient.
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog: https://t.co/tzOmB7gdZP
Available now across all official platforms:
Weights: https://t.co/9LRMahY9Wa
API: https://t.co/VcaQnzYmS9
Coding Plan: https://t.co/Nk8Y98HNhU
ZCode: https://t.co/Peepqv4XSx
Chat: https://t.co/WCqWT0qCQb
AutoClaw: https://t.co/aGEG5HqTTb