All my feed is people posting quant version of glm-5.3-flash (mostly FP4). Little do they know, they can run these models without compression, at same benefits of memory saving.
Two models are set to be released today that should make every DGX Spark owner very happy
- GLM 5.3 Flash
- Qwen3.8 Flash
Qwen3.8 Flash should run on a single Spark using NVFP4. GLM 5.3 Flash will probably require two DGX Sparks.
Huge day!