@TechMDAI 100% the biggest problem we have with glm 5.2 is the speed
It's by far our favorite model but we can only run it around 25-35 tps on the sparks
A glm 5.3 flash would be a complete game changer
SGLang is the first engine to develop and land Breakable CUDA Graph (BCG), the full CUDA Graph, and graph memory reuse.
- BCG drops torch.compile for faster setup and broader compatibility
- Full graph capture brings prefill latency down on dynamic workloads
- Memory reuse keeps the graph footprint fixed as coverage grows.
Read the design details in the blog 👇