Kimi and GLM are some of the best open models available. They're also some of the hardest to serve efficiently. Check out this writeup to learn how we're making them faster and cheaper without losing accuracy.
https://t.co/X1ak2QsGeC
This is exactly why we’re going open-weight — only open-weight lets us rally a broader community to build together. 🤗
@ComfyUI community’s quantized builds already support a wide range of GPUs, including the RTX 6000, dual 3090s, 5060, and 4060 Ti.
More GPU models are currently undergoing testing. Feel free to share any benchmark tests under this thread.