Excited to share our new preprint on VoxServe, a serving system for Speech Language Models 🎙️⚡️
Speech model serving is challenging: complex pipelines, diverse architectures, and strict real-time requirements for low-latency + streamable inference.
(1/2)