I couldn’t find any public doc confirming that fast uses a separate serving path with different routing, scheduling, or inference optimizations. Fireworks also offers dedicated deployments, an 8× B300 instance costs $96/h. So the premium could reflect real infrastructure and capacity costs, along with some margin. I might try deploying it myself tomorrow.
Kimi K3 is only one day into its open-source release, and providers like Baseten, Nebius Token Factory, Fireworks, DigitalOcean, and Together are already hosting it on OpenRouter. Aside from Fireworks’ Fast tier, which carries a 50% premium, the other providers are charging the same rates as Kimi’s official API.
The model weighs in at roughly 1.56 TB, has 2.8T total parameters, uses MXFP4, and activates around 104B parameters per token. Based on publicly validated deployments, the minimum setup appears to be 8× AMD MI355X or 8× B300 GPUs. For individuals, private deployment is basically out of the question. And that’s only the bare minimum. Kimi’s technical documentation recommends 64+ accelerators connected within the same high-speed interconnect domain for production-grade deployment.