Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro ⚡
Meet Inco Splash: our open-source inference engine, built around the model and around Apple silicon.
Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.