@JamesMcPherson@che1i0s@dlenrow@Apple There is a difference between hardware limits and software limits you may not be fully appreciating. Just look at the progress MLX serve is making at concurrent sessions. As of now there might not be a difference, but that is primarily due to unoptimized software.
@ViC305@justinn_builds MLX-serve is rapidly pushing the concurrency frontier on Apple silicon. In the next few days we should have qwen 3.8 flash next at concurrency 3 and with 150 tok/sec aggregate.
@ddalcu Will put on queue. I’m running extensive benchmarking right now on a parallel mtp improvement that shows promise of increasing code generation by 13-19% and prose benefits of 2-23%.
@volatilemarkts I love tensorfold as much as the next dev, but don’t you think this comparison is a tad unfair. Why not show omlx with mtp and apc on vs tensorfold for a fair engine to engine comparison.
@qadir5000@NetworkChuck May I recommend the Bambu X2D. Dual nozzle, $650, 20+ sensors, developer mode… the list goes on and on. I had an agent build a mcp for mine. https://t.co/5PzG7LaZtN
MLX-Serve 26.10.1 is out: speed across the board.
Exactly faster. Qwen3.8 27B with its drafter
+66% on an M5 Ultra,
+28% on an M4 Max,
+37% on an M1 Pro.
Qwen Flash Next:
+11% decode on M4 Max,
+51% prefill on M5 Ultra. (up to ~5400 PP!!!!)
Gemma 4, Qwen, LFM2.5, Spark, Bonsai 2 all faster, output byte-identical to 26.9.6 (across 18 models).
New:
* GGUF models run on our own MLX engine, written in Zig and Metal, instead of the llama.cpp fallback. (experimental)
* 🍣 Sushi-format Flash Next packs run directly.
* Qwen-Image 2.1 edits pictures from instructions.
* Drafting is on by default for every model that supports it.
Huge thanks to everyone who filed, tested and fixed this one: @STRML_, @sbusso, @wutang_superfan , lborloz, LXD-8 @alinselea zeeshanhaque21 @Lojza3D, h9q2cyxvgm-ui @kennethrdegraff, @CowboyCoderHQ , codysk, @Beamsters1, @CerebralCoding_, @AjAbsaki , and more.. THANKS ! Please comment if I did not include your name, for visibility, I know people mostly by their GH handle.
Special thanks to @ashxhart and @Spangler3000 for pushing !
Benchmarked and tested across 5 machines: https://t.co/BDeBKRq6Kd
GH Release Download + Changelog:
https://t.co/Yx9rrU6mcm
Website: https://t.co/qxuPXH75o1
MLX-Serve 26.10.1 is out: speed across the board.
Exactly faster. Qwen3.8 27B with its drafter
+66% on an M5 Ultra,
+28% on an M4 Max,
+37% on an M1 Pro.
Qwen Flash Next:
+11% decode on M4 Max,
+51% prefill on M5 Ultra. (up to ~5400 PP!!!!)
Gemma 4, Qwen, LFM2.5, Spark, Bonsai 2 all faster, output byte-identical to 26.9.6 (across 18 models).
New:
* GGUF models run on our own MLX engine, written in Zig and Metal, instead of the llama.cpp fallback. (experimental)
* 🍣 Sushi-format Flash Next packs run directly.
* Qwen-Image 2.1 edits pictures from instructions.
* Drafting is on by default for every model that supports it.
Huge thanks to everyone who filed, tested and fixed this one: @STRML_, @sbusso, @wutang_superfan , lborloz, LXD-8 @alinselea zeeshanhaque21 @Lojza3D, h9q2cyxvgm-ui @kennethrdegraff, @CowboyCoderHQ , codysk, @Beamsters1, @CerebralCoding_, @AjAbsaki , and more.. THANKS ! Please comment if I did not include your name, for visibility, I know people mostly by their GH handle.
Special thanks to @ashxhart and @Spangler3000 for pushing !
Benchmarked and tested across 5 machines: https://t.co/BDeBKRq6Kd
GH Release Download + Changelog:
https://t.co/Yx9rrU6mcm
Website: https://t.co/qxuPXH75o1
@andyreed Have you considered redirecting all questions to a cross family multi-agent council during your sleeping hours? I implemented it recently and so far experienced excellent results.