@ddalcu Will put on queue. I’m running extensive benchmarking right now on a parallel mtp improvement that shows promise of increasing code generation by 13-19% and prose benefits of 2-23%.
@volatilemarkts I love tensorfold as much as the next dev, but don’t you think this comparison is a tad unfair. Why not show omlx with mtp and apc on vs tensorfold for a fair engine to engine comparison.
@qadir5000@NetworkChuck May I recommend the Bambu X2D. Dual nozzle, $650, 20+ sensors, developer mode… the list goes on and on. I had an agent build a mcp for mine. https://t.co/5PzG7LaZtN
MLX-Serve 26.10.1 is out: speed across the board.
Exactly faster. Qwen3.8 27B with its drafter
+66% on an M5 Ultra,
+28% on an M4 Max,
+37% on an M1 Pro.
Qwen Flash Next:
+11% decode on M4 Max,
+51% prefill on M5 Ultra. (up to ~5400 PP!!!!)
Gemma 4, Qwen, LFM2.5, Spark, Bonsai 2 all faster, output byte-identical to 26.9.6 (across 18 models).
New:
* GGUF models run on our own MLX engine, written in Zig and Metal, instead of the llama.cpp fallback. (experimental)
* 🍣 Sushi-format Flash Next packs run directly.
* Qwen-Image 2.1 edits pictures from instructions.
* Drafting is on by default for every model that supports it.
Huge thanks to everyone who filed, tested and fixed this one: @STRML_, @sbusso, @wutang_superfan , lborloz, LXD-8 @alinselea zeeshanhaque21 @Lojza3D, h9q2cyxvgm-ui @kennethrdegraff, @CowboyCoderHQ , codysk, @Beamsters1, @CerebralCoding_, @AjAbsaki , and more.. THANKS ! Please comment if I did not include your name, for visibility, I know people mostly by their GH handle.
Special thanks to @ashxhart and @Spangler3000 for pushing !
Benchmarked and tested across 5 machines: https://t.co/BDeBKRq6Kd
GH Release Download + Changelog:
https://t.co/Yx9rrU6mcm
Website: https://t.co/qxuPXH75o1
MLX-Serve 26.10.1 is out: speed across the board.
Exactly faster. Qwen3.8 27B with its drafter
+66% on an M5 Ultra,
+28% on an M4 Max,
+37% on an M1 Pro.
Qwen Flash Next:
+11% decode on M4 Max,
+51% prefill on M5 Ultra. (up to ~5400 PP!!!!)
Gemma 4, Qwen, LFM2.5, Spark, Bonsai 2 all faster, output byte-identical to 26.9.6 (across 18 models).
New:
* GGUF models run on our own MLX engine, written in Zig and Metal, instead of the llama.cpp fallback. (experimental)
* 🍣 Sushi-format Flash Next packs run directly.
* Qwen-Image 2.1 edits pictures from instructions.
* Drafting is on by default for every model that supports it.
Huge thanks to everyone who filed, tested and fixed this one: @STRML_, @sbusso, @wutang_superfan , lborloz, LXD-8 @alinselea zeeshanhaque21 @Lojza3D, h9q2cyxvgm-ui @kennethrdegraff, @CowboyCoderHQ , codysk, @Beamsters1, @CerebralCoding_, @AjAbsaki , and more.. THANKS ! Please comment if I did not include your name, for visibility, I know people mostly by their GH handle.
Special thanks to @ashxhart and @Spangler3000 for pushing !
Benchmarked and tested across 5 machines: https://t.co/BDeBKRq6Kd
GH Release Download + Changelog:
https://t.co/Yx9rrU6mcm
Website: https://t.co/qxuPXH75o1
@andyreed Have you considered redirecting all questions to a cross family multi-agent council during your sleeping hours? I implemented it recently and so far experienced excellent results.
Local ai is the way to go, the main downside in my eyes is that you pay every second you aren’t running max tok/sec, because your hardware is an asset, not just your ticket to intelligence, freedom, and privacy. Who else feels my pain?
@sudoingX Local ai is the way to go, the main downside in my eyes is that you pay every second you aren’t running max tok/sec, because your hardware is an asset, not just your ticket to intelligence, freedom, and privacy. Who else feels this pain?