🚀 Rapid-MLX 0.14.2 is live—our biggest local-AI performance leap yet.
⚡️ The Speed Run:
• Qwen3.6 35B: +57% throughput (native MTP)
• GLM-5.3 Flash: +34% faster speculative decoding
• DeepSeek V4.1: 2x faster decoding, 28% less peak RAM
• Code edits: Up to 2x faster via context reuse
🏔 Spotlight: K2 Horizon 7B
We now support IFM's new 7B model from @mbzuai. It's a compact agentic powerhouse featuring a massive 512K context window and scoring a crushing 68.4% on SWE-bench Verified.
The best part? It runs at ~50 tok/s on an M4 Pro.
🤖 The Smarts:
• NEW: Experimental Local Agent Mode (MCP tools + approval gates)
• Zero-prefill web searches
• Auto-recovery for malformed tool calls
100% local on Apple Silicon.
👇 Download: https://t.co/RMVDzn2Agg
🚀 Rapid-MLX v0.14.3 is live!
The star: Bonsai 2 Hadamard 27B. A multimodal 27B model packed into just ~8GB, blazing at 34 tok/s on an M4 Max. Full support across CLI, Server, and Desktop for vision, reasoning, and tool calling.
Try it: rapid-mlx serve bonsai2-27b-2bit
Huge shoutout to @Mieluoxxx (Morgan Woods) for the original Bonsai 2 loader contribution! 🙌
On top of that ⚡️ Multimodal conversations are now much faster:
• 54.8% lower 2nd-turn media TTFT
• 35.4% higher throughput
• seed, top_k, and min_p support added
Everything stays local on your Mac https://t.co/rmBN6TJoxj
@official_Gegeh I wonder how some people on social media think.
How can gehgeh get married on a low? He's not a coward— the last time I checked.
See his face seff. Does he come off as someone who's "happily married?"
Wrap it up, mbok.
@femibond007@SavvyRinu I think you need time to think things throughout before talking
The first lady's statement is not timely, and it's very insensitive.
Periodtttt.
@NigeriaStories This woman has become something else. Her comment about selling akara and all that was so insensitive.
Yes, she had a point. But question is: how well did she communicate her mind?