We swapped a 122B for a 27B in our recommended models.
Qwen3.8-27B scores 52 on Artificial Analysis - the 122B it replaced scored 33.
It peaks at 20GB, so every Mac with 32GB or more gets the same pick. It's the best-scoring open-weights model we serve.
$ rapid-mlx serve qwen3.8-27b-4bit
Apache 2.0. Fully local.
Open models are getting powerful enough to become truly personal.
Models like DeepSeek, Qwen, and Llama are bringing frontier-level intelligence closer to everyone.
The next challenge:
Making that intelligence fast, private, and always available on your own devices.
That’s the future of local AI.
Meet the fresh new look of Rapid-MLX ✨
Your personal, private AI running seamlessly on Apple Silicon. Get chat, coding agents, and voice directly on your Mac. No cloud, no data sharing—just pure local power. 🚀
Join our growing community on X and Discord to ask questions and shape what's next! 👇
🌐 https://t.co/aCMi87z9db
💬 https://t.co/pbSCQ9YSkj
🚀 Rapid-MLX 0.12.18 just dropped.
Check out this video—generated 100% locally on a Mac using LTX-2.5 with synchronized audio. It’s the ultimate local video model for Apple silicon.
We also packed this release with serious speedups (M3 Ultra):
• Qwen3.6-35B MoE: +5.8% decode (86.1 tok/s)
• GDN prefill: up to 2.13x faster
• Transcription: 30s of dictation in <0.5s!
• 215 model aliases + Experimental MTP validated on 18 Qwen configs
Everything runs locally on Apple silicon 👇
Make it. Edit it. Own it. 🎨
Create and reimagine stunning images, then turn them into video. The best part? It runs completely locally on your Mac. Your data never leaves your machine.
"This couch needs a cat." Done, right on-device. 🛋️ 😼
🖼️ Images Models: FLUX · Z-Image · DiffusionGemma · Qwen-image
🎬 Video Models: LTX · Wan · CogVideoX
Meet Rapid-MLX: The ultimate open-source local AI ecosystem built exclusively for Apple Silicon. 🚀
We've been building quietly, and today we're rolling out v0.12.16. Chat, coding agents, voice, and images—everything runs natively on your Mac. Absolute privacy. Nothing leaves it.
🎙️ System-Wide Private Dictation
Tap the right Option key anywhere (Mail, Slack, VS Code). Speak, and your words land instantly at your cursor. Transcription runs entirely on your local server. Your voice never touches the cloud. It works seamlessly in the background, even with the app window closed.
🧠 Hardware-Aware Intelligence
Rapid-MLX automatically matches your Mac with the smartest model it can run. Rocking a 32GB+ Mac? You get Qwen3.8-27B (GPT-5.6-class intelligence, Artificial Analysis index 52) pushing a blazing ~40 tok/s. ⚡️
🛠️ Built for Power Users
• Built-in web search (Zero API keys)
• Native math rendering
• Chat folders & Markdown export
Free. Open source. Powerful AI that belongs directly in your hands.
Terminal: curl -fsSL https://t.co/5i5Bz4DybP | bash
Mac App: https://t.co/rmBN6TJoxj
rapid-mlx 0.12.15 is out 🐆
⚡️ We just shipped lossless MTP speculative decoding for Qwen3.8-27B, 1.25x faster! Just one flag: rapid-mlx serve qwen3.8-27b-mixed-3.5bpw --speculative-config '{"method":"mtp"}'
Other major updates:
🛠️ Coding agents fix: If you use Claude Code, Aider, etc., upgrade now. v0.12.5–0.12.14 could silently shorten what your agent wrote to files. Fixed. (Worth a re-check if your agent edited anything important recently).
🚀 Blazing downloads: Grab a model at 80–90 MB/s. A 9B is ready to chat in ~2 mins, with auto-rerouting if it slows down.
🎙️ Better voice: Whisper no longer makes up words when you pause or the room is quiet.
🛡️ Safer testing: Model downloads can't run code on your machine anymore, and every file is integrity-checked.
Full changelog: https://t.co/DE5FDIaE6n