Bypassed the base M4 memory wall on Qwen2.5-Coder-32B: Skipped 20 layers, jumped from 6 to 8.2 tok/s with zero quality drop.
I’m running Qwen2.5-Coder-32B (3-bit quantised with MLX) on a stock base-model M4 iMac with 24 GB RAM.
For the kind of stuff I actually use it for—heavy refactoring of ancient codebases, complex math, and tough logic problems—local is awesome… until it’s slow. And let’s be real: 6 tokens per second (the usual community number on this hardware) feels completely unusable when you’re trying to do any kind of agentic loop or back-and-forth coding session.
A ~12 GB 3-bit 32B model on the M4’s 120 GB/s memory bandwidth should theoretically top out around 10 tok/s. You literally can’t pull the weights from RAM to the cores any faster than that.
So instead of just grinding away with speculative decoding (which I’m still tweaking because the MLX verifier overhead is annoying), I went a different route: custom static layer skipping.
I spent a bunch of time profiling which transformer blocks actually do the least work during generation, then built a router that just permanently drops the least useful ones. It completely sidesteps the memory bandwidth wall.
Yeah, I know—layer pruning can turn models into lobotomised 2+2=potato if you’re not careful. But the results here are genuinely great.
I’m now seeing 9.05 tok/s in the absolute fastest config, and a very usable 8.20 tok/s in the “I actually trust this for real work” config.
The 8.20 tok/s version is the one that matters: on strict math and coding benchmarks, the output is indistinguishable from the full dense model. Same reasoning, no extra hallucinations, no quality drop. Just pure 32B intelligence running right up against the physical limits of the silicon.
For a regular all-in-one desktop with only 120 GB/s bandwidth, hitting 80–90% of that ceiling and turning a sluggish 6 tok/s experience into a snappy 8+ tok/s coding partner feels borderline magical. Obviously it’s not M4 Max or dual-4090 territory, but for the rest of us stuck with the base chip, this is a game-changer.
Anyways the job ain't done yet, but this is already proving some real life benefits isn't it?
#localllm #openclaw #AI #Qwen
@daumenxyz Good products remove friction
Better ones make progress feel natural
Momo feels like both
Build easily like never before
Play store coming soon!
https://t.co/n0aPRaV7T3
@MrWhale Send $Momo so much higher
Utility peaked with this
there’s something clean about a tool that helps without making everything complicated
Momo keeps it simple
@Momo_Agent_
MOMO App will soon be available on the Google Play Store. ⏳
We are finalizing the last steps to make the app accessible to everyone.
Stay tuned for the official release announcement.
What's next for Momo. A quick look at where we are and where we're going.
Already live on Android:
Your AI reads your emails overnight, sorts by priority, and gives you a clean summary when you wake up. Need to reply? Momo drafts it. You approve and send.
All your messengers. WhatsApp, Telegram, Discord. Already connected. One AI assistant that manages every conversation in one place.
Voice messages, app builder, life dashboard. All working.
What's coming next:
→ iOS app
→ Company formation in Germany. Building this for the long term
→ More integrations, smarter proactive assistant
The vision: one AI that connects your entire digital life so you stop juggling 12 tools.
We're just getting started → https://t.co/Ei6B48fl5d
@GodsBurnt Yo guys check out $Momo
Based dev, strong and active community
Build, set reminders, solve complex problems among much more
Millions primed
@Momo_Agent_