Based on the likely 4-bit quant sizes for
@Alibaba_Qwen
Qwen 3.8-Flash-Next (~95-105GB depending on quant recipe), I'm not convinced this model was made for consumer-friendly hardware the way the 35B version was. Even on a 128GB Spark or Ultra you're left with very little headroom for any serious workload.
@ivanfioravanti Do you know if they provide the config of the benchmark tests? I am a little suspicious of the x9.8 figure as Lmstudio is not the most optimized environment for mlx.
📢Meet Qwen3.8-Max — our most capable model to date.
Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉
Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:
- Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:https://t.co/iVHZWQoeSo
- Real work, real results: Production-quality deliverables across hundreds of professions.
- Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy.
- Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction.
💰Pricing:
Input: $2.0 / M tokens
Output: $6.0 / M tokens
Implicit Caching: $0.25 / M tokens
Start building with Qwen3.8-Max! 🚀
📖 Blog: https://t.co/iwjmQxLBof
✅ Qwen Studio: https://t.co/4V2pFvDovG
⚡ API: https://t.co/gAGqaLQGbN
Everyone knows I'm super AI-pilled. But one trend I'm noticing as I talk to more and more companies: the personal productivity gains are not translating into organizational growth and efficiency as expected.
It's an odd dichotomy. Individuals are more productive. Engineers are writing an insane amount of code. But the gains aren't showing up in the numbers yet.
Not sure how universal this is.
First attempt at attaching MTP tensors to Ornith-1.0-35B (@ornith_) for Apple Silicon, based on Qwopus3.6-35B-A3B 8-bit.
It's not tuned on Ornith's hidden states, but still produces a 20–30% speedup over the original.
https://t.co/d39mTyK3ys
Ornith-1.0 35B by @ornith_ on MLX (using omlx) is very promising so far. Compared to 3.6 I see better prefill, almost the same inference speed even without MTP, stable tool use, and responses on Hermes feel sharper. Probably too early to tell without real-world benchmarks though.
Meet Gemma 4 12B!
A unified, encoder-free multimodal model designed to bring high-performance intelligence directly to your laptop, and released under an Apache 2.0 license.
Bridging the gap between edge efficiency and advanced reasoning. Here is what’s new with Gemma 4 12B: 👇