Latest Exl3 quants on Apple Silicon. Successfully convert Qwen27B model from BF16 to Exl3 on Mac, without touching cuda, quality is now on par with using cuda conversion.
Now everyone can convert Huggingface's Weight to EXL3 format using Apple Silicon. I landed the beta version comfirmed with MiniCPM5-1B, quality on par or above the Exl3 using CUDA. Check https://t.co/Rn3YqtZuvw for v0.2.0!
I ported EXL3 to run well on Apple Silicon. The best quant has arrived in MacOS! - M5 Max can do 2700t/s, 65+t/s on Qwen3.6-35B-A3B and 600 t/s, 38 t/s on Qwen3.6-27B + greedy dflash (20-25tok/s on temp tuned)
https://t.co/1bgN0gFyip
My new chat module landed in oMLX 0.4, support multi-chat, samplings, timeline navigation, branching, model variants, profile, inline perf monitoring and svg rendering right inside web ui.
@jundotkim#omlx#mlx
New open model: MiniMax M3 by @MiniMax_AI is live in the Arena!
Find it across Text, Vision, Document and Code Arena: Frontend. Bring your toughest prompts and vote. Scores incoming soon!
@antirez I'm cooking something amazing. 1000+ Prefill broken in M5 Max 128GB. Gotta refine the edges in a few days. Running ds4-eval to check 15 questions all passed.
@antirez Thanks, I tried it and you know 8Gb for 500k context is a steal, compare to 96Gb for minimax. Now you have M5 Max, maybe you can play with M5 GPU Neural Accelerators⦠this can unlock extra horsepower.