Really interested to see how this ~74b MoE lands once RL is completed, as it seems like ~30b is the current local hotspot. Moreover, I'd presume they are going to convert to TiDAR for specdec speedup. Could be interesting!
@DevaBuilds@Sapient_Int Naive goal but was trying to see if I could get the base model + HRM architecture to push into to Bonsai / LFM / Gemma territory in terms of performance (based on the simple benchmarks I did on them). I wasn't eval'ing on latency, purely curious about accuracy / answer quality.
I messed around with fine tuning @Sapient_Int 's HRM 1B model. I find HRM fascinating! Goal was to learn how fine tuning works and see if I could modestly improve off the base model.
Note the benchmarking is crude as I'm basing it off my workflows and severely hardware limited.
Really impressed with @opencode / @MiniMax_AI M3 combo. Used it for the first time and it pulled Nvidia Diffusion 3B 4 bit MLX out of a ditched by writing a custom MLX loader, ran it on my tiny M2 mini 8gb, and came up with some great insights.
The market's giving MiniMax-M3 a vote of no confidence, as shares tumble down a massive 15.71% after running up into the M3 release.
Hearing from more and more people that this feels like a benchmaxxed model.
Again, hope to be proven wrong. But this seems quite bad.
@PavloMolchanov My apologies! I meant does it *only* apply to smaller deployments w/ less concurrency (and more free cores). But I think you answered both!
@OnlyTerp Currently trying to have Nemo 3 Ultra (Hermes harness) simply update a blog post, write a new blog post, and upload to my website and it is looping badly with a failed tool calls. DSv4 flash / Mimo 2.7 etc handled this repeatable task no problem.
@OnlyTerp Currently trying to have Nemo 3 Ultra (Hermes harness) simply update a blog post, write a new blog post, and upload to my website and it is looping badly with a failed tool calls. DSv4 flash / Mimo 2.7 etc handled this repeatable task no problem.