@XiaomiMiMo Ran MiMo-V2.6 Flash through a generic-use battery this morning: constraint writing, summarization, JSON extraction, code, math, tone rewrite, Hindi translation, long-context recall. 8 tasks, 8 usable answers, 2 to 10 seconds each. Findings: https://t.co/DzFTjKJx1Y
Same model drafted five 3,500-word SAP articles this morning at the same speed, so it is not only a short-answer machine. Free tier, generic use: in range.
Took MiMo v2.6 Flash, a free model on the opencode route, through a generic-use battery this morning: strict constraints, summarization, JSON extraction, code, math, tone rewrite, Hindi translation, and long-context recall. All 8 came back usable. Findings:
Weak spot: the Hindi translation is fluent but reaches for loanwords written in Devanagari instead of native terms. Normal in business Hindi, still a watch item.
@mizorewww Ran laya-mlx through real work on an M4 Air this morning. Honest numbers: triage holds up, tagging depends on the input (56 to 69%), blog categorization fails at 47%, prefilter vs hosted saved 77% of calls but confidently-wrong items pass the gate. Notes: https://t.co/OqIs2uHWNb
Spent a morning testing a small open-weights local model (Laya-MLX, fully offline on a MacBook Air, $0 per call) on real tasks: inbox triage, tagging 100 SAP decks, categorizing blog posts, and a prefilter pattern. Honest numbers in this thread.
Spent a morning testing a small open-weights local model (Laya-MLX, fully offline on a MacBook Air, $0 per call) on real tasks: inbox triage, tagging 100 SAP decks, categorizing blog posts, and a prefilter pattern. Honest numbers in this thread.
All measured on an M4 Air, 16GB. Local model: 656MB RAM, tens of ms per decision, $0. Same rules as always: measure on your own data, engineer the input before blaming the model, never trust a confidence gate alone.