Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run: https://t.co/ZSxahzJRju
At this point DwarfStar contains many fast fused kernels for important model families: feel free to steal everything you want from there according to the MIT license, in order to improve your own implementation.
I wanted my Mac Studio and two Sparks to work together, so I built a macOS RDMA driver.
Metal and CUDA shared-buffer transfers now pass in both directions.
Here’s why I built MCDMA, what works today, and the 25 experiments I want to try next. https://t.co/gCRYWdioOY
Anyone with 64GB Ram want to test Qwen Flash Next on MLX-Serve ? should work with existing release, but next one will be faster. https://t.co/sAKK5jnoir