Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run: https://t.co/ZSxahzJRju
I asked Astra (GPT-6) and Fable 5.1 both to build GTA VI like game researching about it, and the results are below.
Same prompt. Fable took ~2 hours; Astra took ~90 minutes.
Watch the full gameplay side by side below, with audio switching every 10 seconds