This new post from @neuralink combines POYO-style spike tokenisation with a Mamba backbone, akin to our POSSM architecture, and explores self-supervised learning for these models, as we did recently with MOJO. Exciting to see these ideas supporting real-time BCI use! ๐งต
Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts ร 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run: https://t.co/ZSxahzJRju
Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts ร 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run: https://t.co/ZSxahzJRju
@hamishivi Thanks for the comprehensive reply! Value-pretraining seems to be (atleast one of) the best reasons to reconcile this - vapo ablations show it is the single most important part. Would be fun to try out :)
PS: I read sae incorrectly as sao, mb
@hamishivi Also, SAE << GRPO from the plots in your tweet compared to the original SAE paper (Qwen3-30B; see fig below) which beats GRPO by a significant margin on a bunch of tasks (including AIME '25)
These plots suggest GRPO is doing much better than VAPO, but from the VAPO paper the converse is true. How to reconcile this?
Is it just model size (Qwen 2.5 32B in the VAPO paper vs Qwen 3 4b in these plots) or are there other confounders? If the former, a model size ablation would be needed to understand the relative performances of GRPO/VAPO/SAE + * variants?
When working on RL with llms you start realizing that youโre rediscovering a lot of human nature and social constructs from first principles.
Words of encouragement actually matter!
Dinitz-Garg-Goemans conjecture is false. This graph theory problem was open for ~30 years.
The graph below has fractional flow cost 58. Any unsplittable flow (with capacity violation <=15) has cost at least 60.
Chat with GPT 5.6 Pro where this was found: https://t.co/Oi2PQoab2h
One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.
you can now train:
- a 100B-parameter reasoning model
- for 40-turn SWE agent tasks
- in your own coding harness
- for 1000 RL steps
- on just 6 H200 nodes
- in under 2 days
infra co-design is magical
prime-rl can now train 1T parameters MoE blazingly fast, under 5 minutes per step, or 1k steps in ~3 days
To achieve this we shipped in our latest prime-rl 0.6.0:
* inference: wide-ep, fp8 inference, llm-d router, mooncake, kv cache cpu offloading
* training: fsdp2, deep-ep expert parallelism, dsa cp, fp8 training, router replay
* agentic rollout: we rewrote the core of our rollout orchestrator for better scalability
support for glm5, kimi, nemotron, ...,
prime-rl is open source but also end to end optimized to run on our dedicated RL infra and compute layer