How to define Diversity in the context of CodeLMs and Programming Languages ?
1. Diversity is positively correlated with Performance in solving a problem.
2. Shortcomings of diversity in small codeLMs.
3. Code Embedding models don't capture semantics.
https://t.co/aNGhSpxz48
At humans&, we train models from the long-term impacts of their interactions with people. This requires prioritizing long-horizon multi-agent RL. We've developed and are excited to share an open-source, hardware-native 4-bit RL recipe, significantly accelerating training
My June essay is out now!π
It is titled: Itβs time to do the harder things
I discuss how media companies like the Discovery Channel once captured beautiful wildlife shots but now create unscripted reality TV shows like Deadliest Catch because it is cheaper and faster to produce
But with AI, I argue that now it is time for us to go back to doing the harder things
Weight sync in RL post-training looks RDMA-bound. But >99% of BF16 weights are byte-identical step to step β ship only the delta: ~100β1000Γ less traffic, bit-exact. Hardware bottleneck β software problem: RL over commodity ethernet.
π§΅ https://t.co/ivniFyzPo1
These are small scale GRPO runs with Qwen-1.7B-Base on MATH task wherein I consistently observed optimizers with no velocity like Muon, Signum work comparably. Is this something studied in scale ? Are there public info on such large scale runs ?
We're excited to share Stable-Layers!
We train Qwen-Image-Layered further with RL for improved layerization,
using only feedback from a VLM β no paired supervision required!
Paper: https://t.co/WktmmXGNLh
Project Page: https://t.co/WnEV76afQp
π¬ Introducing Stable Cinemetrics, to be presented at NeurIPS 2025.
We present the first taxonomy of professional controls to systematically study and control video generative models through the lens of filmmaking.
Interactive webpage with paper link: https://t.co/Eh4Hw3hBZl
π§΅
πββοΈ Can RL training address model weaknesses without external distillation?
π Please check our latest work on RL for LLM reasoning!
π― TL;DR: We propose augmenting RL training with synthetic problems targeting modelβs reasoning weaknesses.
πQwen2.5-32B: 42.9 β SwS-32B: 68.4