Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run: https://t.co/ZSxahzJRju
Self-improving agents are a top research topic right now.
This new survey is a good map of the area.
(bookmark it)
It splits recursive self-improvement into stages of autonomy.
An agent first executes improvements someone else designed. Then it chooses its own improvement strategy, collects its own experience, adapts to new environments, and finally improves the process of improvement itself.
That staging makes claims easier to check. When a paper says its agent is self-improving, you can ask which of these stages it actually automates.
The survey also uses a Headroom-Closed Index to show where current LLMs fall short, and compares requirements across scientific discovery, embodied agents and software engineering.
Paper: https://t.co/93OctNiROf
What would truly open-source AI look like? Not just open weights, open code/data, but *open development*, where the entire research and development process is public *and* anyone can contribute. We built Marin, an open lab, to fulfill this vision:
‼️huge ssi news.
ilya is about to take his first tentative steps out of the age of research and back into the age of scale. it’s time to smell what ssi is cooking.
ssi have built a small reasoning engine that can compete with much larger training runs because his data is better curated to meta learning. but, more importantly.
we’re about to step into the era of TTT (test time training. gradient descent happening in real time to solve your problems). so instead of a context window you get actual learning.
and because it’s so sample efficient it can be trained on hard to verify tasks that other paradigms can’t touch. everyone else’s weights are frozen, they struggle out of distribution. ssi have created something that has a bundle of knowledge but can truly learn in real time and use that to your advantage.
current approaches are trying to hack their way to ‘learn’ with memory tricks, this thing will updates its weights, remember key lessons, and finally feel like a human level reasoner. this is a huge paradigm shift from the king. early results are very impressive. we can stop watching memento on repeat.
it’s learning all the way down, the descent is real.
Dreamina Seedance 2.5 is now live!
From creators to enterprises, a new era of AI video creation begins.
Try Seedance 2.5 on Dreamina today. Enterprise API access via BytePlus is coming soon.
Create longer videos with greater control:
- Native 30-second generation
- Precise video editing
- Support for up to 50 multimodal references
- Multilingual video generation
More consistency. More control. More possibilities.
🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta!
🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇
🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex!
Check out the configuration details in our official API docs: https://t.co/smCwQZMeiq
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
Macaron-V1 is now available on ModelScope!
A step forward toward more capable AI agents —
with stronger coding, reasoning, and long-horizon workflows. 🚀