sth new | Postdoc Princeton | Incoming Assistant Professor Columbia | Prev Bytedance Seed, Tsinghua, NVIDIA, IDEA Research, Microsoft | Views are my own.
When somebody told you, “We are only second to Fable 5,” we can expect he was just being modest.
Qwen 3.8: 2.7T parameters.
Kimi K3: 2.8T parameters.
It turns out that scaling laws are still very much alive and working. The era of trillion-parameter models may have only just begun.
Qwen3.8 is launching and going open-weight soon!🌐
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.
Can't wait to hear what you build. Stay tuned! 🚀
Token Plan
international:https://t.co/YRvcGdB9Bv
China:https://t.co/PKMUNwUuRp
Thanks for sharing. This is very interesting.
One useful way to think about it is to zoom out from yoyo itself and look at the downstream task it is ultimately serving.
If yoyo is eventually used to create software for users, then the user-facing software is the artifact. In that view, yoyo’s self-written codebase is more like the evolving harness that improves its future ability to solve such tasks. If those files are the final deliverables for users, then this would look like artifact evolution.
So I think the classification depends less on whether files are generated, and more on what role those files play in the larger system.
This is also where the boundary becomes blurry, as I mentioned in the blog. In many current systems, improving artifacts can also lead to changes in the harness itself.
Self-evolving agents have recently attracted a lot of attention.
However, when people talk about “self-evolving” systems, are they always referring to the same thing? Many works describe themselves as self-evolving, self-improving, continually learning, or capable of test-time adaptation. But what exactly are they improving, and how are these ideas related?
In this blog post, we try to provide a simple taxonomy of current self-evolving systems.
We view such systems as consisting of three key components: the model, the harness, and the artifacts produced by the agent. This naturally leads to three types of self-evolving work: systems that optimize artifacts, systems that improve the agent harness, and systems that update the model itself.
For each category, we summarize its motivation, typical methods, and emerging trends. We then take a step back and discuss how these directions are connected, where their boundaries become blurry, and what this may suggest about the future development of self-evolving agents.
Another interactive model. Full-duplex interaction is the future.
It is useful to read GPT-Live’s blog together with Thinking Machines’ post on interaction models:
https://t.co/pQngUkIYlq
They seem to point to the same direction: moving beyond turn-based chat, and building fast and slow systems together. Smaller models handle real-time interaction. Stronger models handle deeper reasoning.
Thinking Machines’ post goes into more technical details on how this can be done.
Introducing GPT-Live, a new generation of voice models for natural human-AI interaction.
Rolling out in ChatGPT starting today.
You’ll want to turn the sound on for this one.
Many people have asked me about academia vs. industry, especially as I’ve seen many people move from universities to companies.
I like this analogy:
Academia should be like special forces: exploring uncertain territory and finding new directions.
Once a direction becomes clear, like LLMs today, industry should be the regular army that pushes it forward at scale.
We’ve received notice that the Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5.
We'll begin restoring access tomorrow, and will share an update soon.
We’re grateful to our users for their patience, and to everyone who worked with us on redeploying the models.