Thrilling to share our latest work VLAct, a VLA-oriented VLM backbone built with representation-centric continued pre-training.
With just 16 GPUs, lab-scale resources, you can build a frontier VLA model from scratch!
Data, code, models and pipeline fully open 🤗
Great team work with @YSenqiao and other starVLA team members!
🦾Scaling robot data is essential. But as we built VLAs, we kept asking a complementary question: beyond data scaling, how can we make the backbone learn more transferable knowledge from the same trajectories?🤔
From the StarVLA team, meet VLAct — a VLA-oriented VLM backbone built with representation-centric continued pre-training.
🚀 92.5% on RoboTwin 2.0
🌍 Ahead of all World Action Models on RoboDojo
⚡ Full continued pre-training on 16 GPUs
🔓 Data, code, models & training pipeline fully open
More results & insights in the thread 🧵👇
VLAct: Beyond data scaling
Representation-centric continued pre-training for VLA models. Turns limited robot data into transferable action knowledge. Achieves 92.5% on RoboTwin, #6 on RoboDojo, and beats full-data baselines with 20% data. Fully open-source, 16-GPU training.
VLAct: Beyond data scaling
Representation-centric continued pre-training for VLA models. Turns limited robot data into transferable action knowledge. Achieves 92.5% on RoboTwin, #6 on RoboDojo, and beats full-data baselines with 20% data. Fully open-source, 16-GPU training.
🦾Scaling robot data is essential. But as we built VLAs, we kept asking a complementary question: beyond data scaling, how can we make the backbone learn more transferable knowledge from the same trajectories?🤔
From the StarVLA team, meet VLAct — a VLA-oriented VLM backbone built with representation-centric continued pre-training.
🚀 92.5% on RoboTwin 2.0
🌍 Ahead of all World Action Models on RoboDojo
⚡ Full continued pre-training on 16 GPUs
🔓 Data, code, models & training pipeline fully open
More results & insights in the thread 🧵👇
🚀 Hy4 preview is here.
770B, 49B active, 1M context.
Built for productivity.
Open source frontier.
Consistent affordable price.
Use it. Tell us what breaks.
More on Hy blog:https://t.co/rbl1IWRk3C
HuggingFace:https://t.co/mE9wevH5XR
Github:https://t.co/pyl9zckpoL
Introducing GEN-1.5, a one-shot learner.
It can learn new tasks in a few seconds. Show it what to do, and it generalizes.
This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
Thoughts About Scaling Law
Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions.
The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed.
Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter.
Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it.
This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count.
Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
Why can Kimi ship K3? Let me tell my story.
Earlier this year, I left academia for industry. I talked to a lot of companies along the way. Here's what I saw:
1��Arrogance. They believe the AI war is over, and they won. No hunger for the future, and no hunger for talent.
2⃣Restlessness. Young labs short on foundation, either rushing to catch the frontier or pivoting away from the competition.
3⃣Fear. Strong teams with real experience, but from the second tier, they can't quite bring themselves to aim for #1.
4⃣Misalignment. Everyone is optimizing for their own credit, but nobody really cares whether the company can reach AGI.
Kimi was different.
Over many conversations with the founders, the same thing came through every time: a raw, genuine hunger for AGI.
I joined. The hunger was real.
We shipped K3. This is only the beginning.
Thrilled to share that our MGM-Omni has been accepted to ECCV 2026! 🎉🚀
As co-author, super excited about this early omni-modal LLM exploration that achieves strong zero-shot voice cloning 🎤✨ and expressive long audio generation 🔊 — and it’s fully open-source! ‼️🔓
#ECCV#AI
MGM-Omni is accepted in ECCV 2026 finally. See you in Sweden!
In this work, we release an omni model for text, image, audio processing and text, audio generation. Moreover, we release a benchmark Long-TTS-Eval for long speech generation evaluation.
Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software.
It’s powered by our newest frontier model, Claude Mythos Preview, which can find software vulnerabilities better than all but the most skilled humans.
https://t.co/NQ7IfEtYk7
I think visual instruction tuning is actually a compromise born from limited compute and data in academia.
In the face of large-scale native multimodal pre-training, vision SFT might indeed become unnecessary.
Kimi K2.5 tech report just dropped!
Quick hits:
- Joint text–vision training: pretrained with 15T vision-text tokens, zero-vision SFT (text-only) to activate visual reasoning
- Agent Swarm + PARL: dynamically orchestrated parallel sub-agents, up to 4.5× lower latency, 78.4% on BrowseComp
- MoonViT-3D: a unified image–video encoder with 4× temporal compression, enabling 4× longer videos in the same context
- Toggle: token-efficient RL, 25–30% fewer tokens with no accuracy drop
Here's our work toward scalable, real-world agentic intelligence. More details in the report 👉https://t.co/N5pwm0M4jm