VLAct has been accepted to NeurIPS 2026!
Hope to see you in Sydney this December!✈️
One tiny bittersweet note: we received two 6s (the highest reviewer score), but ended up with a poster😭. Still, the stars and the sea lie ahead — keep pushing forward.💪
#NeurIPS
🦾Scaling robot data is essential. But as we built VLAs, we kept asking a complementary question: beyond data scaling, how can we make the backbone learn more transferable knowledge from the same trajectories?🤔
From the StarVLA team, meet VLAct — a VLA-oriented VLM backbone built with representation-centric continued pre-training.
🚀 92.5% on RoboTwin 2.0
🌍 Ahead of all World Action Models on RoboDojo
⚡ Full continued pre-training on 16 GPUs
🔓 Data, code, models & training pipeline fully open
More results & insights in the thread 🧵👇
Everyone is testing Opus 5.5 to create impressive videos. The agent's ability and outcome is even more surprising than using a video‑gen model.
Just as our OneVision‑Encoder was recently accepted by NeurIPS, we would like to produce a promotional video to showcase what we consider the most elegant MLLM visual‑compression solution. Accordingly, we provided Opus with our two papers (OV‑Encoder and OV‑2).
It came back with a 3-minute launch film: every frame rendered from code, every note synthesized, every number traceable to the paper, and even the music is directly from code and with matched beats.
What a surprise to me!!! I was slumped on the sofa as if I had witnessed an atomic bomb explode😜.
The two papers👇
A 3.5-month-old knows that a ball rolling behind a wall is still there.
Object permanence is one of the earliest pillars of biological intelligence: the seed of intuitive physics, mechanical reasoning, and much of what comes after.
Today we're releasing a full-stack data infrastructure to teach it to AI:
🧠 150 cognitive tasks
🎬 A Blender generator for each, scaling to 10K+ diverse samples
📦 A 1.5M-sample training set
🚀 A fine-tuned 16B model that validates the data
Website: https://t.co/cLLKJ22n0c
Paper: https://t.co/QM1J94RvYm
Code: https://t.co/skGEFCDm0l
Training data: https://t.co/jPCoXOU67D
Eval data: https://t.co/KRzqlHXSXX
Model: https://t.co/uxj4vROPyM
Leaderboard: https://t.co/PttRjyPXli
GPT-6 Astra is the best vision model we have tested
I put together our results on detection, segmentation, box prompting, counting, reasoning and video
full post: https://t.co/70tEoHqu0k
↓ examples and trade-offs
Meta’s Segment Anything Model (SAM) 3.1 is now available on Meta Model API, giving developers a fast and lightweight model for detection, segmentation and tracking in a single call on inference tuned for SAM 3.1's architecture.
Use a short phrase to find objects in images and video. One API call returns detections, pixel-precise segmentation masks, and identity-preserving video tracks.
Learn more and start building: https://t.co/AfUFEnU2i7
How a robot arm is controlled, explained for ML people new to robot learning. First of the explainers from my own speedrun.
Next one will be on ACT. Follow me to catch it. https://t.co/g25ayJDVhe
What is the role of academic computer vision research in the age of increasingly powerful large models? Is GPT-6 Astra a step change? How can a researcher have an impact today in academia?
These are the questions I ask myself as I head off to ECCV 2026, a conference I’ve attended since 1992. One of my papers this year is VIGA, a method that takes an image as input and outputs a 3D Blender scene that represents that image. This is a classical inverse-graphics task and VIGA was the first method to solve it using an agentic approach.
The idea is now several years old and the first version of the paper was rejected. This delayed publication significantly. After it was accepted at ECCV, it was quickly surpassed by people using Claude Code for the same purpose. Today GPT-6 Astra blows away all previous results. But we still head off to ECCV to tell the community about our invention that is now fully out of date.
The way academic work often progresses is that one reads recent papers, notices that they have limitations, comes up with a new idea, explores this, publishes it, etc. Any published paper I read today is based on ideas that are at least a year old. And those ideas were based on the literature of the time, which was also a year old. That means that any paper I see at ECCV is likely two years out of date. In AI today, two years means your work is likely irrelevant.
At CVPR this summer I noticed that many authors have not gotten the message. They continue to work on “old” problems that have a long history. This history is based on assumptions about how the “vision problem” will be “solved”. The truth is that it is being solved in a very different way and many of these problems are no longer relevant. Another group of papers focuses on very niche problems where large models likely fail because of insufficient data or lack of business interest. The impactful papers were largely from industry and had long author lists and massive data+compute behind them. These papers were also out of data, describing systems that had been released months before, but at least they served to provide the community with more complete documentation and analysis of commercial systems.
So what should academics do? First, we need to put aside the tools we’ve used for years and start from scratch. Every project should start by trying really hard to solve the problem with existing tools. I would like to see every paper begin with a detailed experimental analysis of how existing models perform and why they fail (if they do). This gives the kind of insight we need today. Then, assuming current models fail, the solution should provide some fundamental insight that will outlive the next release of such models.
Reviewers today still focus on technical novelty. This pushes people to focus on tweaking architectures rather than clearly moving the field forward. Papers need to be judged based on their novel insight and not their novel technical contribution. This is a real shift in thinking but it focuses us on what matters - progress of the field.
If we want there to be a “field” of computer vision, then it can’t become a marginal backwater, focusing on esoteric problems. If you haven’t tried using Astra (or whatever comes next) to solve your problem, then you have not done your homework. This omission should be seen as negatively as not having a previous work section.
Concretely, I think papers should include a new section analogous to “Related Work” where that related work is current models and how they perform on the task. Reviewers should start asking for this and expecting authors to be able to articulate their insights about the limitations of existing large models.
I'm interested in your thoughts.
First person at ECCV — literally. 😎
Walked in at 3 PM sharp and got the first badge. Did I take “first-person intelligence” a little too seriously? 😆
Guess what we’re bringing to the ECCV community this year?👀
Find us at Booth #44! @eccvconf
Wow, GPT-6’s ball tracking is actually that good.😝
Just… 7.8M tokens. Maybe don’t feed every full frame — Mage-VL / LLaVA-OneVision-2 already read video like a codec. This court barely moves.🫣
I used GPT-6 Astra Ultra to track a tennis ball. The results are impressive until I tell you the following.
Ball annotation, using 3 parallel agents took 8m 35s and 17,611 tokens.
Complete workflow: 11m 49s and 30,481 tokens.
The complete workflow also processed 313,914 uncached input tokens and 7,528,320 cached input tokens.
That's a total of 7.87 million total, mostly context reused across frame-inspection calls.
Thanks for sharing. I look forward to continued collaboration in building the community, and I hope to be able to contribute a bit to the physical AI.😼
VLAct: Beyond data scaling
Representation-centric continued pre-training for VLA models. Turns limited robot data into transferable action knowledge. Achieves 92.5% on RoboTwin, #6 on RoboDojo, and beats full-data baselines with 20% data. Fully open-source, 16-GPU training.
🦾Scaling robot data is essential. But as we built VLAs, we kept asking a complementary question: beyond data scaling, how can we make the backbone learn more transferable knowledge from the same trajectories?🤔
From the StarVLA team, meet VLAct — a VLA-oriented VLM backbone built with representation-centric continued pre-training.
🚀 92.5% on RoboTwin 2.0
🌍 Ahead of all World Action Models on RoboDojo
⚡ Full continued pre-training on 16 GPUs
🔓 Data, code, models & training pipeline fully open
More results & insights in the thread 🧵👇
Thrilling to share our latest work VLAct, a VLA-oriented VLM backbone built with representation-centric continued pre-training.
With just 16 GPUs, lab-scale resources, you can build a frontier VLA model from scratch!
Data, code, models and pipeline fully open 🤗
Great team work with @YSenqiao and other starVLA team members!
[8/8] Finally, we want VLAct to be useful beyond this paper.
⚡ Full continued pre-training: 16 GPUs
🔓 Data, models, training & fine-tuning pipeline, evaluation — all open
We hope VLAct can serve as a strong VLM starting point for VLA research.🤗
Page: https://t.co/lDwLE9oHhz
Paper: https://t.co/3vjqO6Zg16
Code: https://t.co/XplyDfdmS6
[7/8] RoboDojo gives us another test of cross-robot adaptation.
VLAct is continued-pretrained on Franka + AgileX, then adapted to ARX X5, a different dual-arm embodiment.
On RoboDojo’s 42 challenging tasks across generalization, precision, long-horizon, memory, and open-ended manipulation, VLAct ranks #6 by Success Rate🚀 and leads all explicitly designated World Action Models.
Another encouraging signal that the backbone transfers beyond the robots seen during continued pre-training.