🚨 What makes good world-model training for MLLMs that actually transfers across tasks?
🧶 Excited to share our new work, WOVEN!
🧩 Spatial, physical, temporal, and embodied reasoning all require understanding how the world changes.
🔁 We view world modeling through state transitions p(s′ | s, a), and study visual transition reasoning: inferring what is unobserved in (s, a, s′).
🤔 Can visual transition reasoning be learned as a shared primitive that benefits multiple downstream tasks?
🤔 If so, how should we train it for better downstream gains?
Major takeaways:
➡️ Visual transition reasoning is a shared, teachable primitive that transfers widely downstream: ~2K-example subsets collectively improve 22 of 26 downstream benchmarks by up to 27.3 points.
🧪 A data-centric training recipe: to teach this shared primitive for downstream tasks, select supervision by the reasoning operation it teaches, rather than scene/action/domain similarity.
Thread🧵👇
🤔 What makes visual transition supervision transferable across downstream tasks?
Controlled comparisons, with prospective validation on held-out benchmarks, reveal a data-centric training recipe:
1️⃣ Select by reasoning operation, not by the actions, scenes, or domains it shows.
2️⃣ Prefer larger changes to the visual state for robustness
3️⃣ Supervise temporal and agent-driven tasks directly
4️⃣ Keep dedicated supervision beyond tasks involving state transitions (e.g., static perception).
Let’s say 50,000 ICLR submissions × 3 reviews × 2 hours = 300,000 hours, or 150 person-years of expert labor.
Let’s use AI to fight AI slop. I would suggest adding an AI-only screening round, then send only the papers that pass to human reviewers.
Let’s not break the system.
I have been wondering how to better define "continual learning" and really like the definition here: "the process by which an apprentice develops expertise on the job".
One finding particularly caught my attention: "human testers got faster with practice, while agents generally slowed down as their memory notes grew. "
To me, this highlights how simply accumulating more skills can sometimes just burden agents with their own memory.
So multi-scale abstraction matters: abstracting experience into reusable knowledge at the right level for each task.
How to quantify continual learning has long been a challenge toward RSI, especially measuring what a model truly learns through experience. Making that distinction measurable in a realistic setting is exciting. Great work @NeoCognition@ysu_nlp !
HumanCLAW: Can VLMs act through a body?
Give VLMs a body, ask it to walk over and sit on the couch.
VLMs always fail, in 1,218 long-horizon tasks across 41 indoor scenes:
walls block...
feet catch on furniture...
objects get knocked aside...
Real robot testing fuse decision and joint control into one system. When they fail, we may not tell what failed. So HumanClaw strips low-level control error out of the result: the body never falls.
In practice: a frozen VLM picks one parametric whole-body skill every 0.5s from egocentric view. A pretrained motion generator turns it into continuous motion.
So here, failure = decision failure.
Where it breaks:
- 1. inefficient exploration: once a target is in view, recognizing it is mostly fine. The more common failure is that the model never brings the target into view at all.
- 2. it sees the couch and still can't get there. The model does not know how far away it is or when it has arrived. It stops early, or walks straight past.
- 3. it is standing in contact with the couch and still does not sit. Sit success ranges from 90% down to 3.5% across models.
- 4. the clearest signal is where the collisions land: mostly on the body parts the model can't see, like legs and feet. VLMs describe a wall correctly, then walk a leg into it.
A clear gap in embodied self-awareness!
Everything is open:
- paper https://t.co/50sfmcdDTd
- benchmark https://t.co/8jzzOaNHaj
- code https://t.co/3NvPM2YFDT
- leaderboard https://t.co/JTz8KnuE4G
Done by the amazing @siyao94@Kuvvius 👇