RL2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
Overview of RL2. RL2 relies on base VLA samples to approach the task object, then adaptively applies RL compositional steering to diversify actions toward successful states β when failure is preemptively detected.
π¦ΎAction Distribution Visualization PCA reveals that RLΒ² shifts the VLA action distribution closer to the ground-truth action during failure states.
π― Most VLA steering methods break down exactly when they're needed most: under out-of-distribution failures.
We uncover distinct test-time scaling laws ππ for success and failure states, revealing why.
This leads to ππΒ²-πππ π β an adaptive RL compositional steering framework that intervenes only when the VLA is likely to fail.
π§΅π
πExtensive Evaluation We benchmark RLΒ² on OpenVLA, Οβ, and Οβ.β across simulation (SIMPLER / PolaRiS) and real-world tasks under OOD prompts and task environments. RLΒ² achieves up to +17.5% average improvement over strong steering baselines.
π οΈRLΒ² Framework We train a lightweight offline RL flow-matching policy conditioned on VLA action expert latents, then compose its flow velocity with the frozen VLA at every flow-matching step. SAFE, a SOTA multitask failure detector, activates steering only upon failure.
πSuccess Scaling Laws The opposite holds during success states. Diversity-inducing steering can unnecessarily perturb already-accurate actions, leading to poor scaling compared to baselines. This motivates adaptive steering only when the base VLA is failing.
πFailure Scaling Laws During failure states, action diversity is most beneficial. Diversity-inducing steering methods exhibit the strongest scaling behavior, with RLΒ² compositional steering performing the best by generating actions beyond dominant VLA failure modes.
π§ The OOD Challenge
While VLAs excel on in-domain tasks, their performance drops sharply under unseen language instructions and task environments.
π‘ Recent work tackles this with inference-time steering, enabling better OOD generalization without costly retraining or additional data collection.
π How can we steer pretrained VLAs beyond dominant failure modes in OOD scenarios?
We uncover distinct test-time scaling laws π for steering approaches under success and failure states.
This motivates ππΒ²-ππππ«, an adaptive RL compositional steering framework.
π§΅ π
We're growing the @cortexairobot team across San Francisco, Singapore, and Malaysia.
We're hiring Robot Operator Managers, Robot Operators, and Software Engineers to scale real-world robotics data, deployments, and infrastructure for physical AI.
Check out our open roles: https://t.co/UzGimN9E9X
One year at MARMot Lab, Guillaume Sartoretti's group @NUSingapore, as I wrap up here. Probably the most important arc for me.
Got to work on test-time steering for VLAs, adapting a pretrained policy at inference without touching the weights, and figuring out the hard parts of that.
Thanks to Guillaume for giving me the opportunity, and to @derektan1995 , genuinely one of the best people I've gotten to work closely with. And @sood__shivam , William Teo, Jeric Lew, Srikrishna Iyer for all the great discussions we had along the way. π
VLA-JEPA just dropped in LeRobot π€
What makes this model special is that it does not just learn what action to take from a given observation, it also leverages a JEPA world model to learn action-relevant dynamics.
During training, the VLA leverages V-JEPA2 by conditioning its predictor. This clever trick adds a world modeling objective to the training, which also allows pretraining on human videos.
At inference, the world model is dropped entirely, keeping only a standard VLA architecture: Qwen backbone and action head.
The demo here was only fine-tuned on 13 examples, showing great pretraining capability and running in real time on @NVIDIARobotics DGX Spark!
VLA-JEPA is the first world model to be ported to LeRobot, and I feel like it won't be the last π
@Thom_Wolf@ClementDelangue