LLM post-training used to mean fine-tuning to a downstream task
Robotics has been stuck in this setting, needing task-specific fine-tuning for best performance
π07 changes this: It works out of the box & outperforms fine-tuned specialists
Details: https://t.co/QbO3E4D3QN
New blog post with @perryadong on what we need for robots to be broadly useful in the real world.
https://t.co/jxsNeMUhFg
Reliability is the one of the biggest open challenges in AI right now. Current models work out okay if a person will be reviewing the outputs (eg drafting code), but it will become more of a bottleneck as we want systems to act with more autonomy and more trust.
What will be the “RLHF” moment for robotics? What will it take to get robotics to where LLMs are today and beyond?
New blog post with @chelseabfinn sharing some thoughts on the state of RL for frontier robotics models and what's missing 👇
Blog: https://t.co/crBN1jL5HH
Asynchronous VLA inference reduces inference delay, but breaks the Markovian assumption necessary for RL fine-tuning.
How can we enable RL fine-tuning of VLAs with async inference? We introduce ARLI: Asynchronous RL with Intermediate Information!
https://t.co/5A2OrJcaUH
(1/n)
RL for dynamic tasks requires handling latency.
Building on EXPO-FT, we let the base VLA operate on older images while a small edit policy operates in real time.
Outperforms RTC by a significant margin, across the board.
Paper: https://t.co/TkydzocalS
Introducing Real-Time EXPO-FT – Fast and Reliable RL for Real-Time VLA Policies!
Real-Time EXPO-FT unlocks π0.5 on challenging dynamic tasks, such as balancing a ball on a plate and striking a ball into the goal
(1/6)
A video from a Pi robot deployed at Dandelion Chocolate, fully autonomous w/ no interventions. 🤖
Deploying robots has taught us surprising lessons about the gap between proof-of-concept (i.e. building one box) and real-world utility (productively building boxes for hours).
I expected that the hard part is building the box, since it’s the most dexterous, but that wasn't what we found. Counterintuitively, the hardest part was reliably stacking the boxes. Our original table-mounted robot had poor visibility of the stack without special separately-mounted cameras. Plus, stacking requires more generalization (each box is placed in a different location), and an imprecisely-placed box can lead the entire stack to collapse many boxes later.
We recently switched this deployment over to a mobile robot, and it's fun to watch the robot being completely self sufficient for multiple hours. 🙂 Data and feedback from real world use-cases like these are quite valuable for the π pre-trained model as we scale!
Introducing OM-1, our first robot foundation model, zero-shot generalizing to any robot: table-top arms, industrial arms and humanoids.
- learned directly from human manipulation data
- no teleop/robot data
- close to human-level dexterity and efficiency
- multi-robot collab
Harness optimization is sample-efficient but plateaus. What should you do if you can afford to update the model too?
Introducing WHALE: a simple recipe for jointly optimizing an LLM's weights and harness.
Blog: https://t.co/5La2WaoW11
Paper: https://t.co/xbc9bNEmT0
One of the most important aspects of scientific discovery is deciding where to draw insights from.
While LLMs are promising tools for science, we lack datasets & evaluations for this step.
Help contribute to a public dataset for exactly this: https://t.co/X8NPVnVIwm
We’re asking the research community to help us build a benchmark for research taste.
Scientific discovery starts with a fundamental step: which prior work is worth building on? We want to capture this undocumented layer through our collective knowledge.
Please sign up: https://t.co/T4vfwQnHUY ↓
World models have emerged as one of the biggest directions in physical AI. At the same time, RL fine-tuning is unlocking capabilities in frontier models beyond what pretraining can achieve on its own
Can we get the best of both worlds?
We propose Q-Learning with World Models (QWM)
(1/7)
Robots can already fold laundry, make espresso, clean kitchens, and assemble things. The harder problem is getting them to do those tasks reliably, for long periods of time, without a human babysitting them.
At Startup School 2026, @physical_int cofounder @chelseabfinn explains what it takes to build general-purpose robots that work in the real world.
She shares how reinforcement learning pushed robot throughput up 2x, how their systems can run autonomously for hours, and why she believes robotics is entering its GPT era: moving from specialized models toward general-purpose systems that can work across tasks, robots, and environments.
00:00 — The State of Physical Intelligence
01:23 — What It Takes to Make Robots Useful
05:11 — The Reliability Problem
07:43 — Reinforcement Learning for Robotics
09:35 — Learning From Failures
12:43 — Training Robots to Improve Themselves
14:21 — Can a Robot Work for 13 Hours Straight?
17:36 — Why Robots Need Memory
21:22 — Building a General-Purpose Robot
25:02 — From Fine-Tuning to Out-of-the-Box Models
27:35 — Training on All the Data
30:20 — One Model That Beats the Specialists
31:21 — Compositional Generalization
37:49 — The GPT Era of Robotics
39:49 — Q&A
Pretraining a Q-function often doesn’t actually help RL finetuning, compared to initializing Q from scratch.
We find that pretraining Q-functions on data from diverse policies is critical to see improvements from pretraining.
Paper: https://t.co/dlj5RXVFED
Pretraining has worked remarkably well across domains
We show this doesn’t hold for Q-functions in online RL from a pretrained policy — and propose IPE, a more effective way to learn Q-functions for online RL fine-tuning
(1/6)
I'm giving a talk tomorrow at ICML on emergent physical generalization, including π0.7 🤖
3:15 pm @ SCALE workshop in Ballroom 201
https://t.co/xppCOkktOJ
RL is hitting a ceiling with human feedback. What if the world itself becomes the signal?
Join us at the RLxF: RL from World Feedback 🌍 workshop at ICML 2026 @icmlconf tomorrow (July 10th)!
Web page: https://t.co/cN0itnL1yI
Project led by @marceltornev, @anubhamahajan01, @AbhijnyaBhat
Paper: https://t.co/2dyFJpxwU7
Code & videos: https://t.co/KBAPSbr6a8
Check out Marcel’s thread for more details!
https://t.co/tlUBH8VJqY
We should stop optimizing robot policies against a single overall reward. Trajectories differ along many axes, such as speed, precision, and subtask completion, and one can be better on some while worse on others. If we collapse all of that into a single overall axis we lose this structure making the reward ambiguous and harder to optimize.
Blog: https://t.co/WXWue03RVq
Paper: https://t.co/AvJ904Xt9S
Freeform preference learning has multiple nice properties:
(1) It works better
When controlling for the number of preference queries, learning with multi-axis preferences yields far more performant policies than single-axis rewards.
(3) Long-horizon credit assignment
Most robot RL focuses on short horizon tasks b/c dense temporal rewards are hard to get.
Freeform preferences yield dense rewards for subtasks without subtask segmentation.
With freeform preference data, we train a lang-conditioned reward that captures all axes of a task.
We then train a policy conditioned on each reward axis and the corresponding reward.
FPL allows the robot to maximally leverage and learn from each axis of supervision.
Freeform preferences let the supervisor define relevant axes and then specify preferences along those axes.
Axes can be either a fixed rubric or freeform language.
This eliminates ambiguity, allows for thorough coverage of all axes, and provides more dense supervision.