Set for ~20 minutes. Landed ~10 seconds.
The headstrap pressed the phone’s side camera control and stopped the take. The contributor kept going. The episode was already gone.
Hardware buttons must not kill a recording. Stop only from the screen, with accidental-tap protection.
30 minutes of robot video every second, and still they signed a $3.5B compute deal. That tells you where humanoids are headed. The bottleneck is data plus GPU, not hype.
MIT defenders are now talking about simulation-ready worlds. That is the real bottleneck in robot learning. If you cannot generate good worlds, you cannot test at scale.
Operator hesitation is poison in robot data. Clean the pauses, the static frames, the wobble, and you get training sets a robot can actually learn from. That is the real product.
50,000+ teleop trajectories and 60,000+ scene variants. That is the part people keep missing. Robot data gets real when coverage is wide enough to beat hand-curated demos.
Unpopular opinion: the data pipeline is already here. Factory workers are wearing cameras so robots can learn from their hands, not a lab demo. The real fight is consent and job pressure.
100,000 GPUs. That is the tell. Humanoid training is becoming a data center problem, fast. The teams that win will look less like robot shops and more like infra companies.
700 hours of mocap, one model for humans and humanoids. That is the interesting part. The best robotics data story is not more labels, it is cleaner motion that transfers across bodies.
A human is still the best robot sensor. NVIDIA open-sourced Isaac Teleop, and that matters because the fastest way to train a robot policy is still a person driving the edge cases. That is the data moat.
1M+ real robot trajectories. That’s the part people miss.
Robots are starting to train on shared episode data, pooled across labs. It feels a lot more like text corpora for LLMs than old-school hand built demos.
Before the AI, build the safety layer. Robotics fails fast when people skip the boring hardware, control, and safety work. The model can learn later, but it cannot fix a bad base.
NVIDIA is turning one motion stack into two products. One for animation, one for humanoids. That is the real tell. The bottleneck is not model size, it is clean motion data at scale.
We prefer a global shutter. A 120 degree fisheye is a weak default. Wide does not mean usable.
LAWM-3D this month: single-view pixels do not give you 3D for free. We need a clip that can hold geometry. The lens is a collection decision.
50 homes, 3 takes, and a bunch of robots. Synthetic data only works when it feels like production, not a demo. That is the point of this Isaac Sim workflow.
Bad angle is a reject. Object in view, hands out of frame. Same as a missing hand.
AtlasVLA this month: a wrist camera forgets what leaves the field of view. We fail the take earlier. If the clip never held the grasp, there is nothing to remember.
Current videos fail review when the hands leave the frame. The task can look done. The take does not count.
IROS Physical World Models notifications land today. Data quality is on the program, not just more rollouts.
If the hands are gone, it is not a take.
A public write-up last week audited a few thousand local training episodes and found more sitting in review or exclude than pass. Idle stretches were the usual reason.
On our days, five or six hours of footage is usually about ten hours of work. Recording is the fast part. Review decides whether a take counts, including takes that are mostly waiting.
A paper this month ran eleven video models on one job: turn a first-person human demonstration into a robot video. The leaders still lost contact with the object.
On our days, five or six hours of footage is usually about ten hours of work. A lot of that is watching hands. Review is where we set a take aside when the contact fails.
WorldArena is still open this week. The IROS challenge is scoring whether a predicted next moment stays physically usable.
On our collection days, five or six hours of footage is usually about ten hours of work. Recording is the fast part. Review decides whether a take counts.
We do not know yet how many submitted hours survive that pass. We know that is where the day goes.