One week of our video data: 16,190 hours. A full-time worker would need 8 years to watch it.
Our agent-native infra analyzes it in a day, with inference on @inco_ai: 31× the throughput per GPU at 1/15 the cost per hour of video.
@inco_ai How the system figures it out:
https://t.co/4U2aWGm7KO
A technician carries a tire across the shop. Whether it's part of the job depends on where the tire came from and where it ends up.
One week of our video data: 16,190 hours. A full-time worker would need 8 years to watch it.
Our agent-native infra analyzes it in a day, with inference on @inco_ai: 31× the throughput per GPU at 1/15 the cost per hour of video.
@inco_ai A real shift isn't all action. Machines run their cycle, people take breaks, parts move between stations. The system finds the work in it, so what reaches training is hours of people actually doing the job.
We're Linewise again. The line is where things get made, and wise isn't only knowing things. It's knowing the right way to do them.
Machines are learning to build in the physical world. Almost everything they need is already on the floor, in the hands of people who do the work every day. We're making sure none of it gets lost.
@visionlabAI has raised a $6M seed round, led by Race Capital, with continued participation from Y Combinator and new backers including Foothill Ventures, 500 Global, and other incredible investors.
At Vision Lab, we are building the real-world data infrastructure for Physical AI: capturing, structuring, and annotating how skilled humans perform real work in the physical world, so robots can learn from it.
We are actively looking for data partners across manufacturing, hospitality, agriculture, logistics, construction, repair, and other economically valuable physical work. If you operate in these environments or know potential partners, please dm me.
Full article on our raise can be found in the comments.
What if a robot could watch a million hours of video and learn to teach itself?
1XWM makes the robot imagine doing the task before it acts, and that imagination is good enough to bootstrap real-world self-improvement.
How they accomplished this:
- A 14B video diffusion model pretrained on web-scale video, mid-trained on 900 hours of egocentric human data, then fine-tuned on just 70 hours of NEO robot data. Given a text prompt and a starting frame, it generates 5 seconds of future video.
- An Inverse Dynamics Model (Depth Anything backbone + flow matching head) trained on 400 hours of unfiltered robot data converts generated frames into executable actions via sliding window prediction.
- Caption upsampling, using a VLM to rewrite brief task descriptions into detailed captions, improves video quality and downstream task success across all evaluation splits.
From 1XWM, we learn that if you can generate physically accurate video of a task being done, extracting actions becomes the easy part. And because NEO's humanoid body closely matches human form, manipulation priors from internet video stay in distribution. What the model can visualize, the robot can usually do.
Robot demos are expensive. Human video is everywhere. EgoScale shows that 20K hours of people just doing stuff with their hands follows a predictable scaling law that directly translates to dexterous robot performance.
How they accomplished this:
- They extract wrist motion via SLAM and retarget 21 human hand keypoints into 22-DoF robot joint angles using a classical optimizer, turning raw egocentric video into dense action supervision.
- A three-stage recipe where we pretrain on 20K hours of in-the-wild human video, mid-train on just 50 hours human + 4 hours robot data with matched cameras and tasks, then post-train on ~100 task-specific robot demos.
- Validation loss which follows a near-perfect log-linear scaling law with data volume (R² = 0.9983), and that loss directly predicts real-robot performance. No saturation at 20K hours.
Results: 56% average success rate on five dexterous tasks (syringe transfer, bottle unscrewing, card sorting) vs 2% from scratch. One-shot shirt folding hits 88% with a single robot demo.
From Egoscale, we learn that human video becomes useful for dexterous control only when you supervise in joint-angle space and bridge embodiment with a small aligned dataset. Without either piece, 20K hours of video does nothing. With both, human data becomes the most scalable motor prior we have.
Day 4 of highlighting fundamental VLA papers.
GR00T-N1 was the first open foundation model for humanoid robots, and its results make a strong case that data quality matters more than architecture.
The backbone uses a Vision-Language Model (Eagle-2) for reasoning at 10Hz and a Diffusion Transformer for generating continuous actions at 120Hz via flow matching. Both are tightly coupled through cross-attention and jointly trained end-to-end. 2.2B parameters, works across GR-1, NEO, and other humanoid platforms with swappable embodiment-specific encoders/decoders.
How they accomplished this:
- A "data pyramid" combining three tiers of data with real teleoperated demonstrations at the base, synthetic trajectories in the middle, and internet-scale human video at the top for VLM pretraining.
- The GR00T Blueprint synthetic pipeline generated 780,000 trajectories in 11 hours using Omniverse simulation and Cosmos rendering. Adding this synthetic data to real data boosted performance by 40%.
- Cross-embodiment generalization comes from training on heterogeneous data across multiple robot platforms. The shared VLM + DiT trunk stays frozen, and embodiment-specific encoders/decoders handle the hardware differences at the boundaries.
The 40% jump from better data on the same architecture clearly demonstrates the importance of data. For robot foundation models, the ceiling is set by the diversity and quality of data.
NVIDIA's newest robot world model, DreamDojo, doesn't pretrain on robot data. It pretrains on 44,000 hours of humans doing things, and that turned out to be the whole point.
Robot demonstrations are expensive, narrow, and embodiment-locked. DreamDojo sidesteps all of this by pretraining entirely on egocentric human video, then snapping onto specific robots with minimal post-training.
How they accomplished this:
- A curated egocentric dataset (DreamDojo-HV) spanning 44k hours, 6,015 tasks, 9,869 scenes, and 43,237 objects. That's 15× longer, 96× more skills, and 2,000× more scenes than the largest prior world model dataset. Scale and diversity of the egocentric footage was the single biggest factor in generalization.
- A Latent Action VAE that extracts 32-dimensional action vectors from consecutive frame pairs. No motor commands, no action labels, no robot in the loop. This makes any egocentric video "robot-readable" by representing "what changed" in a hardware-agnostic latent space.
- Two-phase pipeline: pretrain on massive human video with latent actions to learn general physics and manipulation priors, then post-train on a small amount of robot-specific data to map real motor commands into the latent action space.
The path to generalist robot intelligence runs through egocentric human video.
To make training scalable, VLA should be able to learn from millions of videos that have zero action labels, and UWM proves it.
The bottleneck for scaling VLAs has always been action-labeled demonstrations (expensive to collect, narrow in coverage). Meanwhile, YouTube alone has more robot-relevant videos than every teleoperation dataset combined. UWM cracks this open by making action-free video data directly usable for policy training, no pseudo-labeling or separate inverse dynamics model required. (1/3)
From UWM, we learned that world modeling and imitation learning don't need separate architectures. By decoupling the noise schedules, a single model implicitly learns both, and the dynamics objective acts as a regularizer that teaches visual invariance, not just action copying. The result is policies that are meaningfully more robust to distribution shifts.
Day 1 of highlighting fundamental VLA papers. Paper in comments. (3/3)
They accomplished this with:
- Two independent diffusion processes: one for actions, one for video, both sharing a single transformer. Both modalities also get their own noise timestep, sampled independently during training. The model learns to denoise both simultaneously in one forward pass.
- At inference, you control what the model does by clamping timesteps. Set video timestep to 0 (clean input) and denoise actions for VLA deployment. Set action timestep to 0 and denoise video for video generation.
- For action-free video, missing actions are simply set to maximum noise (t_a = T). The model still trains on the video denoising objective, forcing the shared transformer to learn dynamics and scene understanding from unlabeled footage, which transfers directly into better action predictions after fine tuning. (2/3)
We first started off by building a "FactoryGPT" system to capture manufacturing knowledge and help frontline operators navigate the complexity of factory environments. What we didn't expect was that this same knowledge would become incredibly valuable for training the next generation of robots.
Proud to be part of the robotics revolution. We're already working with two frontier AI labs - if you're curious to learn more, feel free to reach out.