We have 2M hours of egocentric human experience recorded. While it's the largest library of datasets, it is also the least informative number we publish.
Two recordings of equal duration can differ by an order of magnitude in what a model can learn from them. One holds repetitive tabletop pick-and-place under fixed conditions. The other holds bimanual manipulation across changing states, with precise geometry, varied contacts, coordinated fingers and tool use. Equal duration, nowhere near equal learning value.
Effective experience ≈ hours × information per hour.
At ten thousand hours that difference is noise. At ten million it is the difference between a corpus and an archive.
One thing the video doesn't explain: where the legs come from.
Pull one frame out of it. The camera sits on the person's head. It never sees their legs, not once. The lower body comes out of the fit.
A humanoid policy needs it. A head-mounted camera never shows it to you.
Richer data. Stronger supervision. More reliable robot actions.
This video showcases our precise full-body pose tracking—including individual hand joints—alongside tracking of objects being manipulated and heatmaps of contact regions.
Together, these complementary data modalities provide richer supervision for training humanoid robot models, helping them learn more effective action plans and execute tasks with higher success rates.
Full releaseon on Monday.
This is the input. One head-mounted camera, one person, one ordinary task.
Hand joints, the object, and where the hand is making contact, all from that single view.
Our CTO shows you what comes out of it on Monday.
Maxinsights is heading to #IROS2026!
Meet our team at Booth 630 to talk real-world data for robotics and explore ways to collaborate.
Real-world data. Real conversations. See you there!
The overlays are the output the rig choice produces. Fine assembly in an electronics lab, box folding, gloved hands at a sink. Each one is a case where a single head-mounted view loses the contact, and nothing downstream reports it missing.
This is the depth axis: not more kitchens, but more of each interaction preserved inside the kitchens already recorded.
Every element of the physical state left unrecorded is one the model has to infer from pixels.
Sometimes that is fine. Sometimes it is the difference between learning contact physics and memorising appearance. Either way the decision is made at capture, by whoever writes the protocol, and no downstream process reverses it.
🧵
Single-view head-mounted collection is enough for a large share of tasks.
A lot of it is everyday manipulation, like the outdoor food work here, captured single-view. Running the full rig everywhere would be slower without making most of that footage better, so we reserve it for the contact-critical subset. Matching capture to the task, not maximising sensors, is part of what makes 450K hours a month possible.
End-to-end yield is the one metric of the three that is well defined today.
Of the hours recorded, how many survived to training. Not per stage, end to end. It is answerable now, which is what separates it from marginal information gain and state coverage, and it is the number worth asking any supplier for.
A pipeline is a series of gates, and yield multiplies.
Four gates at 95% give 81% end to end. Push every gate to 99% and yield reaches 96%. At ten-million-hour volumes those fifteen points separate a corpus that gets delivered from one that gets reprocessed.
🧵
Underneath it, the ingestion layer: 84 batches, 799,946 files, 12.84 TB across 100 nodes on a 100 Gbps cluster.
This is the least interesting thing to publish and the part that decides how many recorded hours become training hours.
Two hands working, left on the pot, right on the faucet, water running. MaxVLM returns "Fill the pot with water" with both hand roles assigned, human approved. Both baselines return non_action.
A visible two-handed action, scored as nothing happening.
MaxVLM 1.0 leads all five reported quality metrics on HD-EPIC, evaluated against Gemini 3.7 Flash, GPT-5.6 Terra and Qwen 3.7 Plus.
Three things the table shows:
temporal recall is where the separation is real, 0.6429 against a field best of 0.4376
temporal F1 follows it, 0.7277 against 0.5533
semantic precision does not, 3.7062 against 3.6422
And one thing it does not show, which comes first.
🧵
Precision and recall fail differently.
A precision error surfaces on its own. A wrong label is wrong, something downstream trips over it, someone goes looking. A recall miss stays quiet. The interaction happened, the model did not pick it up, nothing reports it was ever there, and no filtering afterwards brings it back.
For a training corpus that is a coverage problem wearing the clothes of a quality metric.