A question I keep coming back to: How can robotics teams get more useful training data for every dollar they spend collecting it?
It’s easy to count hours and episodes. Figuring out what to change is harder. Should you spend more time on certain tasks? Give an operator more practice? Take a closer look at why some demonstrations aren’t working?
Those are the decisions we want Munari to help teams make. Start with operator performance, look at quality across tasks, and use actual evidence to learn how to improve.
Robots learn from human demonstrations. The quality of those demonstrations start with the operator.
Here's a preview of @MunariAI: compare operator performance, inspect data quality by task, and explore the episodes behind every score.
ABC-130K has operator IDs on ~130k episodes. The paper uses them for style. I haven’t seen anyone ask if those operators were actually any good.
We scored kinematics on each episode. Same task, people come out all over a 0-10, and you often can’t tell watching at normal speed. Then we train on the whole pile anyway.
We’re improving the human layer behind the robots. Our platform scored the operators behind data collection and teleop tasks, so companies can get more useable data per teleop hour
If you want to see how your operators perform, try uploading some episodes or reach out to us!
Every robotics company says data quality matters, but very few objectively evaluate the human operators responsible for collecting that data.
Behind every training set are a LOT of humans working to collect enough episodes. Each with different speed, smoothness, precision, hesitation, and recovery ability. All of it flows straight into the model through a single gate, pass or fail. But operator quality isn't binary, and someone had to put a score on it.
So we did.
Today we’re launching a demo of what we've been building at Munari. We score episodes on the kinematics that separate a “good” demonstration from a “bad” one, roll them up to a 0-10 score per episode and per operator, and put it in a easily queryable view to analyze findings.
Our demo looks at ABC-130K, XDOF's open bimanual teleop dataset: roughly 130,000 episodes that ship with anonymized operator IDs nobody has published a look at. The paper reports scale, task diversity, and policy results. Nothing about the people who made it.
What's interesting is how two episodes of the same task, in a clean dataset, can still sit far apart on the scale. At normal playback speed you often can't tell which is which. But when you drill in, you can uncover significant variance. And it’s clear that certain operators are better equipped at these tasks than others.
For us, kinematic quality is a starting point, not the final rubric for what makes good training data. That answer is different for every company and every policy, which is why we built our scoring to be adapted, reweighted, and validated against what data actually improves your policy.
And of course, please reach out if you want to connect or share feedback! Demo linked below.