Founder in San Francisco.
Every robotics company says data quality matters, but very few objectively evaluate the human operators responsible for collecting that data. Behind every training set are a LOT of humans working to collect enough episodes. Each with different speed, smoothness, precision, hesitation, and recovery ability. All of it flows straight into the model through a single gate, pass or fail. But operator quality isn't binary, and someone had to put a score on it. So we did. Today we’re launching a demo of what we've been building at Munari. We score episodes on the kinematics that separate a “good” demonstration from a “bad” one, roll them up to a 0-10 score per episode and per operator, and put it in a easily queryable view to analyze findings. Our demo looks at ABC-130K, XDOF's open bimanual teleop dataset: roughly 130,000 episodes that ship with anonymized operator IDs nobody has published a look at. The paper reports scale, task diversity, and policy results. Nothing about the people who made it. What's interesting is how two episodes of the same task, in a clean dataset, can still sit far apart on the scale. At normal playback speed you often can't tell which is which. But when you drill in, you can uncover significant variance. And it’s clear that certain operators are better equipped at these tasks than others. For us, kinematic quality is a starting point, not the final rubric for what makes good training data. That answer is different for every company and every policy, which is why we built our scoring to be adapted, reweighted, and validated against what data actually improves your policy. And of course, please reach out if you want to connect or share feedback! Demo linked below.
Read the post on X
Looking for
Can help with
Not shared yet.