Even with attempts to standardize robot learning data, things remain too heterogenous. ex - since LeRobot v3 concatenates episodes into shared MP4s, treating the end timestamp as inclusive would append the first frame of the next episode to the previous. Whereas MCAP stores timestamped messages across separate channels, meaning camera and state streams must be aligned explicitly.
Beyond faithful episode reconstruction, metadata fields vary massively both across data types (UMI/teleop/human ego) and within them. ex, Galaxea records mobile base state/chassis IMU, while say build100k is effectively pure RBG.
The goal with Argus is to provide a unified annotation harness over these datasets.
Our belief is that as (a) VLMs continue to push the Pareto frontier of multimodal understanding/cost (as they've demonstrated over the last 3 months) and (b) there exists a harness to provide said VLMs with all of a dataset’s relevant context, dense annotation is a solved problem.
Argus: an open-source robotics data annotation and quality pipeline.
With new frontier VLMs like GPT-6 Astra, we can generate rich, high-quality annotations for robotics data for both training and dataset analysis. Argus delivers detailed, timestamp-level annotations while catching issues like mislabeled instructions, sped-up recordings, swapped camera streams, and unflagged operator mistakes.
To support better data quality for the robotics community, we’re open-sourcing Argus for anyone to use 🧵
We (@PantheonInc) have been working on a robotics data quality pipeline that uncovered a series of major problems in public robotics datasets, especially for world modeling.
To improve the quality of data available to open-source robotics, we're publishing annotations for four of the most popular datasets. Some examples of issues, and our report 🧵