@toolshed_labs@notgiannei Yep. Which is why you must either fine-tune them with your specific use case plus rely on some 3D geometry from your sparse/dense Landmarks map + Camera poses to fix your depth (this is still open ended and is actively being researched in our lab).
@toolshed_labs@notgiannei Stereo depth itself is pretty sparse + sometimes noisy to work with
Recovering scale from IMU should never be done if you cant reliably calibrate or initialize your IMU. It'll only worsen your reconstruction.
State of the art depth estimators + depth completion algorithms ;)
@pablovelagomez1@notgiannei Currently the fast mode operates around 8-10 fps. The system is heavy because accurate reconstruction of the environment is what we care about the most since it'll be the core for downstream tasks
Today we’re announcing microSLAM, the new state-of-the-art in monocular SLAM from human egocentric video, built by our Zürich-based computer vision team in collaboration with the ETHZ Computer Vision and Geometry Lab, @rohamzn and @Ace3Zi
We are nr. 1 on the LaMAria Benchmark for monocular SLAM, and one of the best SLAM systems built so far, when taking into account any amount of cameras and any amount of IMUs (with just monocular RGB)
microSLAM turns plane monocular video into fully interactive 3D environments, so called RL environments, which teach state-of-the-art behavioral foundation models to understand their surroundings, navigate across factory floors and become experts at the industrial use cases our clients care about.
These results mark a foundational step toward our goal of deploying a billion robots worldwide, bringing the value from foundational models to the physical world.
Check out the article on our website to dive deeper into microSLAM’s design and performance.
We will release the microSLAM paper and code in the near future, providing the community with a reproducible foundation to evaluate, extend, and build on this work.
Project Update: I’m training a World Action Model for general bimanual manipulation.
Here’s everything I’ve learned so far, including the CPU contention bug that was silently corrupting my data. Thanks to @nvidia, @AWS for the compute and @ethroboticsclub for the hardware! (1/8)