A first look at what we’re building at Sudo AI. Introducing #sudo R1.
We train robots entirely in simulation, then deploy them directly to hardware - no data collection, per-object tuning, or calibration.
In this first release shared in our blog, #sudo R1 runs 60 minutes uncut and reliably picks objects it has never seen before: glass, fabric, metal, irregular shapes - under changing lighting, clutter, and physical interference.
Simulation-trained systems rarely hold up in practice. We think that’s starting to change.
If you’re building in robotics or thinking about where this is going, follow along at https://t.co/lR8iDEhTW7
Yesterday the hyped Genesis simulator released. But it's up to 10x slower than existing GPU sims, not 10-80x faster or 430,000x faster than realtime since they benchmark mostly static environments
blog post with corrected open source benchmarks & details: https://t.co/7f183ZXVGv
Two years ago today, we had DreamFusion. Since then, the field has advanced rapidly, and it’s been great to take this wild ride. At #ECCV2024 , I will present Sparp onsite with a short oral pre on CV4Metaverse today, a poster on Wild3D tomorrow, and a main conf poster on Friday.
Single-image-to-3D is ill-posed. What if we had a few more views? We introduce 𝗦𝗽𝗮𝗥𝗣, a framework that understands spatial relationships between 𝘂𝗻𝗽𝗼𝘀𝗲𝗱 views to predict camera poses and reconstruct the underlying 3D objects. #ECCV2024
Project: https://t.co/BFzGs6FPbw
Single-image-to-3D is ill-posed. What if we had a few more views? We introduce 𝗦𝗽𝗮𝗥𝗣, a framework that understands spatial relationships between 𝘂𝗻𝗽𝗼𝘀𝗲𝗱 views to predict camera poses and reconstruct the underlying 3D objects. #ECCV2024
Project: https://t.co/BFzGs6FPbw
Thanks to @_akhaliq for featuring our #ECCV paper!
🌋Project: https://t.co/BFzGs6Gn14
🤗Demo: https://t.co/rnlfFzh2mh
A quick intro to the paper in my pinned post: https://t.co/bCKflLlQvP
#AIGC#GenAI#Imageto3D#Hillbot#sudoAI
SpaRP
Fast 3D Object Reconstruction and Pose Estimation from Sparse Views
discuss: https://t.co/tzzJQDLHnl
Open-world 3D generation has recently attracted considerable attention. While many single-image-to-3D methods have yielded visually appealing outcomes, they often lack sufficient controllability and tend to produce hallucinated regions that may not align with users' expectations. In this paper, we explore an important scenario in which the input consists of one or a few unposed 2D images of a single object, with little or no overlap. We propose a novel method, SpaRP, to reconstruct a 3D textured mesh and estimate the relative camera poses for these sparse-view images. SpaRP distills knowledge from 2D diffusion models and finetunes them to implicitly deduce the 3D spatial relationships between the sparse views. The diffusion model is trained to jointly predict surrogate representations for camera poses and multi-view images of the object under known poses, integrating all information from the input sparse views. These predictions are then leveraged to accomplish 3D reconstruction and pose estimation, and the reconstructed 3D model can be used to further refine the camera poses of input views. Through extensive experiments on three datasets, we demonstrate that our method not only significantly outperforms baseline methods in terms of 3D reconstruction quality and pose prediction accuracy but also exhibits strong efficiency. It requires only about 20 seconds to produce a textured mesh and camera poses for the input views.
SpaRP generalizes well in the real world, accurately predicting camera poses and reconstructing 3D objects effectively from both Amazon product images and real-life photos captured on mobile devices.
Thanks for sharing! 🚀 Our latest project is a cutting-edge framework born from the embrace of the multiview-based feed-forward paradigm in One-2-3-45 for Image-to-3D. Upgraded multiview and 3D recon modules. A perfect fusion of speed and quality – the best of both worlds! #AIGC
One-2-3-45++: Fast Single Image to 3D Objects with Consistent Multi-View Generation and 3D Diffusion
paper page: https://t.co/r7ylXY0Tlv
Recent advancements in open-world 3D object generation have been remarkable, with image-to-3D methods offering superior fine-grained control over their text-to-3D counterparts. However, most existing models fall short in simultaneously providing rapid generation speeds and high fidelity to input images - two features essential for practical applications. In this paper, we present One-2-3-45++, an innovative method that transforms a single image into a detailed 3D textured mesh in approximately one minute. Our approach aims to fully harness the extensive knowledge embedded in 2D diffusion models and priors from valuable yet limited 3D data. This is achieved by initially finetuning a 2D diffusion model for consistent multi-view image generation, followed by elevating these images to 3D with the aid of multi-view conditioned 3D native diffusion models. Extensive experimental evaluations demonstrate that our method can produce high-quality, diverse 3D assets that closely mirror the original input image.
📢Excited news from sudoAI (@sudoAI_): The interactive demo of our 3D Generative AI model is online! (alpha test for desktop, mobile / pad coming soon)
✨Transform images & text into stunning 3D models in just 60 seconds!🚀
Try it now!
👉https://t.co/hW1nZshwqK