1/
Super excited to announce that our work, Depth Any Camera (DAC) has been accepted to CVPR 2025!
A zero-shot metric monocular depth estimation framework that generalizes seamlessly to different camera types – no retraining needed! 📸
@Bosch_AI @BoschGlobal
🧵👇
Announcing Series C
We’ve raised $1.4B, valuing the company at over $14B
With this capital, we will accelerate our mission to build omni-bodied intelligence 🚀
https://t.co/q0ArkRa8L8
🔉 Introducing SAM Audio, the first unified model that isolates any sound from complex audio mixtures using text, visual, or span prompts.
We’re sharing SAM Audio with the community, along with a perception encoder model, benchmarks and research papers, to empower others to explore new forms of expression and build applications that were previously out of reach.
🔗 Learn more: https://t.co/FPnfv66UCP
😠💢😵💫Tired of endless data collection & fine-tuning every time you try out VLA?
Meet RDT2, the first foundation model that zero-shot deploys on any robot arms with unseen scenes, objects & instructions.
No collection. No tuning. Just plug and play🚀
Witness a clear sign of embodied superintelligence
- 7B one-step diffusion → 23 Hz inference⚡
- Re-designed UMI @chichengcc@SongShuran and manufactured 100 portable devices
- Trained on 10K-hour UMI data on 100 real houses
- Zero-shot: pick, place, press, wipe… open-vocabulary
- Demos: block 30 m/s arrows in 500 ms🛡️; first to play ping-pong with an end-to-end model 🏓; extinguish burning incense by shaking quickly🥢
Fully open source at https://t.co/juzBEY4H45
Project page: https://t.co/z8ce3sAvFs
Thanks to awesome collaborators @bang_guo96535@D0g4M74794@EthanNg51931527
We built a robot brain that nothing can stop.
Shattered limbs? Jammed motors? If the bot can move, the Brain will move it— even if it’s an entirely new robot body.
Meet the omni-bodied Skild Brain:
A robotic ballet! 🩰
Coordinating multiple robot arms on a busy factory floor is notoriously complex.
Each arm needs to move without colliding with its neighbors or the surrounding equipment, and today that planning is still mostly done by hand, a process that takes specialists hundreds of hours.
Researchers at @ucl, @GoogleDeepMind, and @intrinsic have introduced RoboBallet, a new AI system that tackles this challenge head-on.
Using a combination of graph neural networks and reinforcement learning, RoboBallet learns how to coordinate many robots in real time, generating smooth, collision-free motion plans in seconds instead of days.
In tests, the system handled up to 40 tasks with eight arms working together, far beyond the limits of traditional planners.
More importantly, it showed strong generalization: RoboBallet could adapt instantly to new factory layouts or recover if one robot failed, something previous methods struggled with.
For manufacturers, this could mean faster deployment of automation, less downtime, and the ability to reconfigure production lines on the fly. Applications range from automotive welding to electronics assembly, and even large-scale construction.
The current system focuses on reaching tasks, but the team plans to expand it to more complex operations like pick-and-place or painting.
Here's the paper: https://t.co/5RyIJCLvKM
Ever wish a robot could just move to any goal in any environment—avoiding all collisions and reacting in real time?
🚀Excited to share our #CoRL2025 paper, Deep Reactive Policy (DRP), a learning-based motion planner that navigates complex scenes with moving obstacles—directly from point cloud input.
w/ @Jiahui_Yang6709
(1/N)
Rugged by design. Elevated by nature.
The #LucidGravityX concept redefines what a trail-ready adventure vehicle could be.
Read more about our new bold concept: https://t.co/BxgKiyfG8g
We’ve all seen humanoid robots doing backflips and dance routines for years.
But if you ask them to climb a few stairs in the real world, they stumble!
We took our robot on a walk around town to environments that it hadn’t seen before. Here’s how it works🧵⬇️
AI that truly understands the physical world should not be limited by robot type or tasks.
We tackle robotics in its full generality @SkildAI.
The goal is to build a continually improving, omni-bodied brain that can control any hardware for any task.
https://t.co/TGYGRHPgEq
TRI's latest Large Behavior Model (LBM) paper landed on arxiv last night! Check out our project website: https://t.co/AV2cmfeX40
One of our main goals for this paper was to put out a very careful and thorough study on the topic to help people understand the state of the technology, and to share a lot of details for how we're achieving it.
https://t.co/EVFLJAY6Zu
The neural network objective function is a very complicated objective function. It's very non convex, and there are no mathematical guarantees whatsoever about its success. And so if you were to speak to somebody who studies optimization from a theoretical point of view, they would tell you that there is no theoretical reason to believe that the optimization will succeed. And yet it does. And this is an empirical fact.
-- Ilya Suskever in 2015
Mastering deep learning = gaining intuition into why this in fact succeeds.
Excited to share that TokenVerse won Best Paper Award at SIGGRAPH 2025! 🎉
TokenVerse enables personalization of complex visual concepts, from objects and materials to poses and lighting, each can be extracted from a single image and be recomposed into a coherent result. 👇
Log-linear attention — a new type of attention proposed by @MIT which is:
- fast and efficient as linear attention
- expressive as softmax
It uses a small but growing number of memory slots that increases logarithmically with the sequence length.
Here's how it works:
I'm excited to share our new work Diffusion as Shader (DaS), a versatile controllable video generation method for various tasks: object manipulation, camera control, mesh-to-video, and motion transfer.
Project page: https://t.co/KMkUnCkwCd
Github: https://t.co/4IkUlzkIqF
We move our eyes actively—driven by survival and efficiency—but we still don’t fully understand how. That makes supervised learning hard.
In our new work, we explore how to train VLMs to reason visually using RL. ViGoRL offers a glimpse into how models like o3 might be trained.