Gemini Robotics 2 is here, with our new suite of models, robots can now reason through every movement to manage tasks that weren’t possible before, like tying delicate knots - and even team up to solve complex workflows. Huge congrats to the robotics team on this great milestone!
Here’s a first look at Gemini Robotics 2 from @GoogleDeepMind on the FR3 Duo: 20 minutes of uninterrupted, real-time tool kitting.
Notice the emergent recovery behaviours - generalized dexterity & precision at a whole new level.
More demos dropping over the next days! Stay tuned!
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. It's kind of janky but fun. But it's a bit mindboggling that the LLM has to place and orchestrate various polygon assets in (x,y,z) coordinates and write code that animates it all, and that it even does anything at all.
I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". There might be a lot more. But I'm excited about creating hyper custom worlds that you can imagine dropping players into, e.g. here to participate in the LoTR story as a spectator NPC, or one of the characters, or etc. Something like an ephemeral GTA of X on demand.
Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Here, Opus 5 had to very slowly and painstakingly take screenshots at different points, and it messed up a few times and created a bunch of jank. An example of raw capability (multimodal, gameplay) that I think is still quite lacking.
Combining real-time interactivity, task understanding, and full-body action prediction on a humanoid is so, so hard.
Here's an example where we bring all of these together in Gemini Robotics 2 🤖🧠
Gemini Robotics 2 is here, with our new suite of models, robots can now reason through every movement to manage tasks that weren’t possible before, like tying delicate knots - and even team up to solve complex workflows. Huge congrats to the robotics team on this great milestone!
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
Ctrl-World is a controllable world model that generalizes zero-shot to new environments, cameras, and objects.
Paper: https://t.co/Bog5hk0h28
Model & code: https://t.co/3OFHRAXv2O
The results are exciting — a short thread on why. 🧵
Unitree Introducing | Unitree H2 Destiny Awakening!🥳
Welcome to this world — standing 180cm tall and weighing 70kg. The H2 bionic humanoid - born to serve everyone safely and friendly.