@sama@OpenAI
My 93-year-old grandfather discovers ChatGPT Voice Mode for the first time, and the results are nothing short of amazing. He loved it, and it made him so happy. Affective computing and AI like this will transform life for older adults. #MerryChristmas
NVIDIA Cosmos Lab is looking for PHD interns to help build Open Physical AI foundation models. The job is in Santa Clara, CA. Below is the application link.
https://t.co/vz21MpQEwQ
NVIDIA Isaac ROS 5.0 brings AI agents to robotics development. 🤖
Announced at #ROSCon, the release adds agentic workflows, Isaac Skills, ROS 2 Lyrical support, GPU-accelerated libraries and NVIDIA Jetson support from Orin Nano to Thor.
Open source and available now 📖 https://t.co/6idfj67AZE
Nvidia $NVDA CEO Jensen Huang could pay around $8 billion in taxes over the next 5 years.
“I’m not afraid of paying taxes. I’m just afraid of being poor.”
“Instead of paying people as little as possible to get their work, I try to pay people as much as possible.”
Move the target. The robot finds it again.
This visual grasping pipeline combines #IntelRealSenseD405 + #reBotArmB601-RS + #ROS2 + #YOLOE:
Detection → Depth → Hand-Eye Calibration → Grasp Pose → Pick & Place
Move the pen, and the system detects its new position, calculates a new grasp pose, and sends the arm to pick and place it.
The Seeed Wiki walks through the full #visualgrasping workflow — from camera & ROS setup to detection, calibration, grasp validation and the final grasp-and-place pipeline.
https://t.co/xHYzTqvJsC
How can we improve robot generalization beyond collecting more demonstrations?
A robot can use inference to work out how to solve tasks it was never explicitly trained to perform. With learned models of the world, we can plan our future actions and goals before acting.
I wrote a research perspective, Generalization by Construction, on how learning can be combined with inference for flexible generalization.
https://t.co/MfNWubAp4l
For me, with the rise of agents, AI has become a trade-off between doing more extensive reasoning with fewer iterations, or doing less reasoning per step but relying on more iterations in the loop. That’s why I feel models need to become much faster and more efficient.
MINT is now open source.
From first-person RGB video, MINT reconstructs 3D camera trajectories and bimanual hand motion in a shared world coordinate system. Reconstructed human hand motion can also be retargeted to WUJI Hand.
The release includes model weights, training and inference code, visualization and evaluation tools, EgoPipeline reference code, and structured annotations for 560,649 episodes covering 1,021 hours of first-person video.
Project:
https://t.co/9imtdqIHHH
Code:
https://t.co/NEyO1G9W5b
#MINT #EmbodiedAI #Robotics #WUJIHand
#WRC 2026 Day 1 was a blast!🤖
At Booth A435, @seeedstudio × @RobStride_com are showcasing how #reBot B601-RS goes beyond scripted actions with #VLM Skills—turning natural-language goals into closed-loop Perceive → Reason → Plan → Act → Recover workflows, making reBot more agentic!
We’re also launching our new Physical AI course, collaborated with the #NVIDIA, taking developers from simulation to real-world #VLA deployment with reBot arm + NVIDIA #saac. Check full course at: https://t.co/COF5uchtI6
Huge thanks to the @NVIDIARobotics team for the continued support! 🙌 Come see Agentic #PhysicalAI in action at WRC 2026, Booth A435, Aug 19–23 in Beijing, China!
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
🤖 Introducing RoboTok, the “TikTok for robots.”
Just as TikTok recommends videos to people, RoboTok recommends relevant human demonstrations for robot learning.
Robot learning needs broad and diverse demonstrations, but collecting robot data is expensive. RoboTok is an internet-scale data engine that uses web video as a scalable and continuously growing source of demonstrations for dexterous manipulation learning.
Given one human demonstration video as a query, RoboTok retrieves other web videos with similar underlying manipulation motions.
Rather than matching videos by labels or visual appearance, RoboTok compares how the hands move over time.
💡 The key idea is to represent each video with canonicalized 3D hand trajectories. Each trajectory is expressed in an estimated actor-centered reference frame, so the movement is described relative to the person rather than the camera. This makes similar manipulation behaviors easier to compare across different viewpoints and scenes, even when the actor is partly occluded.
In our experiments, RoboTok retrieved more manipulation-relevant demonstrations than existing robot-data retrieval methods. When those videos were used to guide robot training, the simulated robots completed manipulation tasks more successfully.
I’m sincerely grateful to Howard Qian and Kaiyu Hang for leading this project. I also want to thank Yiting Chen, Kejia Ren, Podshara Chanrungmaneekul, Gaotian Wang, Bowen Wen @bowenwen_me, and Chen Wei @_Chen_Wei_ for their guidance and collaboration.
Project site: https://t.co/v43CCf0040
Paper: https://t.co/WZGwQLhDTK
Code: https://t.co/MfYLTxHopH
Data and models: https://t.co/i0s7RcrXvX
The ICML 2026 oral recordings are now live:)
Check out our talk on motion attribution for video generation, where we study how different training examples influence generated motion and what this can reveal about video generation models.
🔗 https://t.co/5r8JmnQQ0N