Almost all bimanual robots include wrist cameras. They add bulk, cost, and complexity, but are critical for performance. Are they really needed?
Excited to finally share EyeRobot 2.0, where we use active gaze from an ego stereo camera instead 🧵 (with @kushtimusprime)
Can a robot design a tool from scratch? Introducing HOT: Robot Tool Design from Scratch via Behavior-Aware Hierarchical Optimization.
Check out our project page for more details, videos, and code! 👇
https://t.co/1YLAThub0W
turns out when you have too many GPUs for your own good you can do stupid stuff such as pre-training a foundation model with constrained decoding on PTX -- that is, it does not speak natural language, it's tokenizer is PTX, it's thoughts are PTX, it's all PTX 🧵
Yes, GPT Astra can control a humanoid through long-horizon tasks in the real world!
HomeBody asks a simple question: can a frontier VLM skip learned VLAs entirely? @giohuh_ engineered an incredible harness that lets VLMs like Astra use persistent spatial memory to directly orchestrate composable humanoid skills — no learned VLA in the middle.
A few of my favorite moments beyond the main demo ↓
Gio is one of the most cracked guys I know. It was amazing to see him work on this. The website is super clean too.
You can plug in any model and any skills and let his harness cook.
What can Astra do when given a humanoid embodiment?
We built HomeBody to find out. Controlled by GPT Astra, it carries out long-horizon tasks in a previously unseen kitchen—from tidying up across the room to retrieving remembered objects from ambiguous requests—without environment-specific training data or additional policy learning.
Here's how we did it 👀: https://t.co/FFGS2F7VtI
What can Astra do when given a humanoid embodiment?
We built HomeBody to find out. Controlled by GPT Astra, it carries out long-horizon tasks in a previously unseen kitchen—from tidying up across the room to retrieving remembered objects from ambiguous requests—without environment-specific training data or additional policy learning.
Here's how we did it 👀: https://t.co/FFGS2F7VtI
1/ Introducing PhilosophyBench from @StanfordAILab@StanfordHCI, the first independent, large-scale benchmark for evaluating AI’s philosophical capabilities.
https://t.co/csZ86gugw6
I put TypeSafe’s JEV in control of G1 humanoid in my virtual kitchen.
The mission: reach the stove with obstacles on the way.
A VLM and depth information helped map the kitchen.
A path planner produced possible ways around the obstacles, and JEV chose which move to make next.
A pretrained walking policy handled locomotion, with calibrated controls for turns, forward movement, and curves.
The robot reached the stove in about 64 seconds of simulation time. This video shows the highlights.
Still a prototype, but pretty cool watching perception, decisions, and movement come together.
Deploying DiffusionGemma-Jev (djev) just got a lot easier. You can now spin up a Jev API-compatible endpoint on Google Cloud Run using a single command.
Performance is solid: ~35-60 ms for single step latency and batch@32 is ~100-123 requests/sec.
It's a straightforward way to experiment without needing your own GPU. Runs at roughly $3/hr and drops to $0 when idle.
Get the code and instructions here: https://t.co/E3agzrW2hZ
Introducing Xiaomi MiMo-V2.6 — Pro & Flash.
Frontier intelligence, all the modalities, built in public.
🔹 Two omnimodal models, advancing through scaled reinforcement learning
🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks
🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models
🔹 Stronger coding, computer use, 3D reasoning and creative capabilities
🔹 Open model weights, technical report, RL environments and training code
Blog:https://t.co/JGdwC9Ymvj
Happy to announce the full ABC release! We’re also excited that ABC was accepted to CoRL 2026!
Check out our website for code, 400+ hours of sim data on 24 tasks, and 5,850 labeled policy-evaluation episodes.
@arthurallshire@Cinnabar233@ritvik_singh9@redstone_hong@davidrmcall
@xereeto@MatternJustus I suppose it is a failure to read intent. Like the intent was to come up with a compression algorithm and it instead did something very hacky
We’re disclosing HEIF Heist, a months-long investigation into libheif that allowed us to hack OpenAI, Slack, Meta, GitHub Ent, Rails, Next.js, ImageMagick, and many more.
It was literally xkcd #234, one obscure image library beneath a huge number of apps. 🧵
@MTSlive@stanfordnlp@chrmanning This seems true when considering Stanford as an organisation but the reports etc will be made by individuals within that organisation. These individuals are rarely independent from influence from the top labs (huge amount of Stanford PhDs end up in frontier labs)