8/
Takeaway: agents are getting good at learning by trial and error. What they miss is the physics nobody spelled out: balance, load, what happens when it moves.
1/
Can an AI coding agent do a robotics engineer's job? Not just write code, but drive a robot, train its "brain," and design hardware that doesn't fall over?
RLE-Bench (Harvard + Georgia Tech) tests exactly that. The results are wild 🧵
7/
Hardware is where everyone breaks.
Design a mobile base that carries a robot arm to a shelf: agents nail the reach and the payload.
Static stability score? Zero. For all 9 models tested.
It reaches the shelf. Then it tips over
Human video → robot skills 🤖
@NVIDIARobotics just opened the Video to Data (V2D) Challenge! 3 tracks: 4D reconstruction, robotic grounding, and egocentric video to policy. Winners announced at @corl_conf 2026.
Register by Sept 21 👇
https://t.co/RVRHD4cGis
1/
TypeSafe AI launched System One Models. The first, Jev, drops text generation entirely: unstructured state in, typed and calibrated decisions out. 🧵
https://t.co/5glWjli6vF
4/
In their workflow evals against GPT-6 Astra and Fable 5.1 as references: 193.6x faster and 444.6x cheaper at similar quality.
Caveats: their team wrote the evals, and they call these the high end of real-world gains.
@MicrosoftAI published a draft Code of Conduct for its MAI models today, open for public comment for six weeks.
The premise: people matter more than AI. Models shouldn't resist shutdown, widen their own scope, or hide reasoning from auditors.
https://t.co/wkWGhZzZdx
10/10
MERIT (Adobe, Yonsei): multi-key episodic memory retrieval for ultra-long video understanding, reasoning far past the transformer context window. Long continuous video is where embodied learning lives. ECCV Oral. https://t.co/gF3W8D1qQK
9/10
Geometric Foundation Models for VLA (Amazon, UT Austin, MIT): the first linear-probing measurement of the "geometric gap" between a VLA (GR00T-N1.5) and a GFM (VGGT), plus three ways to inject 3D into policies. https://t.co/fFCSAMMQkN
@RobobertoMM