Awesome-Astra-Embodied-AI:
https://t.co/SKxQTtiMU0
Key paradigms:
• Zero-shot control
• Real2Sim & data rollout
• Astra-driven RL training
• Agentic skill calls
🚀 We introduce Neural Theorizer (NEO) — a new type of world model that learns to theorize the world from observation, without language or LLM supervision.
Selected as an ICML 2026 oral presentation — 0.7% of submitted papers.
The paper asks:
"What does it mean to understand the world and build a world model?"
Today’s world models are often trained to predict the future: the next frame, next latent state, or next observation.
But is prediction enough?
We argue that a world model should be a theory-building system: one that discovers reusable primitives, composes them into executable explanations, and transfers those explanations to novel phenomena.
NEO is our first step toward this vision — a World Theory Model that learns explicit, compositional theories from raw observation.
This work was led by my wonderful students: Doojin Baek*(@doojin_a_baek), Gyubin Lee* (@gyubin0521), Junyeob Baek (@JunyeobB), and Hosung Lee (@HosungLee_).
For more details, take a look at the paper — and if you’re attending ICML, let’s talk there!
📄 arXiv: https://t.co/TGMXLLfzP7
🌐 Project page: https://t.co/aLJywp8rfq
Also- we’re hiring! The team at Apptronik is world-class and care deeply about building a technology that makes people’s lives better.
Two open listings on the controls team:
- Senior Reinforcement Learning Engineer
- Software Engineer - Human Motion Data
- and more!
Flow-based VLAs are powerful, but they struggle with one key question: "How sure am I?"
This work introduces VFD (Velocity Field Disagreement), a principled uncertainty estimator for flow-matching VLAs. By measuring disagreement between ensemble velocity fields, it captures epistemic uncertainty and helps identify where the policy needs more data.
Our post-training pipeline is a substantial redesign from Super.
The core idea: don't rely on stacked RL stages alone. We do SFT, multi-environment RLVR across a huge mix of agentic/reasoning/code/safety environments, then Multi-teacher On-Policy Distillation (MOPD). 10+ domain-specialized teachers, merged into the student via dense token-level guidance on its own rollouts. See Figures below for overview and tech report for all the details. 2/4
Not every latent is reasoning.
We’re releasing Continuous Reasoning for Vision-Language-Action.
Our claim is simple: if a robot policy is truly reasoning, another policy should be able to decode the same thought, verify it, and use it to make better action predictions.
(1/6)
On-Policy Distillation is the most active new research direction being explored in RL for LLMs. Had the chance to discuss how it works with Dwarkesh and why it fits so nicely into large-scale pipelines.
Introducing Vero, the strongest fully open RL recipe for training next-generation visual reasoners.
From charts to spatial to open-ended tasks, Vero sets a new bar.
• sota 8B VLM across 30 benchmarks
• +4.4 avg over four base models (30 evals)
• beats prior RL datasets
🧵👇
The lack of formalism wrt broadcasting in deep learning models annoyed me so much I learned category theory. Weaves, Wires, and Morphisms is now out on arXiv! First step to using the Yoneda lemma to automatically derive fused kernels.
https://t.co/bzR6LeZweZ
We made Muon run up to 2x faster for free!
Introducing Gram Newton-Schulz: a mathematically equivalent but computationally faster Newton-Schulz algorithm for polar decomposition.
Gram Newton-Schulz rewrites Newton-Schulz such that instead of iterating on the expensive rectangular X matrix, we iterate on the small, square, symmetric XX^T Gram matrix to reduce FLOPs. This allows us to make more use of fast symmetric GEMM kernels on Hopper and Blackwell, halving the FLOPs of each of those GEMMs.
Gram Newton-Schulz is a drop-in replacement of Newton-Schulz for your Muon use case: we see validation perplexity preserved within 0.01, and share our (long!) journey stabilizing this algorithm and ensuring that training quality is preserved above all else.
This was a super fun project with @noahamsel, @berlinchen, and @tri_dao that spanned theory, numerical analysis, and ML systems! Blog and codebase linked below 🧵
Our recent findings on World Action Models (WAMs): the core advantage of WAMs is not test-time “imagination” of futures, but the training-time supervision from future video prediction.
We propose Fast-WAM, which makes inference simple, fast, and policy-centric.
someone should train an AI on a blender RL env and use demonstrations from blender experts (blender can actually record all the python api actions!) Train a VLM to make photorealistic scenes
would be hilarious if this works better than raw pixel generation + 3d memory that the top video gen models do now
Robotics as a field faces a large data gap, and one of the most promising directions for solving it is sim-to-real robot learning. We talked with @Stone_Tao about ManiSkill3: a powerful, easy-to-use framework for generalizable sim-to-real learning.
In particular, we learned about how with maniskill + @LeRobotHF you can easily train your own sim-to-real manipulation policies on a low-cost robot!
Co-hosted by @micoolcho and @chris_j_paxton
I've written the full story of Attention Sinks — a technical deep-dive into how the mechanism was developed and how our research ended up being used in OpenAI's new OSS models.
For those interested in the details:
https://t.co/0EAi2KQMMx