Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions.
CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks.
With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks.
We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed.
Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size.
📄 Blog: https://t.co/zwi9JOHKGx
💻 Code: https://t.co/rsHRYCGR8I
🗣️ Discord: https://t.co/Uqtdefvo3J
🤗 Data & Models: https://t.co/wdSWGGO3hu
More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
@YuXiang_IRVL@SteveTod1998 In practice, feeding in the entire video introduces a lot of redundancy, so we use keyframes generated by GPT itself as the context.
[1/2] GPT-6 Astra might be the craziest robot policy I’ve seen. 🤯
Give it ONE demo per task, and it can control a real robot to unscrew a bottle cap, insert a plug into a power strip, clear obstacles, and even imitate a human dance.
One context. New behavior. 👇
🚀 GPT-Policy code + new results are live!
Put GPT-6 to work on real robots: tackle tasks zero-shot or learn from just ONE DEMO through in-context learning. 🤖
💻 [Code](https://t.co/j3B9QRpQ3f)
📄 [report](https://t.co/CAg2OmN0Lk)
🎬 [demos](https://t.co/OccIupWiic)
[1/2] GPT-6 Astra might be the craziest robot policy I’ve seen. 🤯
Give it ONE demo per task, and it can control a real robot to unscrew a bottle cap, insert a plug into a power strip, clear obstacles, and even imitate a human dance.
One context. New behavior. 👇
[1/2] GPT-6 Astra might be the craziest robot policy I’ve seen. 🤯
Give it ONE demo per task, and it can control a real robot to unscrew a bottle cap, insert a plug into a power strip, clear obstacles, and even imitate a human dance.
One context. New behavior. 👇
@k7agar Really interesting exploration—we’ve also been testing GPT-6 Astra’s in-context learning on real robots, using just one demonstration to guide new behaviors. https://t.co/lAghJoHEGW
[1/2] GPT-6 Astra might be the craziest robot policy I’ve seen. 🤯
Give it ONE demo per task, and it can control a real robot to unscrew a bottle cap, insert a plug into a power strip, clear obstacles, and even imitate a human dance.
One context. New behavior. 👇
[2/2] GPT-6 Astra brings memory and reasoning to mobile manipulation. 🚀🚀🚀
It learns new behaviors in context from human videos, spots fine details, understands complex scenes, interacts naturally and remembers past interactions.
Embodied intelligence in motion. 👇
[1/2] GPT-6 Astra might be the craziest robot policy I’ve seen. 🤯
Give it ONE demo per task, and it can control a real robot to unscrew a bottle cap, insert a plug into a power strip, clear obstacles, and even imitate a human dance.
One context. New behavior. 👇
What if RL didn’t choose one action at a time?
Reinforcement Learning with Action Chunking
Q-Chunking: run RL directly over short sequences of actions.
Instead of:
aₜ → aₜ₊₁ → aₜ₊₂
the agent chooses:
[aₜ, aₜ₊₁, …, aₜ₊ₖ]
This simple change turns out to be powerful for long-horizon robotic tasks.
🤖 One brain, for any robot. But how does a generalist robot achieve mastery on a specific task?
🚀 Introducing CLIFT (Closed-Loop Iterative Fine-Tuning), a framework that turns Gemini Robotics On-Device into a humanoid specialist using only a managed fine-tuning API. [1/6]
🌐 https://t.co/Uts9KDEY8s
📄 https://t.co/Pjvjcr1rSN
📚 @saturdayrobotic Robotics & World Models Reading Club 18 Recap: keynote @Guanming717 (@gen_instinct), @ffeng01 (@UCSanDiego, @AetherLab_AI).
Scaling video generation alone will never produce reliable embodied intelligence. The next generation of world models will be defined by causality, latent actions, verification and task alignment — not bigger diffusion models.
DreamZero, ImageWAM, FastWAM, LeWorldModel, TC-WM, WAV and Unified Latent Action Models all point in the same direction.
DreamZero jointly models video and actions instead of treating control as an afterthought. Frames are VAE-encoded into latents that enter Causal DiT blocks together with action noise and proprioception+language. Training uses joint video-action flow matching under teacher forcing. Inference keeps a KV cache, samples action chunks autoregressively from real observations, executes them asynchronously in closed loop, and decodes future frames only on demand. The joint prediction is deliberately factorized into a video term plus an inverse-dynamics term.
ImageWAM turns image-editing foundation models into world-action models. A frozen LLM encodes the language instruction; the current observation is VAE-encoded with noise and fed into a tunable image-editing backbone. A lightweight tunable Action Expert then reads the edited future observation and outputs the full action sequence, leaving the powerful visual priors frozen.
FastWAM asks a blunt efficiency question: does action prediction really need to attend to future video? Training still mixes three masking regimes — joint video-action denoising, video denoising plus inverse dynamics, and action conditioned only on current observations. At inference the winning strategy predicts actions without ever attending to future video tokens, treating video prediction merely as an auxiliary training signal. Joint training still helps representation learning; decoupling at inference delivers much faster real-time control.
LeWorldModel learns planning-oriented latent dynamics instead of pixels. An encoder maps observations to latents; a predictor rolls those latents forward for many steps under actions; a cost module compares the trajectory against a goal latent. An auxiliary regularizer projects the latents onto random univariate directions and forces the distribution toward normality, producing a compact space directly usable for model-based control.
“Is latent all you need?” The answer is nuanced. Asymmetric denoising, sparse future imagination and dense action refinement let the model dream only when necessary. Heatmaps show predicted action hotspots tightly aligned with objects and robot end-effectors, proving that latent representations can focus computation on causally relevant regions while ignoring irrelevant background.
Classic video generators fail three basic tests: precise action control, object consistency and physical consistency. Choppy knife-cutting sequences and physically implausible 3D navigation are not edge cases — they are symptoms of missing causal structure.
Three interlocking questions therefore dominate:
Can we recover the hidden state factors behind observations?
Can we recover the latent actions that actually drive system dynamics?
How do we use those representations to build self-improving systems?
Hu & Shum (2013) give a concrete answer to the first: under mild assumptions a short temporal block of observed trajectories is already sufficient to recover the latent context up to an invertible transformation. When clean latent factors are absent, stochastic residuals still matter and the model becomes a pseudo-Bayesian filter. Empirically this approach ranks at or near the top on Kitchen (~70 %), Maze2D (~160 %), Walker (~118 %), LIBERO-object (~93 %) and LIBERO-long (~62 %), beating DD, DF, LDCQ, Diffuser and DP.
Latent-factor identification proceeds by feeding raw trajectories through a sequential encoder to obtain latents, then a sequential decoder that reconstructs the trajectories. These latents are subsequently used by Ada-Diffuser-Planning and Ada-Diffuser-Policy modules that inherit the same causal inductive bias (diffusion I/O, masks, inverse dynamics).
Task-centric world models go further. Five architectural families are contrasted; the winning TC-WM injects a task signal that aligns and splits latents during training. History is per-patch encoded and aligned with the current embedding; actions condition a latent-dynamics transformer (positional embedding + transformer + per-patch decode) that predicts future embeddings and proprioception. A trainable encoder-decoder sits on top of frozen vision foundation models. Planning uses either cross-entropy method elite selection or latent diffusion guided by inverse dynamics. The same models outperform TD-MPC2, DreamerV3, MuZero and DINO-WM on CEM (Maze 100 %, Wall 100 %, Push-T 92 %, Cheetah 292) and LDP (Lift 60+ %, Can 62 %, Square 40 %, Hopper 46) while producing markedly more physical contact and fewer floating artifacts than Cosmos3-Nano.
Identifiable world models can also serve as verifiers. Generative models produce blurry objects, blurry arm motion and interaction hallucinations. Identifiable latents are fed through an Ada-Diffuser; Temporal Difference Verification then applies gradient guidance that penalizes physically inconsistent regions, pushing trajectories onto the valid dynamic manifold for reliable test-time guidance.
Unified Latent Action Models abandon robot-specific action spaces. Instead of recovering every hidden factor they recover only the shared latent actions that explain how the world changes across embodiments (partial identifiability, Kong & Xie 2022). A video foundation encoder produces latents; an inverse-dynamics stack of spatio-temporal ViTs and a forward stack of DiTs are trained inside a diffusion process with AdaLN, timestep conditioning, embodiment-ID classification and gradient reversal. At inference a single frame yields a transferable latent action that can be executed zero-shot on a new body.
Zero-shot transfer experiments confirm the point: LAD collapses into incoherent frames, LVP produces plausible video that ignores the source actions, while the latent-action approach successfully executes the same intrinsic behavior on the target embodiment.
Self-supervised skill learning needs neither demonstrations nor rewards. A skill embedding is contrastively aligned with state-action pairs: positives that move toward a goal are pulled together, negatives are pushed away (InfoNCE). Interaction-weighted resampling from the replay buffer focuses learning on meaningful interactions. Locomotion often follows linear-Gaussian dynamics; manipulation exhibits discontinuous “jumpy modes”. Local causal structure learning is the key that bridges these islands.
Hallucination has a clear data-allocation root. Policies only need the narrow distribution of optimal actions; world models need the broad distribution of suboptimal and exploratory actions. Action-free internet video is abundant for learning general dynamics, yet action-labeled robot data remains scarce. On-policy collection (Sailor, VLAW, World-VLA) limits generality; information-maximizing exploration runs into the information paradox where model uncertainty does not correlate with useful learning progress.
The practical learning framework therefore combines an adaptive curriculum that generates increasingly difficult tasks, a skill library of latent skill variables, and MIST-style masking: states are randomly masked, the model must maximize mutual information between the masked states and its predictions of observation and reward. This forces compact, task-relevant, causally structured representations.
WAV (World Action Verifier, Liu et al., arXiv:2604.01985) reframes the remaining problem as verification. Three core ideas: (1) semi-supervised setting that exploits more data, (2) decomposed verification that replaces one hard check with two easier ones (state plausibility + action reachability), (3) goal-oriented cycle consistency that couples inverse and forward models. Diverse subgoals are sampled from action-free video; sparse inverse dynamics extracts action-relevant features; the agent actively seeks the hardest plausible subgoal. On MiniGrid the method achieves the highest correlation with true error, lowest prediction error and strongest action following (near Oracle). On real robots it adapts to novel appearance, novel objects and shifts in policy optimality with only 200 target samples.
The high-level pipeline is now clear: Environment (with exploration that discovers novel representations) → Representation Learning (compress and distill from foundation models) → Structure Learning (incorporate domain knowledge and task feedback) → Decision-Making (adaptivity, compositionality, controllability).
Scaling alone makes world models broader but not necessarily physical or controllable. Causal hidden representations identify what the world is; latent actions identify where the control signals come from; simple causal principles enable models that are more physical, controllable and self-improving.
The ultimate question left open by all of this work remains: can we build a world model that not only dreams of the world, but also lets agents interact inside it, experiment, discover goals, and continuously refine the model itself?