Top Tweets for #Robometer
Next-Gen Robotic Models & Embodiment: The Dexterity Frontier by DeepReach & Realhand in Palo Alto
Formal Interface Between Evaluation and Generation in Robot Foundation Models — #Robometer × #DreamZero Toward Trajectory-Space Intelligence
The bottleneck in robot foundation models is not planning or data scale, but the lack of a stable interface between evaluation (System 2) and action generation (System 1) under real-world dynamics.
Robometer strengthens explicit evaluation through video-language reward modeling (progress estimation + trajectory preference), but remains outside the policy computation graph. This creates a fundamental optimization interface mismatch with three key challenges:
🔷 Gradient-based optimization is ill-posed: reward models are not directly coupled within the policy’s computation graph, making signal propagation into continuous action spaces inefficient and unstable.
🔷 Co-training is unstable: joint optimization leads to non-stationarity, reward hacking, and severe distribution shift.
🔷 Temporal abstraction mismatch: rewards operate at trajectory-level semantics, while actions occur at high-frequency continuous control, making long-horizon credit assignment fundamentally difficult — far more complex than discrete next-token prediction in LLMs.
DreamZero collapses policy and world model via generative rollout (implicit evaluation through joint video + action prediction), enabling strong zero-shot and cross-embodiment generalization. Yet, despite supporting real-time closed-loop control at 7Hz via autoregressive updates with observation correction, it remains susceptible to residual drift over long horizons and lacks explicit preference grounding.
Crucially, current world-action models remain coarse in temporal and control granularity, limiting precision in fine manipulation.
The primary failure mode is not high-level planning, but decoding abstract intent into physically valid, high-frequency control under contact-rich dynamics (friction, deformation, slip) and partial observability.
This is further constrained by a control frequency gap (∼10–30Hz perception/generation vs. ∼100–1000Hz control) and by latent action representations that are not fully physically grounded.
Data is not merely low-resolution — it is fundamentally misaligned with control (vision vs. force/torque/impedance). For example, grasping involves sub-frame contact events that standard video cannot capture.
The emerging paradigm is reward-guided generative control: world models propose diverse trajectories, while reward models shape them at the trajectory level via re-ranking or guidance (e.g., MPC-style selection or diffusion guidance), avoiding brittle gradient coupling and unstable co-training.
Robotics is not converging to next-token prediction, but to trajectory-space intelligence: generation proposes, evaluation shapes, and control executes — all grounded in closed-loop interaction with real physical dynamics.

Robotics World Model Reading Club discussions on DreamZero with @Voyager_Yu
👉🏻Future events: https://t.co/mc6grS9AiU
DreamDojo vs. DreamZero differ in drift and control. DreamDojo mitigates drift via post-training + real-time distillation (~10.8 FPS) for stable closed-loop control. DreamZero (WAM) relies on closed-loop correction: real observations are injected into the KV cache, so imperfect predictions are corrected online. This constrains observable drift, but latent/belief drift can still accumulate over long horizons.
WAM ≠ plain video diffusion. It combines causal autoregressive rollout (temporal consistency, enabling control) with per-step diffusion (spatial generation). Controllability comes from causality, not diffusion. Sim-control interface: video tokens encode state; actions (explicit or latent) are rolled forward autoregressively; actions are either explicitly predicted or decoded from latent/video tokens via an action head or IDM, then executed—making WAM a unified planner–policy.
Latent action is 2D, egocentric, thus entangled with background motion → noisy labels. This is a trade-off: 2D video scales best; scale > label purity. Disentanglement helps but isn’t required. Pseudo labels can be noisy; with broad coverage, they form a strong initialization. Post-training denoises and adapts across embodiments.
Human–robot gaps (head motion, intrinsics) induce distribution shift in next-frame prediction, biasing toward human-like dynamics. Mitigations: filtering to match embodiment, mixing robot data, or mid-training alignment (EgoScale). Fundamentally a distribution matching problem; if data isn’t invalid, wider coverage helps.
IDM depends on action labels, not raw pixels. If labels exclude background-induced motion, IDM ignores it. WAM→IDM reduction holds only if state (video tokens) is action-sufficient.
Vs. VLA: WAM improves data efficiency and generalization by avoiding high→low mapping and reducing pretrain gap; VLA often needs tuning (e.g., AR decoding). Cost: WAM is slower (denoising all tokens) vs. lightweight VLA heads. Best practice: video backbone + VLM for high-level planning/long horizon.
Video supervision > action supervision because it provides pixel-level spatiotemporal gradients that are action-causal (not just dense), while actions are low-dim/sparse → better sample efficiency and less forgetting. Pure RL is hard for long horizons; diffusion RL scales poorly; within ~1k steps, video signals dominate.
EgoScale adds a mid-training stage for human–robot distribution alignment, requiring slow, high-quality demos; large dynamics gaps limit transfer. Long term, as embodiments converge, transfer eases and data distribution design dominates; short term, startups face mismatch.
Most world models start single-agent, egocentric; multi-view/third-person and ego-4D reduce scalability and are rarely first-order.
👉🏻More fun pics: https://t.co/4k9v8jNx9m

Last Seen Hashtags on Sotwe
latenightFun
Seen from South Africa
nolimit()+filter:native_video
Seen from Italy
anal
monkeyapp #flash
Seen from Germany
NewZealand
Seen from United States
คู่มหาชัย
Seen from Thailand
wariabogor
Seen from Indonesia
SilverbulletDay
Seen from United States
漏奶
Seen from Singapore
กกนใช้แล้ว
Seen from Thailand
Most Popular Users

Elon Musk 
@elonmusk
241M followers

Barack Obama 
@barackobama
119.2M followers

Cristiano Ronaldo 
@cristiano
112M followers

Donald J. Trump 
@realdonaldtrump
111.8M followers

Narendra Modi 
@narendramodi
107.1M followers

Rihanna 
@rihanna
98M followers

NASA 
@nasa
92.2M followers

Justin Bieber 
@justinbieber
91.2M followers

KATY PERRY 
@katyperry
88.4M followers

Taylor Swift 
@taylorswift13
82.3M followers

Lady Gaga 
@ladygaga
73.8M followers

Virat Kohli 
@imvkohli
71.1M followers

Kim Kardashian 
@kimkardashian
70.2M followers

YouTube 
@youtube
68.7M followers

Bill Gates 
@billgates
64.3M followers

Neymar Jr 
@neymarjr
64M followers

The Ellen Show
@theellenshow
62.4M followers

CNN 
@cnn
61.8M followers

Selena Gomez 
@selenagomez
61.5M followers

X 
@x
60.8M followers






