🚀 Introducing Streaming-WAM: Streaming Your World-Action Model for Real-Time Robot Manipulation.
🧠Streaming-WAM lets the robot act while the WAM thinks ahead, predicting the future during execution through action-conditioned world prediction.
🌐 Project: https://t.co/jdmWklV3AJ
🦿 Real-robot evaluation
On a single RTX 5090 at 30 Hz, Streaming-WAM achieves 90.0% success (27/30), cuts Chunk Time from 682.1 ms to 122.62 ms (5.6× faster), and reduces rollout time from 90 s to 38 s (2.4× faster).
🚀 Introducing Streaming-WAM: Streaming Your World-Action Model for Real-Time Robot Manipulation.
🧠Streaming-WAM lets the robot act while the WAM thinks ahead, predicting the future during execution through action-conditioned world prediction.
🌐 Project: https://t.co/jdmWklV3AJ
Streaming-WAM also transfers to RoboCasa, showing the approach is not tied to a single WAM family.
It reaches 75.35% average success, comparable to X-WAM at 75.42%, while reducing Chunk Time by 3.2× and Total Time by 1.8×.
On RoboTwin 2.0, Streaming-WAM reaches 91.74% total success, improving over FastWAM-Joint at 87.00%.
At the same time, it delivers 12.0× faster Chunk Time and 1.4× faster Total Time—speeding up control without sacrificing task performance.
On LIBERO, Streaming-WAM reaches 98.20% average success while delivering 12.0× faster Chunk Time and up to 3.0× faster Total Time than FastWAM.
Without action conditioning, success drops to 96.25%—showing the value of aligning prediction with ongoing actions.
⚡ Faster generation alone is not enough.
Streaming-WAM moves prediction into the robot’s execution window. Across LIBERO, RoboTwin 2.0, and RoboCasa, it reduces both Chunk Time and end-to-end rollout time—turning model speed into faster robot execution.
🔄 Streaming-WAM overlaps each Stream Update with the execution of the current action chunk. Shared actions provide temporal continuity across adjacent chunks and condition the next visual prediction, aligning world prediction with the robot’s ongoing motion.
Robots that continuously perceive contact changes mid-action — introducing TacForcing, a streaming action generation framework built for contact-rich dexterous manipulation tasks.
Special thanks to Sharpa Robotics for their systematic support. Paper releasing soon — stay tuned!
Robots that continuously perceive contact changes mid-action — introducing TacForcing, a streaming action generation framework built for contact-rich dexterous manipulation tasks.
Special thanks to Sharpa Robotics for their systematic support. Paper releasing soon — stay tuned!
Wave Forcing is now open source 🌊 To our knowledge, it is the world’s fastest streaming video generation system: 117.7 E2E FPS (126.5 steady-state), 7.9× faster on 8×H200, and ~7.5× at 14B. Code + preview: https://t.co/gHZzsq4ZyZ #videogeneration
Comparison with commercial models
Compare with Kling 2.5 Turbo, Hailuo 2.3, Vidu Q3, and Wan 2.7: CineMobile keeps the subject sharp and the camera path stable across the full sequence, while several commercial models drift or blow out the scene at later frames.
CineMobile: On-Device Image-to-Video Generation for Structured Camera Motion
📱 We introduce CineMobile, an efficient on-device image-to-video generation framework for structured camera-motion videos.
arxiv: https://t.co/2IN38lTLHr
VBench quantitative comparison
At 1.2B params and only 4 steps, CineMobile scores 88–89 on VBench across bullet time, dolly zoom, and slow motion — within ~0.5-0.9 points of the 14B, 20-step Wan2.1 teacher. 10% of the parameters, 20% of the steps, nearly all of the quality.
Qualitative results across effects
Sniper stakeouts, guzheng by the lake, street style, motocross jumps, hurdles — CineMobile holds subject identity and camera trajectory steady across wildly different scenes and motion effects, all generated on-device in 4 denoising steps.
Method: SFT warm-up + step distillation pipeline
How we got there: depth-pruned DiT → supervised fine-tuning warm-up → adversarial 4-step distillation with a real/fake score pair (DMD) plus a discriminator feeding GRPO rewards.
Efficiency: 40.11x speedup
From 97 seconds to 2.4 seconds. We took a 14B teacher DiT and compressed the whole pipeline — pruning, 4-step distillation, hybrid quantization — down to a 1.2B model, while running at ~20s/step on a MediaTek Dimensity 8400 phone.