15s of MiniMax-H3 video + audio, generated in ~72s on one RTX 5090.
Our baseline took ~26 minutes. We wrote up how LightX2V got it down to 72s ๐
768ร1344, steady-state benchmark. MP4 export and network transfer excluded. https://t.co/0sMSdJfcIQ
๐ค What if various environments became interactive LLM agents?
๐ Introducing Trace2Env: a training-free framework for agentic language world modeling, supporting stateful and long-horizon environment simulation.
๐น๏ธ Below is a โfakeโ Terminal, observations of it are purely generated.
#WorldModels #AIAgents
๐ Excited to release RealtimeWAM! The first one-step asynchronous World Action Model with superior task performance.
โก ~12 ms action generation with a 6B model on a single H100 GPU โ ~25ร faster! โก
โจ Supports any MoT-based WAM (e.g., Fast-WAM and Faster-WAM), with less than 1% success-rate loss on diverse benchmarks (e.g., RoboTwin 2.0 and LIBERO-Plus).
๐งฉ Algorithmโsystem co-design: few-step distillation, blockwise asynchronous inference, and efficient kernelsโall implemented in LightX2V.
๐ Welcome to check it out, try it, and share feedback!
๐ Paper: https://t.co/aE80M7nM9H
๐ค Model weights: https://t.co/qzpOuBNbk1
โญ Code: https://t.co/h6KnAkbkiZ