#NiuLai#MiniMaxH3
Prompt 👉
Style: The first 2 seconds imitate a high-fidelity ink-wash-style wuxia film, then abruptly cuts to low-poly 3D animation — a jarring visual split with harsh, clashing color tones. Duration: 10 seconds.
Scene: The opening shot is steeped in classical Chinese aesthetic, with layered mountain ridges receding into the distance. Suddenly the image stutters/glitches, cutting to a patch of rough, low-detail grassland. Several cartoon cows, crudely assembled from basic geometric shapes, stand stiffly like puppets. The lead character, a little yellow calf ("Niu Lai"), opens its rectangular mouth — textured with only a couple of low-res image maps — and, in a flat, emotionless AI-generated voice, shouts loudly: "Mama——!" The background audio is rough and low-fidelity, with an audible electrical hum/noise floor underneath. The camera holds static, capturing this awkward moment with no attempt at cinematic polish.
Rendering plausible video ≠ knowing what the world is doing. The video shell game is a sharp test of whether world models actually track unobserved state — and the results look revealing.
1/
Can Video World Models Track Unobserved World States?
Video world models increasingly look like simulators. But rendering plausible video is not the same as knowing what the world is doing.
We test this with a video shell game. A ball hides under one of five identical cups, which are swapped N times out of sight. Every frame is easy to render. The hidden arrangement is not.
A standard causal DiT trained on five swaps still renders ten swaps at 40.4 dB PSNR, but lifts the wrong cup. LaCT (TTT-KVB w. SwiGLU) tracks the correct final state with similar visual quality.
@jhshin2000 The video shell game is such a clean test — rendering plausible frames is not the same as tracking hidden world state. This is the kind of evaluation world models actually need. Great work!
Open-source video generation just crossed a real threshold: MiniMax-H3 accelerated 75–90× by Video Delta Net, rendering 14s of 768p in 11s. Faster than playback, without giving up quality.
Open-source video generation is now faster than playback without compromising quality.
Introducing Video Delta Net (VDN): hybrid attention for live text-to-video with near-lossless quality.
VDN accelerates Minimax-H3 by 75 - 90 x, generating 14 seconds of 768p video in 11 seconds on 8× NVIDIA B200 GPUs.
Checkpoints + training/inference code + Technical Blog ⬇️
(1/6)
@HaochengXiUCB 11 seconds to generate 14 seconds of 768p video — faster than playback, with near-lossless quality. Hybrid attention like VDN is exactly the architecture shift open-source video gen needed. Huge work!
@notiansans Scene-referred, half-float EXR straight out of a generative tool changes the game for color-managed pipelines. No more round-tripping through LUT hacks. Nice upgrade.
@Diplomeme The walking shots are the giveaway no more — motion coherence here is genuinely impressive. Feels less like "generated video" and more like real footage.
@xuanmingzhangai 2.4T params with 1M context is a serious jump. The post-training focus on Coding & Cowork feels like the right bet — most enterprise pain lives in long, messy multi-step workflows. Congrats on the release!
Local MiniMax H3 deployment made simple — with Codex doing the heavy lifting, running AI video generation with stereo audio on your own GPU becomes a copy-paste job. Great hands-on tutorial!
@yangqing_66 Love this workflow — letting Codex handle the local H3 deployment and then just collecting the results. Tutorials like this make running AI video generation with stereo audio at home genuinely approachable. Looking forward to the updates!
Open source velocity in action: a TensorRT VAE acceleration node for MiniMax H3, benchmarked and published within days of release. Faster local renders for everyone — this is what a healthy ecosystem looks like.
Open Source Drop! Benchmark results for the newly released #MiniMaxH3#TensorRT VAE acceleration node by lihaoyun6 (Sept 1, 2026).
🔗https://t.co/9Fx60JoM5U
(VAE & LoRA Download Links + full prompt are in the replies 👇)
────────────────────────────
Summary: Generation is nearly 30s faster ⬇️, video bitrate is up by 23% ⬆️, and VRAM remains rock-solid under 9.5GB ⬇️. A massive upgrade for 12GB GPUs.
────────────────────────────
Test Setup: RTX 5070 (12GB) / 8s, 192 frames (1344x752) widescreen video
Workflow Config: Based on 8-step PDD Acc sampling
Compared: Native Standard Encoding vs Latest TensorRT VAE
────────────────────────────
#ComfyUI Benchmark Comparison:
1. Video Bitrate:
• Native Standard Encoding: 2,751.0 kbps (2.62 MB)
• Latest TRT VAE: 3,381.5 kbps (3.22 MB) ➔ ⬆️ +22.9% detail bitrate boost
2. Total Generation Time:
• Native Standard Encoding: 438.5s (~7m 18s)
• Latest TRT VAE: 410.0s (~6m 50s) ➔ ⬇️ 28.5s faster overall
(The custom TensorRT .engine cuts pure VAE decoding from 18s down to 5s, a 1.7x speedup ⬆️)
3. Peak VRAM Usage:
• Native Standard Encoding: Spikes past 11.5GB+ during decoding (hitting 12GB limits, prone to crashes)
• Latest TRT VAE: Peak VRAM strictly locked under 9.5GB ⬇️ (2.5GB+ safe buffer headroom ⬆️)
4. Dynamic Range (Pixel Std):
• Native Standard Encoding: 66.08
• Latest TRT VAE: 69.81 ➔ +5.6% contrast & lighting depth improvement
Which visual look do you think is better? 🫴🫴
@Tomw852 The open source ecosystem around MiniMax H3 is moving incredibly fast — a TensorRT VAE acceleration node already benchmarked and shared with download links and prompts. Thanks for putting this together for the community!
MiniMax H3 Max vs Seedance 2.5 in extreme parkour — a close and fascinating matchup. Head-to-head comparison posts like this are exactly what the AI video community needs.
@renoiseai Extreme parkour is the perfect stress test — physics, motion blur and camera tracking under fast movement. Both models hold up impressively well. Thanks for sharing the prompt for reproducibility!
AI video generation faster than you can watch — 5-second renders in under 3 seconds with MiniMax H3 Max on Morphic. The speed barrier is basically gone now.
@morphic A 5-second video generated in under 3 seconds — AI video generation is becoming truly instant. Serving H3 Max at that speed is a serious infrastructure achievement. Congrats on the launch!
H3-World shows how far a strong video foundation model can stretch: from video generation to world modeling with minimal added parameters. Efficient adaptation like this is the trend to watch in 2026.
We built H3 to generate videos. But the team found a "world" inside it. 🌍
"Only 8K samples + 0.199% trainable parameters."
"Directly turning H3’s existing language understanding into character and camera control."...➡️
The most exciting part isn’t just that "H3 can also become a world-model", it’s how little had to be added.
That’s another thing we love about open models. People don’t just improve what you build. They uncover possibilities you didn’t even know were already there.
Huge shoutout to the team! 💜