🚀 MiniMax H3 Super Acceleration in Sol-Engine🤩
We pushed H3 far beyond our previous 3–4× optimization regime — reaching 22.2× speedup for 5s video and 27.7× for 10s video vs. the published SGLang baseline.
On a single NVIDIA GB200:
• 5s 768p: 152.3s → 6.85s (22.2×)
• 10s 768p: 414.1s → 14.93s (27.7×)
But the more interesting part may be what this means economically.
Using MiniMax’s published H3 API price as a reference, we translate inference speed directly into production economics. Under an ideal fully utilized GB200 scenario, Sol-Super can serve about 525 five-second videos/hour, corresponding to roughly $210/hour of output value at the reference API price.
Assuming $5.50/GPU-hour, that implies a 97%+ GPU-only gross margin in the idealized model.
At full utilization, one GB200 could produce:
• 12.6K 5s videos/day
• 378K videos/month
• equivalent to 525 hours of finished video per month
For us, this is the bigger point of inference optimization: a 20×+ speedup does not just reduce latency — it can fundamentally change the unit economics, serving capacity, and viable business models of video generation.
🔗 https://t.co/dI7uBD1dSo
🚀 LTX-2.5 × Sol-Engine — from data center to desktop.
Just one day after the release of the new open-source LTX-2.5, we’ve brought it into Sol-Engine and optimized it across a wide range of deployment settings:
⚡ GB200 — high-performance data center inference
💻 DGX Spark — compact desktop AI deployment
🎮 GeForce RTX 5090 — consumer Blackwell GPU
LTX-2.5 comes with a more complex multi-stage inference pipeline and multiple generation configurations. Sol-Engine automatically adapts its full-stack optimization recipe across these different workloads and hardware targets — combining kernel optimization, attention acceleration, caching, and system-level optimization where they matter most.
Really excited to see another strong open video generation model released to the community — and to make it faster and easier to run everywhere from GB200 to desktop GPUs. 🚀
🔗 https://t.co/j0NZPddofC
🚀 MiniMax H3, accelerated on Day 1 with Sol Engine!
Within just 4.5 hours, our agent-native Sol Video Inference Engine achieved:
⚡ 3.95× end-to-end speedup over Diffusers
⚡ 2.80× speedup over SGLang
🎬 8× NVIDIA GB200, 1344×768, 24 FPS, 124 frames
The acceleration combines kernel fusion and graph capture, cross-step caching, and training-free sparse attention powered by Sol-Attn—with no distillation, fine-tuning, LoRA, or offline calibration.
We are especially excited to see a powerful open-weight model like MiniMax H3 released to the community. Open models are essential for pushing video generation research, systems optimization, and real-world deployment forward.
We hope this is the beginning of a much more vibrant open-source video generation ecosystem—and Sol Engine will keep working to make the latest models faster and easier to deploy from day one.
🔗https://t.co/XpESoX5OnU
We are starting Intent Lab, building an autonomous team we call "fleet" that turns intent into production software. Today we are sharing some early results: the fastest GLM5.2 inference engine, one shot database creation, and a fully verified agent filesystem.
https://t.co/CtwtypE0rH
Coming soon on Sol-engine: broader compatibility for video/image generative models + cutting-edge optimizations such as multi-GPU distributed inference.
Announcing NVIDIA Nemotron 3 Super!
💚120B-12A Hybrid SSM Latent MoE, designed for Blackwell
💚36 on AAIndex v4
💚up to 2.2X faster than GPT-OSS-120B in FP4
💚Open data, open recipe, open weights
Models, Tech report, etc. here:
https://t.co/CAYpP1iK3i
And yes, Ultra is coming!
@si_pbc fascinating work on the context compression part! Wondering if the model would be able to know what is a good quality action at system 2 level. eg. when to raise/fall on playing card games 😀
> Our video encoder encodes nearly 2 hours of video in the same number of tokens—that’s 50x more token-efficient than the previous state-of-the-art and 100x more token-efficient than OpenAI’s encoder.
Here comes the new pied piper 😍
Computer use models shouldn't learn from screenshots.
We built a new foundation model that learns from video like humans do. FDM-1 can construct a gear in Blender, find software bugs, and even drive a real car through San Francisco using arrow keys.
@alightinastorm haha sure, full-time job is building data centers worldwide at nvida, on the side big fan on gaming and using AI to play games like hearthstone battleground. 🤣