Native 😃 MiniMax-SOL H3 inference acceleration for ComfyUI
integrating Exact Runtime optimizations with Sana Sol-Attn, rectangular Q/KV attention, and composable support for VDN, Spectrum, Untwisting RoPE, Diff-Aid, and Flow mixed-grid workflows.
👇
https://t.co/WnHA60hlpm
What’s New in DaVinci Resolve 21.1! Watch this video to learn about new features in DaVinci Resolve 21.1, including AI assistant integration, expanded still photography support plus over 100 new tools and controls for editors and colorists!
Minimax H3 running locally.
~65GB of models.
12GB VRAM. RTX 4070 Ti.
100% Text-to-Video.
100% ComfyUI.
Pass 1 → 960×544
Pass 2 → 1920×1088
One continuous loop built from multiple context-connected generations.
No cloud. No API.
CONTROL IS THE KEY.
LOCAL IS THE WAY.
We are sharing a major update to our General World Models efforts. GWM Worlds 2 can generate full, interactive, real-time video simulations in one continuous 720p stream at 24 fps and audio at 48000 Hz, responding to your inputs as you explore.
It is the first world model that generalizes to arbitrary actions rather than a fixed set of actions or just navigation. Dynamic actions are what actually matter for building gaming and interactive experiences, and for training agents in these simulations.
We also introduced WorldPrompt, a way to tell GWM-2 that a world can be split into two kinds of state: what persists and what changes over time. For example: persistent gravity, physics, and light, alongside actions that describe movement, gestures, object interactions, speech, and sound.
WeatherNext 3 is a major breakthrough in how we forecast global weather. ⛅
Developed with @GoogleResearch, the model learns directly from real-world, real-time observations to give more localized highly accurate predictions faster. 🧵
We’re introducing a new capability to our latest Gemini models: agentic video understanding.
This allows developers to process long-form video content with more accuracy, while using up to 88% less tokens.
See how it works 🧵
One cool use for Atlas is creating "bullet time"-style effects using only a few cameras.
In this case, I was able to use 3 iPhones + my finely honed juggling skillz to synthesize a smooth camera trajectory between each pair of cameras.
Also I layered in an AI-generated beats because that's just the world we live in now. Deal with it.
We share H3-World 🌍
The first to turn MiniMax-H3 itself into world model.
No new action module. We directly convert H3’s pretrained language understanding into world control.
Only 8K samples + 0.199% Trainable Params.
📄 https://t.co/Bn4M4JwEJs
💻 https://t.co/sT44pUHAl8
Introducing Atlas:
The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time.
I'm particularly excited by Atlas's ability to reconstruct scenes from a very small number of input image -- was able to create this flythrough of London's Natural History Museum by combining 3 input images I found from completely separate sources on Google Images
Today, we're sharing new research on Solaris, our first Interface World Model.
Solaris is a new kind of operating system that generates interactive interfaces frame by frame, in real time, with no code. We find that Solaris outperforms frontier LLMs when generating new interfaces, across structural similarity and information retention. Read more and request early access at the link below.
Trellis.2 and Pixal3D are now native in ComfyUI core
No custom nodes. No CUDA extensions. No PyTorch downgrade. No non-commercial license traps.
Plus a rebuilt 3D pipeline: Load/Preview/Save 3D nodes, mesh post-processing, and PBR texturing that bakes normals + AO for a full material set.
Runs on consumer hardware. Free to use commercially.
Click the link below to run the workflow and learn more ⬇️
SAM 3D Body is open-source, so it can be run locally or on Comfy Cloud.
Extract motion from a recording → bring it into Blender → control the camera → use the result with Seedance 2.5
More control over body motion and camera movement.
To try this workflows, links below 👇