DLSS 5, is it you? 🤔
⬇️
RGBX-Next: Towards Realistic Generative Rendering from G-Buffers
NVIDIA:
Zheng Zeng, Marco Salvi, Lifan Wu, Jan Novák, Daqi Lin, Saeed Hadadan, Yichen Sheng, Robert Pottorff, Shiqiu Liu, Ravi Ramamoorthi, Ling-Qi Yan, Miloš Hašan
https://t.co/cpoUQ4s4PA
Abstract:
Diffusion models have achieved impressive results in image, video, and streaming generation. However, compared to traditional 3D rendering, they still lack precise control over the generated output. We believe a viable path forward is to use generative models as learned renderers conditioned on traditionally rendered G-buffers. We introduce RGBX-Next, a unified generative framework for forward and inverse rendering, which allows estimating G-buffers from images, videos, and streams, and rendering realistic images, videos, and streams from G-buffers. Our key contribution is a general recipe for finetuning diffusion transformer (DiT) models into generative forward and inverse renderers. We show that the resulting models achieve high quality in both realistic generative rendering and intrinsic decomposition. We will make all our models publicly available. We believe that the design principles presented in this paper will benefit future research on controllable generative forward and inverse rendering.
train gaussian splatting straight in the browser with Splat.js. SfM included, open source, and MIT Licensed.
Video from @PixelBilly and @arrival_space_
MiniMax H3 open weights are out, with the nodes already in ComfyUI.
One Blender blockout, run against Seedance 2.0 at a fraction of the cost per generation, leaving more room for experimentation and iteration.
To try this workflow, link below 👇
MiniMax H3 is native in ComfyUI, day zero with the weights.
Model highlights:
→ Text-to-video, prompt only
→ Image-to-video
→ First-and-last-frame - control the opening frame, the closing frame, or both
→ Reference-to-video - carry a subject, a motion, or a voice through the clip
→ In-place editing - modify a shot you already have instead of regenerating it
The capability @MiniMax_AI leads with is multimodal context understanding, and it's what collapses those five tasks into one model.
H3 takes images, audio, and video together and resolves them against a prompt that explains how they relate.
Download the workflows & learn more in our blog 👇
What if you could turn any number of photos (3, 8, 15, or even 60) into one clean 3D surface (pts & mesh) with Flow Matching?
Check out our new work, Surflo: Consistent 3D Surface Flow Model with Global State. 🧵
1/n
🔗https://t.co/lBcJRgpfdg
Modality Forcing - turns FLUX.2.klein into 3D-aware models.
- joint RGB-D, I2D, and D2I synthesis in one model
- pixel-space depth tokenization
- preserves T2I quality via self-distillation
https://t.co/1cODlGjpl4
This one is really neat. Not only can it repose according to a depth, but it preserves the general vibe and feel of the reference image as well.
Gosh, suddenly a lot of neat Flux Klein 9B loras to try ...
https://t.co/u6b1uEUba3
MAMMA!
You can do the Matrix Bullet Time effect with only 4 cameras.
It's a markerless motion capture system that turns standard multi-view video into accurate 3D digital models.
No expensive hardware and physical markers. Tracks complex human interactions like dancing/hugging.
https://t.co/CQewmdvOTS
Pixal3D-ComfyUI
Nice, full-featured nodes for Pixel-Aligned 3D Generation.
- GLB export
- FlashAttention support
- manual camera control
- native model management
https://t.co/2SMpC2azxK
We’ve released the official Blender add-on for ZOZO’s Contact Solver, our open-source physics solver developed at ZOZO, Inc. Blender users, please take a look!
https://t.co/t2kmcaCP2i
I trained an ACEStep 1.5 XL LoRA on "some obscure 60s English rock band". Then I wrote a song about LoRA training and had them play it. Absolutely wonderful experience. I still have some UI work before I can make training public in AI Toolkit, but working on it as fast as I can.
ACE-Step 1.5 XL is open-source! 4B DiT decoder, three variants. ����
Beats Suno v5 on SongEval (4.79 vs 4.72), Style Align 47.9 — #1 across all models tested.
From 12GB VRAM (INT8 offload) to 24GB full quality. MIT license. Commercial-safe training data.
Three variants, three use cases:
✅ XL Base — all tasks (extract, Lego, completion), high diversity, best for fine-tuning
✅ XL SFT — highest audio quality, CFG guidance scale control
✅ XL Turbo — 8-step distilled, fastest, no CFG (early release)
All compatible with LM 0.6B / 1.7B / 4B.
🤖 Base: https://t.co/pmxcjpR4jY
🤖 SFT: https://t.co/8UKuVNioVM
🤖 Turbo: https://t.co/sifKHQmhtY
What is this high quality of LFS new strategy "MRNF"? Unbelievable. I didn't even do any cleanup. Maybe I no longer need Postshot !!
--
This 3DGS was made with OSMO360(only 5min walk), Metashape, my tool and LFS(LichtFeld Studio).
--
See my instructions :) https://t.co/I0h3WdngKY
Release time! The Grove 2.3 grows better 3D trees with much less memory, in Blender, in your browser, and now also in Houdini Indie! — see what’s new at https://t.co/bVgeZxKLm3 or go play with the online demo at https://t.co/kPKKWJVha7 - have fun growing!
Bytedance just dropped a new open source image model.
Better than Qwen & Z-image.
It's autoregressive, so it has a much better world understanding like Nanobanana & GPT-Image.
https://t.co/8Vvmp5daQH