Introducing WorldCrafter: a world model with persistent memory—not rigid 3D geometry, not just latent memory, but a “soft skeleton” of 3D-aware latents. The result: striking consistency as you explore, look away, and return—even in dynamic scenes.
🚀 Thrilled to share our work **CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation**, selected as a **CVPR 2026 Highlight**! 🏆
Video generation currently faces a dilemma:
❌ **2D Diffusion**: Lacks spatial consistency in large camera moves.
❌ **Explicit 3D**: Complex, slow, and prone to reconstruction artifacts.
We bridge this gap by using **Implicit 3D as a Spatial Anchor**, using this representation as **context** into the video generation model. ⚓️
Through **Scene-Decoupled Diffusion**, we represent the static environment via implicit 3D features, decoupling scene priors from dynamic motion. This enables unprecedented scene consistency and precise control over complex scenes, camera paths, and characters. 🎬
CineScene is a versatile toolkit for the future of filmmaking & World Models:
📍 **Virtual Stage**: High-fidelity, consistent environments for virtual production.
🎬 **Scene Blocking**: Directorial control over scene background, camera paths, and prompt-driven foreground character dynamics.
🌍 **World Simulators**: A potential step towards stable, consistent world modeling.
We are also excited to open-source the **Scene-Decoupled Video Dataset**, a large-scale, high-quality collection to empower the community! 🎁
🔗 Project: [https://t.co/ENjoQDdQQf](https://t.co/6tnZY9IbSz)
📊 Dataset: [https://t.co/h7PVoVR97Z](https://t.co/cAWp3P7E0l)
📄 ArXiv: [https://t.co/2c7e38rCCZ](https://t.co/bs5jXAucpc)
Huge thanks to the amazing co-authors! 🙏 @KaiyiHUANG84276@xinntao @yukun6414 @yujiwenHK@jianhongbai@lin_zinan@FiNingm@wanfufeng
\#CVPR2026 #AIvideo #GenerativeAI #FilmGeneration #WorldModels #ComputerVision @CVPR
Excited to attend SIGGRAPH Asia 2025!🎉
Our work "OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion" will be presented at:
🕒1:32–1:43PM (GMT+8), 16 Dec 📍Meeting Room S421, Level 4
Feel free to stop by!
#SIGGRAPHAsia2025#HongKong
(3/3) such as geometry, intrinsic properties, and lighting. Once this is achieved, the boundary between 2D and 3D will be blurred, and whether to generate 2D contents or 3D assets will simply depend on user preference.
(1/3) Excited that OmniX is gaining attention. I would like to share some thoughts behind this work:
The remarkable progress in video generation has caused some "panic" within the 3D research community, namely, do we still need explicit 3D?
Introducing OmniX — a family of flow matching models for panorama generation, perception, and completion, enabling the automatic creation of graphics-ready 3D scenes:
🏠Website: https://t.co/NIj4VUwhH7
📄Paper: https://t.co/2L185nB3Wz
💻Code: https://t.co/bxlr21Lfum
(2/3) I think what we should really care about is, how well the model understands the world? Generating images or frames is just the beginning. A truly intelligent vision model goes beyond just RGB pixels; it understands the underlying factors that makes those pixels—
𝐊𝐞𝐲 𝐜𝐨𝐧𝐭𝐫𝐢𝐛𝐮𝐭𝐢𝐨𝐧𝐬:
- Unified framework for visual generation, perception, completion
- Automatic pipeline for generating PBR-ready 3D scene assets
- High-quality synthetic panorama dataset with geometric and intrinsic annotations
Introducing OmniX — a family of flow matching models for panorama generation, perception, and completion, enabling the automatic creation of graphics-ready 3D scenes:
🏠Website: https://t.co/NIj4VUwhH7
📄Paper: https://t.co/2L185nB3Wz
💻Code: https://t.co/bxlr21Lfum
Our part-aware 3D generation work, OmniPart, is accepted by Siggraph Asia 2025. Code and model released!
Paper: https://t.co/vEAyV5kqD2
Project page: https://t.co/ovnAysSa7I
Code: https://t.co/dToyRki7R8
Demo: https://t.co/9gcBmo2NdP
🎉Our paper DreamCube is accepted to #ICCV2025 ! Thank @_akhaliq for sharing our work!
Project page: https://t.co/tYKCbOvIA7
Code: https://t.co/EQMueFNmF3
Model: https://t.co/P5L3NHgF3i
Video: https://t.co/D8C837yi8d
Special thanks to my co-authors: @XihuiLiu@KaiyiHUANG84276
Excited to release SAMPart3D: Segment Any Part in 3D Objects. It supports zero-shot 3D part segmentation at multiple granularities on diverse data such as Objaverse. Code available!
https://t.co/L3ms8Qp0kr
https://t.co/kLHKWUShFJ
https://t.co/rl3p4CJj5R
https://t.co/2ARvI46z7C
Introducing The Matrix --- a foundation world model for generating infinite-length, hyper-realistic videos with real-time, frame-level control:
- Infinite-length video generation
- 720p high-quality rendering
- Real-time, frame-level control at 16 FPS
- Generalization to real-world video control
🔗Blog: https://t.co/ODmne5rEOu
📄Paper: https://t.co/NgDDGXNlMx
💻Code & Playable Demo: Coming soon!
Key Innovation: A brand new technique called the shift-window denoise process model, enabling auto-regressive generation for diffusion and consistency models in real-time.
Special thanks to project leader Ruili Feng and the entire Matrix team for their dedication and hard work over the year-long project.
🔥New feature of our work DreamWaltz-G! Now we can reenact arbitrary in-the-wild human videos with our generated avatars!
Project page: https://t.co/KhHrltG1l8
Github repo (code avaliable): https://t.co/GKoLCO0e3E
DreamWaltz-G
Expressive 3D Gaussian Avatars from Skeleton-Guided 2D Diffusion
discuss: https://t.co/NOqqk2DoHa
Leveraging pretrained 2D diffusion models and score distillation sampling (SDS), recent methods have shown promising results for text-to-3D avatar generation. However, generating high-quality 3D avatars capable of expressive animation remains challenging. In this work, we present DreamWaltz-G, a novel learning framework for animatable 3D avatar generation from text. The core of this framework lies in Skeleton-guided Score Distillation and Hybrid 3D Gaussian Avatar representation. Specifically, the proposed skeleton-guided score distillation integrates skeleton controls from 3D human templates into 2D diffusion models, enhancing the consistency of SDS supervision in terms of view and human pose. This facilitates the generation of high-quality avatars, mitigating issues such as multiple faces, extra limbs, and blurring. The proposed hybrid 3D Gaussian avatar representation builds on the efficient 3D Gaussians, combining neural implicit fields and parameterized 3D meshes to enable real-time rendering, stable SDS optimization, and expressive animation. Extensive experiments demonstrate that DreamWaltz-G is highly effective in generating and animating 3D avatars, outperforming existing methods in both visual quality and animation expressiveness. Our framework further supports diverse applications, including human video reenactment and multi-subject scene composition.
DreamWaltz-G
Expressive 3D Gaussian Avatars from Skeleton-Guided 2D Diffusion
discuss: https://t.co/NOqqk2DoHa
Leveraging pretrained 2D diffusion models and score distillation sampling (SDS), recent methods have shown promising results for text-to-3D avatar generation. However, generating high-quality 3D avatars capable of expressive animation remains challenging. In this work, we present DreamWaltz-G, a novel learning framework for animatable 3D avatar generation from text. The core of this framework lies in Skeleton-guided Score Distillation and Hybrid 3D Gaussian Avatar representation. Specifically, the proposed skeleton-guided score distillation integrates skeleton controls from 3D human templates into 2D diffusion models, enhancing the consistency of SDS supervision in terms of view and human pose. This facilitates the generation of high-quality avatars, mitigating issues such as multiple faces, extra limbs, and blurring. The proposed hybrid 3D Gaussian avatar representation builds on the efficient 3D Gaussians, combining neural implicit fields and parameterized 3D meshes to enable real-time rendering, stable SDS optimization, and expressive animation. Extensive experiments demonstrate that DreamWaltz-G is highly effective in generating and animating 3D avatars, outperforming existing methods in both visual quality and animation expressiveness. Our framework further supports diverse applications, including human video reenactment and multi-subject scene composition.