Thanks for sharing @janusch_patas !
TL;DR - Sparse Views to Dynamic 3D generation.
Key Insight - Slowing a video (magnifying temporally) improves 3D generation quality.
In-2-4D: Inbetweening from Two Single-View Images to 4D Generation
Contributions:
• To the best of our knowledge, In-2-4D is the first method for generative 4D inbetweening over two distant monocular frames spanning arbitrary motions.
• Our novel hierarchical approach breaks the complex inbetweening into a series of simpler motion estimations via video, followed by 4D (i.e., dynamic 3DGS) generation.
• To generate smooth 3D object and motion transitions, we further optimize the 3D trajectories using a bottom-up merging strategy with smoothing regularization.
• We contribute a new 4D interpolation benchmark, I4D-15, on challenging object motions and real-world scenes.
Here's my talk from the CVPR 2026 "Bitter Lessons" workshop earlier this summer. I've split it up into parts for the sake of discussion.
Part 1: in which I wax poetic about my youth and force the audience to look at my dissertation results.
new post on harness engineering for AI self-improvement: https://t.co/ZYvGfVs61k
It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter models keeps harnesses simple.
Even when many harness improvement get eventually internalized into core model, the need to specify goals and context will not disappear.
طلبت من ChatGPT صورة لـ 7 تفاحات، سواها بس كانت 8.
مشكلة معروفة نماذج ال generation تخطئ في الأعداد الدقيقة،والـ VLMs تكتفي بتقدير تقريبي في المشاهد الكثيفة.
حل ABACUS يعالج الأمرين في unified model واحد 3B، يجمع counting و count-faithful generation، مع self-correction دون Labeling Data.
⚠️ preprint
🔗 https://t.co/tL1o2N7gOk
I’m excited to share that I’ve joined the @wayve_ai Labs team in Vancouver as a Principal Scientist!
https://t.co/xdBMlGt9fB
I was drawn to Wayve Labs for two reasons: 🧵⬇️
📢📢📢introducing 𝐏𝐨𝐰𝐞𝐫 𝐅𝐨𝐚𝐦
A 3D representation that can be ray traced or rasterized in real time, with NO COMPROMISE in quality.
- Project: https://t.co/LkmVQjkIt2
- arXiv: https://t.co/TtMbyKrvrp
Rasterized at 3DGS-class FPS
Ray traced at Radiant Foam speeds
Shared attention is a powerful way to transfer style in diffusion models: let tokens attend to a reference image, and the model can pick up stylistic cues.
⚠️ But in RoPE-based DiTs, this often breaks badly.
Instead of transferring style, the model starts copying the reference content.
In this work, Untwisting RoPE, we explore how positional bias and semantic understanding become twisted together in DiTs, and how frequency control can untwist them.
Excited to share our recent work: Free-Range Gaussians 🥚✨
The core idea: instead of predicting Gaussians on a pixel- or voxel-aligned grid, we let them live freely in 3D space.
🌐 Project: https://t.co/HkwmGam0Pq
📝 Paper: https://t.co/OhHA6VnwZT
Exciting progress in the auto-seam & auto-UV space!🔥
Thrilled to share our recent work MeshTailor.
Huge congrats to @qixuema for driving this incredible project. We finally brought AI to native seam generation, learning directly from professional data.
Check the video below!
Advances in 4D Representation: Geometry, Motion, and Interaction
Mingrui Zhao, @Dumb_Thug, Kai Wang, @anvorain, Guangda Ji, Peter Chun, @arash_mham, @richardzhangsfu
tl;dr: in title
https://t.co/76BsKEXdlh
Advances in 4D Representation: Geometry, Motion, and Interaction
Abstract (excerpt)
Instead of offering an exhaustive enumeration of many works, we take a more selective approach by focusing on representative works to highlight both the desirable properties and ensuing challenges of each representation under different computation, application, and data scenarios.
The main take-away message we aim to convey to the readers is how to select and then customize the appropriate 4D representations for their tasks. Organizationally, we separate the 4D representations based on three key pillars: geometry, motion, and interaction. Our discourse will not only encompass the most popular representations of today, such as neural radiance fields (NeRFs) and 3D Gaussian Splatting (3DGS), but also bring attention to relatively under-explored representations in the 4D context, such as structured models and long-range motions.
Throughout our survey, we will reprise the role of large language models (LLMs) and video foundational models (VFMs) in a variety of 4D applications, while steering our discussion towards their current limitations and how they can be addressed. We also provide dedicated coverage on what 4D datasets are currently available, as well as what is lacking, to drive the subfield forward.
📢📢📢 "ASIA" @ #SIGGRAPHAsia2025. 😉
We segment 3D shapes into possibly non-semantic and non-text describable parts using only a few annotated in-the-wild images as references!
👉Project Page: https://t.co/UZxVrwKWru
📄Paper: https://t.co/UOlEmlv258
🚀 Introducing PhysiX: One of the first large-scale foundation models for physics simulations!
PhysiX is a 4.5B parameter model that unifies a wide range of physical systems, from fluid dynamics to reaction-diffusion, outperforming specialized, state-of-the-art models.
Quick “teaser” for a fun #SIGGRAPH2025 project, led by Hossein Baktash, on optimizing a shape to have the desired rolling statistics.
Basically we can turn arbitrary objects into fair dice, or make dice which capture the statistics of other objects—like several coin flips.
Excited to share that TokenVerse won Best Paper Award at SIGGRAPH 2025! 🎉
TokenVerse enables personalization of complex visual concepts, from objects and materials to poses and lighting, each can be extracted from a single image and be recomposed into a coherent result. 👇