🚀 Thrilled to introduce RayRoPE, a plug-and-play module that improves multi-view attention!
https://t.co/RPLgluV8R9
Special thanks to my advisor @shubhtuls and collaborators: @Minsik_Je0n , Rick, and @OncelTuzel
[1/N] Rotary Position Embeddings (RoPE) are ubiquitous across transformers that process tokens from 1D, 2D, or 3D grids e.g. language, images, or videos. Our RayRoPE formulation extends these to multi-view transformers. Paper and code: https://t.co/abVobLRJxq
[1/N] Introducing Point4D: dense 4D reconstruction across hundreds of frames from a monocular video. Unlike existing feed-forward 4D methods that operate over short windows, Point4D keeps tracking in 3D, even across occlusions and out-of-frame motion.
I will be presenting RayRoPE on Thursday Morning @ #ECCV26
Spotlight presentation: Sep 10th, 10:17 am - 10:22 am @ MalmoMassan AB (The 3D spotlight session)
Poster: Sep 10th, 10:30 am -12:30 pm @ ExHall #295
I am always happy to chat!
RayRoPE is accepted to ECCV '26!
Need to condition your model on camera poses? Give RayRoPE a try!
In our latest version, we add more experiments and show that RayRoPE can be a plug-and-play module that improves performance in various 3D tasks!
[1/N] Large flow policies (e.g., VLAs) have become extremely popular, unlocking scalability. However, they come at a cost: the robot cannot react quickly. π𝐑² enables them to achieve reactive, real-time control.
Introducing MotionForesight:
A very simple approach re-purposing video models for forecasting future 3D scene flow given a short observed video context
Trained with just 40k RGB human videos, it generalizes well to completely new home/office scenes!
1/n
RayRoPE is accepted to ECCV '26!
Need to condition your model on camera poses? Give RayRoPE a try!
In our latest version, we add more experiments and show that RayRoPE can be a plug-and-play module that improves performance in various 3D tasks!
[1/N] Rotary Position Embeddings (RoPE) are ubiquitous across transformers that process tokens from 1D, 2D, or 3D grids e.g. language, images, or videos. Our RayRoPE formulation extends these to multi-view transformers. Paper and code: https://t.co/abVobLRJxq
To be honest, training on handmade 4D asset datasets is a dead-end. Almost all 4D asset data is synthetic and diverse real data barely exists, so models trained on it struggle to reconstruct objects that deform, get occluded, and move freely about the scene.
Our new work, Lift4D, instead lifts 2D & 3D priors into 4D, reconstructing complete dynamic objects from a single in-the-wild video 🧵 (1/n)
🔗Webpage + Demos: https://t.co/XI5jUViTpC
Feed-forward 3D reconstruction methods typically predict pointmaps in camera-centric frames. But why should a camera's arbitrary orientation define the coordinate system?
We introduce G3T, a transformer that predicts pointmaps in gravity-aligned frames. Regardless of input image orientation, our method always produces upright pointmaps (see demo).
We leverage this uprightness to create G3T-Long, a submap-based reconstruction method that improves robustness on long-sequence 3D reconstruction (more on that below).
Interactive demos, code, and model weights are available on our project page.
📢📢📢 Velox 🚀: Learning Representations of 4D Geometry and Appearance
In our #CVPR2026 paper, we introduce a method for learning a native 4D representation, useful for many downstream tasks, such as video-to-4D, 3D tracking, cloth simulation, and others!
🌐: https://t.co/MCkCMEftoJ
📝: https://t.co/iLKgrprXlO
Excited to share that we’ve open-sourced LiTo (ICLR 2026) code + models from @Apple!
Interactive image-to-3D generation from images.
• Apple Silicon demo via MLX
• Full training code
GitHub: https://t.co/wLy5EiHSJW
Paper: https://t.co/Se6k9AWboZ
#Apple#MLX#3D#AI
Two months ago, I vaguely posted a number: 0.9 FID, one-step, pixel space.
Now it is 0.75, and can be even lower.
Many wonder how.
I thought it might end as a small FID prank: simple and deliberate.
It started with one question: can FID be optimized directly, and what does it reveal?
Introducing FD-loss.
Most multi-view reconstruction models need full supervision. We show they can self-improve without any ground truth labels.
Introducing SelfEvo: Self-Improving 4D Perception via Self-Distillation. Up to +36.5% in video depth, +20.1% in camera estimation, zero annotation.
CRISP is accepted at ICLR 2026!!! @iclr_conf
Excited to see more impact of building simulation-ready assets from monocular video on animation / robotics
code is ready (https://t.co/yUAzZrrGJh) with the cleaned-up videos, including several parkour videos clipped from YouTube.
Most 3D representations capture shape or texture, but rarely both, especially view-dependent effects like reflections.
Check out LiTo: a set of latent tokens that capture both geometry and appearance for high-quality image-to-3D generation.
https://t.co/JaHas1gdNw
(1/n)
𝗢𝗻𝗲 𝗺𝗲𝗺𝗼𝗿𝘆 𝗰𝗮𝗻’𝘁 𝗿𝘂𝗹𝗲 𝘁𝗵𝗲𝗺 𝗮𝗹𝗹.
We present 𝗟𝗼𝗚𝗲𝗥, a new 𝗵𝘆𝗯𝗿𝗶𝗱 𝗺𝗲𝗺𝗼𝗿𝘆 architecture for long-context geometric reconstruction.
LoGeR enables stable reconstruction over up to 𝟭𝟬𝗸 𝗳𝗿𝗮𝗺𝗲𝘀 / 𝗸𝗶𝗹𝗼𝗺𝗲𝘁𝗲𝗿 𝘀𝗰𝗮𝗹𝗲, with 𝗹𝗶𝗻𝗲𝗮𝗿-𝘁𝗶𝗺𝗲 𝘀𝗰𝗮𝗹𝗶𝗻𝗴 in sequence length, 𝗳𝘂𝗹𝗹𝘆 𝗳𝗲𝗲𝗱𝗳𝗼𝗿𝘄𝗮𝗿𝗱 inference, and 𝗻𝗼 𝗽𝗼𝘀𝘁-𝗼𝗽𝘁𝗶𝗺𝗶𝘇𝗮𝘁𝗶𝗼𝗻.
Yet it matches or surpasses strong optimization-based pipelines. (1/5)
@GoogleDeepMind@Berkeley_AI
Spatial reconstruction is a long-context problem: real scenes come with hundreds of images. But O(N²) transformer-based models don’t scale efficiently.
Introducing: 🤐ZipMap (CVPR ’26): Linear-Time, Stateful 3D Reconstruction via Test-Time Training (TTT).
ZipMap “zips” a large image collection into an implicit TTT scene state in a single linear-time operation. The state will then be decoded into spatial outputs, and can be queried efficiently for novel-view geometry and appearance (~100 FPS)
ZipMap is not only much faster (>20× faster than VGGT), but also matches or surpasses the accuracy of all SOTA models.
One of the more interesting and thought provoking research papers I've seen in a while. A system for reading and reimplementing NeRF papers, and it seems to work very well. Pretty easy to extrapolate out from here to what CVPR 2027 papers will look like. https://t.co/gokzG27mIT
📢Current world models aren't really modeling the world; they're modeling one agent's view of it. Partial observations ≠ world state.
Future world models will be independent of any one agent's perspective. You will be able to “drop in” any number of agents at any point in time, and a persistent world state will evolve with their interactions. Imagine a neural MMORPG server. 🧵[1/10]
Vincent's post is timely. I wrote a follow up examining some of his arguments and providing an alternative perspective.
https://t.co/D3BBGiV938
Yes, the lesson is bitter, but I believe the flavor is also sweet. Thread 🧵