Where on the floorplan was this photo taken? ๐
Excited to share SceneAligner, accepted to #NeurIPS2026!
We tackle this by reconstructing the scene in 3D and aligning it to the floorplan. Joint work with @ruojin8@ElorHadar [1/8]
Where on the floorplan was this photo taken? ๐
Excited to share SceneAligner, accepted to #NeurIPS2026!
We tackle this by reconstructing the scene in 3D and aligning it to the floorplan. Joint work with @ruojin8@ElorHadar [1/8]
Excited to share our #NeurIPS2026 E&D ๐ Spotlight!
We all agree that memory is essential for modeling a persistent 4D world.
But: ๐ Can 4D Foundation Models Remember?
Introducing ๐ฅ PersistBench: a benchmark for testing visual memory in camera-controllable video models / 4D reconstruction models / video-to-360ยฐ models.
One key challenge in evaluating visual memory is obtaining real-world ground truth for what happens to an object after it leaves the cameraโs view.
In PersistBench, we introduce a simple way to sidestep this challenge. ๐คซ
๐ Project Page: https://t.co/CjedC4pCMT
๐ Leaderboard: https://t.co/r5lBmJkLEn
๐ GitHub: https://t.co/g85Y5USjYP
๐ค Dataset: https://t.co/sSgCprESzH
See how well your models do and submit to our leaderboard today!
๐งต [1/n]
I proudly share this new paper "๐ง๐ฒ๐ ๐๐๐ฟ๐ฒ ๐ฆ๐ฝ๐ฎ๐ฐ๐ฒ ๐ ๐ฎ๐๐ฒ๐ฟ๐ถ๐ฎ๐น ๐๐ถ๐ณ๐ณ๐๐๐ถ๐ผ๐ป" from my NVIDIA mentors: Jacob Munkberg, Peter Kocsis (@Peter4AI), and Jon Hasselgren!
The idea is elegant: instead of generating material views in image space and fusing them, run diffusion directly in PBR material texture space. That makes the result view-consistent by design, while reusing strong pre-trained video diffusion priors to generate complete PBR materials that generalize across arbitrary geometries and UV parameterizations.
Honestly, the 8K-res PBR texture generation is the most impressive I've seen for this task. Enjoy the super-detailed lion statue in the attached video!
I'm also excited about its extension to neural materials, which connects beautifully to my own ๐๐๐ช๐๐๐ฉ๐๐ญ work with the same team on extracting neural materials from images. Two sides of the same story โ pushing the boundaries of what we can do within PBR and with new representations beyond PBR.
Don't miss the stunning visuals on the project page!
๐ Page: https://t.co/WGu6RQ1Ua3
Also check out the updated ๐๐๐ช๐๐๐ฉ๐๐ญ results ๐ค
๐ Page: https://t.co/LzYPpJgNo6
Very excited to share LoGo, my internship project at @theworldlabs on post-training world models! We found that reward design matters a great deal in post-training long-horizon video gen for 3D consistency, echoing what we see in other domains, e.g. LLM reasoning. Check out more ๐ https://t.co/tLW6CVGaAk
G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing
Cornell University
TL;DR: We introduce G3T, a transformer that predicts upright, gravity-aligned pointmaps regardless of input image orientation, and G3T-Long, a pipeline that leverages this uprightness to enable robust long-sequence 3D reconstruction.
Ever suffered from making slides like I do? ๐ญ
Aligning everything, fighting animations, trying to add cool interactive visualizationsโฆwhile AI generated slides still look extremely confusing and overwhelming.
So I vibecoded JostSlides
โญ https://t.co/8m995aAIPL
๐งต (1/n)
@ShamoonMaxxing@ruojin8@ElorHadar Thanks! We tested on in-the-wild floorplans, including fairly complex layouts like this one ๐ We'd love to push this further to even more challenging floorplans!
โจ Really excited to see SceneAligner at #NeurIPS2026!
Love the cross-domain challenge: grounding images, 3D recons & floorplans in one spatial frame, even across scenes with little or no visual overlap.
Amazing work by Jun, and huge thanks to Hadar for all the guidance!