๐๐๐๐ง๐, ๐ผ๐๐ฟ ๐ฟ๐ฒ๐ฐ๐ฒ๐ป๐ ๐ฝ๐ฟ๐ผ๐ท๐ฒ๐ฐ๐ ๐ผ๐ป ๐ต๐ถ๐ด๐ต-๐ณ๐ถ๐ฑ๐ฒ๐น๐ถ๐๐ ๐ฏ๐ ๐๐ฎ๐๐๐๐ถ๐ฎ๐ป ๐ฎ๐๐ฎ๐๐ฎ๐ฟ ๐๐๐ป๐๐ต๐ฒ๐๐ถ๐ ๐ต๐ฎ๐ ๐ฏ๐ฒ๐ฒ๐ป ๐ฎ๐ฐ๐ฐ๐ฒ๐ฝ๐๐ฒ๐ฑ ๐๐ผ ๐๐ฉ๐ฃ๐ฅ ๐ฎ๐ฌ๐ฎ๐ฒ!
๐ก๐ง๐;๐๐ฅ
We study a mutually reinforcing synergy of 2D & 3D face priors (generative & data priors) to tackle the longstanding challenge in monocular synthesis of animatable, photorealistic 3D face avatar โ Balancing between in-the-wild generalization and efficient synthesis.
๐๐๐๐ง๐: ๐๐ณ๐ณ๐ถ๐ฐ๐ถ๐ฒ๐ป๐ ๐๐ฎ๐๐๐๐ถ๐ฎ๐ป ๐๐ฒ๐ฎ๐ฑ ๐๐๐ฎ๐๐ฎ๐ฟ ๐ณ๐ฟ๐ผ๐บ ๐ฎ ๐ ๐ผ๐ป๐ผ๐ฐ๐๐น๐ฎ๐ฟ ๐ฉ๐ถ๐ฑ๐ฒ๐ผ ๐๐ถ๐ฎ ๐๐ฒ๐ฎ๐ฟ๐ป๐ฒ๐ฑ ๐๐ป๐ถ๐๐ถ๐ฎ๐น๐ถ๐๐ฎ๐๐ถ๐ผ๐ป ๐ฎ๐ป๐ฑ ๐ง๐๐๐-๐๐ถ๐บ๐ฒ ๐๐ฒ๐ป๐ฒ๐ฟ๐ฎ๐๐ถ๐๐ฒ ๐๐ฑ๐ฎ๐ฝ๐๐ฎ๐๐ถ๐ผ๐ป
๐ ๐ฃ๐ฟ๐ผ๐ท๐ฒ๐ฐ๐ ๐ฝ๐ฎ๐ด๐ฒ: https://t.co/gsbkL5lYeC
๐ ๐ฃ๐ฎ๐ฝ๐ฒ๐ฟ: https://t.co/sJGGGHRQ7p
ELITE was a project that I've worked on with all my heart, and thanks to my collaborators' tremendous efforts and their patience. Huge congrats to my coauthors: Lee Hyoseok, Subin Park, @GerardPonsMoll1 , @Tae_Hyun_Oh
Now that we have building blocks for "๐ฅ๐๐ค๐ฉ๐ค๐ง๐๐๐ก๐๐จ๐ฉ๐๐" and "๐จ๐ค๐๐๐๐ก" world simulations. What'll be the next steps? ๐ง
3D/4D reconstruction is one of the great achievements of computer vision. With roots in photogrammetry, multi-view geometry was much studied (e.g. Hartley & Zisserman). In the last decade, it was successfully combined with deep learning. Shape priors which come from recognition got incorporated (e.g. for humans, in HMR) and recent systems like VGGT and SAM 3D can do amazing stuff. This progress is directly relevant for robotics, which (duh) operates in the 3D world. Many robotics papers in the learning era weren't exploiting 3D structure, which IMHO is just wasting valuable signal. But there has been a strong real2sim2real tradition for years e.g. here are two projects from my group (and there are many others, I am not claiming exclusivity). I am glad that people have rediscovered this capability with Astra. But please do give credit to human researchers, not just AI models. https://t.co/OeBSOfO6Sg ,
https://t.co/7ZIrP8SPU1
For 3D learning, the โBitter Lessonโ doesnโt properly respect Kolmogorov complexity.
The colloquial version of the bitter lesson is that, in a nutshell, just scaling up data and compute always beats problem-specific modeling (note that this is not actually what Sutton said; more on that later).
But look at the current situation with LLMs and agents generating 3D content: the really impressive examples depend on configuring an API, language, or software package (Three.js, OpenSCAD, Blender, etc.). Each representation impacts the description size for a given objectโand for the classes of object that can be faithfully represented via one of these procedural descriptions, the encoding is vastly more efficient than an explicit description (points, splats, polygons, etc.) or even an implicit description (SDF, NeRF, etc.).
Moreover, the smaller you can make the procedural description, the better the LLM works: it has to predict fewer tokens to describe the same object, because each token does a better job of compressing geometric information. E.g., the words โunit sphereโ are far more compact than a list of points or a collection of voxels.
So, in the procedural setting, building models for predicting shape becomes an exercise in bounding Kolmogorov complexity: whatโs the smallest โprogramโ I can use to describe a given object? (Or more precisely: which language minimizes description length *and* the entropy of the description, for a model trained on that language.)
Thatโs why current tools lean on carefully crafted abstractions (again, Blender, OpenSCAD, Three.js, etc.), for which many examples are available. Astra doesnโt run off and build you a giant neural network, hoping that the โright representationโ will somehow be learned deep within the network weights.
None of this, by the way, contradicts what Rich Sutton actually said: the input/output encoding matters, even if the rest is just scaling up search and learning. And the reason we see examples in Three.js and not a more esoteric language is that there is a large corpus of examples to learn from. But the idea that the network will just โlearn the best representationโ by adjusting its weights is a misguided fiction.
Instead, it really seems like we are entering an era of โcode as geometryโ counterposed to WorldLabsโ โgeometry as codeโ, where design of domain-specific languages is a key task. (And of course, search and learning will inevitably play a role in this designโฆ)
Computer Graphics also plays a central role, because itโs the mechanism for decompressing the procedural description into an explicit representation (building a cylinder, tessellating a curve, etc.). Thatโs the tool calling part of the agentic process. And the better these graphics algorithms get, the more effective the language becomes. People are worried about the death of Computer Visionโbut it seems Graphics is coming back from the grave. ๐งโโ๏ธ
(And, might I add, that differential geometry is a very expressive language that few are taking advantage of in this setting. ๐)
Excited to share our new arXiv paper, "๐๐ ๐๐ฟ๐ฎ๐ฐ๐๐ถ๐ป๐ด ๐ก๐ฒ๐๐ฟ๐ฎ๐น ๐ ๐ฎ๐๐ฒ๐ฟ๐ถ๐ฎ๐น๐ ๐ณ๐ฟ๐ผ๐บ ๐ ๐๐น๐๐ถ-๐๐ถ๐ฒ๐ ๐๐บ๐ฎ๐ด๐ฒ๐."
๐ Project page: https://t.co/LzYPpJhldE
๐ Paper: https://t.co/qK63dbwJWd
๐ก๐ง๐;๐๐ฅ
We tackle the challenging inverse rendering problem of extracting "neural materials" from images by combining:
1๏ธโฃ Large-scale material reconstruction foundation model
2๏ธโฃ Test-time optimization using differentiable Monte Carlo path tracer
3๏ธโฃ Uncertainty-guided material regularization
What is a "๐ก๐ฒ๐๐ฟ๐ฎ๐น ๐ ๐ฎ๐๐ฒ๐ฟ๐ถ๐ฎ๐น"? It's a universal SVBSDF representation that goes beyond classic PBR. The neural specular latent space encodes complex material-light interactions like haze, dust, clearcoat, fuzz, and scattering, all while being path-traceable in real time.
Classic PBR struggles (diffuse + single-lobe specular component) with these complex multi-layered specular effects. To overcome this and the lack of prior knowledge about the neural material latent space, we combine a Large Material Reconstruction Model (LMRM) for robust initialization with uncertainty-guided Test-Time Optimization (TTO).
This project wraps up an intense and amazing 13-week internship at NVIDIA Real-Time Graphics Research (https://t.co/J1cIoMKoiT). Huge thanks to my amazing mentors and the team: Jon Hasselgren, Jacob Munkberg, Peter Kocsis, Andrea Weidlich, and Tae-Hyun Oh. Super excited to share our efforts on next-generation generative content creation.
๐ขNeuMatEx - Material Reconstruction beyond PBR
We extract neural materials from multi-view images using a Large Material Reconstruction Model combined with differentiable rendering,
Great project from @kim_youwang!
https://t.co/awPdliD0EN
Excited to share our new arXiv paper, "๐๐ ๐๐ฟ๐ฎ๐ฐ๐๐ถ๐ป๐ด ๐ก๐ฒ๐๐ฟ๐ฎ๐น ๐ ๐ฎ๐๐ฒ๐ฟ๐ถ๐ฎ๐น๐ ๐ณ๐ฟ๐ผ๐บ ๐ ๐๐น๐๐ถ-๐๐ถ๐ฒ๐ ๐๐บ๐ฎ๐ด๐ฒ๐."
๐ Project page: https://t.co/LzYPpJhldE
๐ Paper: https://t.co/qK63dbwJWd
๐ก๐ง๐;๐๐ฅ
We tackle the challenging inverse rendering problem of extracting "neural materials" from images by combining:
1๏ธโฃ Large-scale material reconstruction foundation model
2๏ธโฃ Test-time optimization using differentiable Monte Carlo path tracer
3๏ธโฃ Uncertainty-guided material regularization
What is a "๐ก๐ฒ๐๐ฟ๐ฎ๐น ๐ ๐ฎ๐๐ฒ๐ฟ๐ถ๐ฎ๐น"? It's a universal SVBSDF representation that goes beyond classic PBR. The neural specular latent space encodes complex material-light interactions like haze, dust, clearcoat, fuzz, and scattering, all while being path-traceable in real time.
Classic PBR struggles (diffuse + single-lobe specular component) with these complex multi-layered specular effects. To overcome this and the lack of prior knowledge about the neural material latent space, we combine a Large Material Reconstruction Model (LMRM) for robust initialization with uncertainty-guided Test-Time Optimization (TTO).
This project wraps up an intense and amazing 13-week internship at NVIDIA Real-Time Graphics Research (https://t.co/J1cIoMKoiT). Huge thanks to my amazing mentors and the team: Jon Hasselgren, Jacob Munkberg, Peter Kocsis, Andrea Weidlich, and Tae-Hyun Oh. Super excited to share our efforts on next-generation generative content creation.
Check out our new arXiv paper, "๐๐ถ๐๐: ๐๐ฒ๐ฒ๐ฑ-๐ณ๐ผ๐ฟ๐๐ฎ๐ฟ๐ฑ ๐ถ๐ป๐๐๐ฎ๐ป๐ ๐๐ฎ๐๐๐๐ถ๐ฎ๏ฟฝ๏ฟฝ ๐๐ผ๐ฑ๐ฒ๐ฐ ๐๐๐ฎ๐๐ฎ๐ฟ๐ ๐ณ๐ฟ๐ผ๐บ ๐ฎ ๐ฆ๐ถ๐ป๐ด๐น๐ฒ ๐ฃ๐ผ๐ฟ๐๐ฟ๐ฎ๐ถ๐ ๐๐บ๐ฎ๐ด๐ฒ."
๐ Project page: https://t.co/4HT4FopLdV
๐ Paper: https://t.co/qadDfyEYB6
๐ก๐ง๐;๐๐ฅ
We can generate an animatable, photorealistic avatar from a single iPhone selfie in 5 seconds, thanks to our carefully designed 2D/3D foundation models:
1๏ธโฃ Human-centric vision foundation model (Fine-tuned Sapiens)
2๏ธโฃ Joint UV texture and geometry diffusion model
3๏ธโฃ Feed-forward mesh detail refinement network
4๏ธโฃ Universal mesh-to-3DGS prior model
Combining 2D (image/UV) space foundation models with a 3D space prior model can solve the challenging inverse problem of "Single selfie โ Photorealistic animatable 3D avatar."
This work was done while I was interning at Meta Codec Avatars Lab in 2024โ2025. I'm really proud of the results, and this work became the core motivation for my follow-up paper at CVPR 2026, ELITE (https://t.co/gsbkL5lYeC).
Special thanks to the amazing team for their invaluable guidance and patience:ย Chen Cao, Zhengyu Yang, Liuhao Ge, Yu Rong, Timur Bagautdinov, Su Zhaoen, Nir Sopher, Jovan Popovic, Deng Teng, @Tae_Hyun_Oh. It was my honor to work with you all, and again, I'm really proud of the results we achieved ๐
#CVPR2026 Highlight
How to make relighting more photorealistic? Make reconstruction happening together!
GeoRelight jointly resolves Geometry, Instrinsics, and Relighting, and proves they have mutual benefit (Geometry helps you know shadow and shading)
https://t.co/fpMUvXctiP
Large-scale Codec Avatars: learning photorealistic avatars from millions of videos.
A massive team effort, and incredibly proud of how it turned out.
- Project: https://t.co/XMZWMDuI0P
- Paper: https://t.co/iIMbGpGoc8 #CVPR2026
๐๐๐๐ง๐, ๐ผ๐๐ฟ ๐ฟ๐ฒ๐ฐ๐ฒ๐ป๐ ๐ฝ๐ฟ๐ผ๐ท๐ฒ๐ฐ๐ ๐ผ๐ป ๐ต๐ถ๐ด๐ต-๐ณ๐ถ๐ฑ๐ฒ๐น๐ถ๐๐ ๐ฏ๐ ๐๐ฎ๐๐๐๐ถ๐ฎ๐ป ๐ฎ๐๐ฎ๐๐ฎ๐ฟ ๐๐๐ป๐๐ต๐ฒ๐๐ถ๐ ๐ต๐ฎ๐ ๐ฏ๐ฒ๐ฒ๐ป ๐ฎ๐ฐ๐ฐ๐ฒ๐ฝ๐๐ฒ๐ฑ ๐๐ผ ๐๐ฉ๐ฃ๐ฅ ๐ฎ๐ฌ๐ฎ๐ฒ!
๐ก๐ง๐;๐๐ฅ
We study a mutually reinforcing synergy of 2D & 3D face priors (generative & data priors) to tackle the longstanding challenge in monocular synthesis of animatable, photorealistic 3D face avatar โ Balancing between in-the-wild generalization and efficient synthesis.
๐๐๐๐ง๐: ๐๐ณ๐ณ๐ถ๐ฐ๐ถ๐ฒ๐ป๐ ๐๐ฎ๐๐๐๐ถ๐ฎ๐ป ๐๐ฒ๐ฎ๐ฑ ๐๐๐ฎ๐๐ฎ๐ฟ ๐ณ๐ฟ๐ผ๐บ ๐ฎ ๐ ๐ผ๐ป๐ผ๐ฐ๐๐น๐ฎ๐ฟ ๐ฉ๐ถ๐ฑ๐ฒ๐ผ ๐๐ถ๐ฎ ๐๐ฒ๐ฎ๐ฟ๐ป๐ฒ๐ฑ ๐๐ป๐ถ๐๐ถ๐ฎ๐น๐ถ๐๐ฎ๐๐ถ๐ผ๐ป ๐ฎ๐ป๐ฑ ๐ง๐๐๐-๐๐ถ๐บ๐ฒ ๐๐ฒ๐ป๐ฒ๐ฟ๐ฎ๐๐ถ๐๐ฒ ๐๐ฑ๐ฎ๐ฝ๐๐ฎ๐๐ถ๐ผ๐ป
๐ ๐ฃ๐ฟ๐ผ๐ท๐ฒ๐ฐ๐ ๐ฝ๐ฎ๐ด๐ฒ: https://t.co/gsbkL5lYeC
๐ ๏ฟฝ๏ฟฝ๐ฎ๐ฝ๐ฒ๐ฟ: https://t.co/sJGGGHRQ7p
ELITE was a project that I've worked on with all my heart, and thanks to my collaborators' tremendous efforts and their patience. Huge congrats to my coauthors: Lee Hyoseok, Subin Park, @GerardPonsMoll1 , @Tae_Hyun_Oh
Now that we have building blocks for "๐ฅ๐๐ค๐ฉ๐ค๐ง๐๐๐ก๐๐จ๐ฉ๐๐" and "๐จ๐ค๐๐๐๐ก" world simulations. What'll be the next steps? ๐ง
ELITE:ย Efficient Gaussian Head Avatar from a Monocular Video viaย Learnedย Initialization and TEst-time Generative Adaptation
This work combines 3D data priors with 2D diffusion generative priors to create 3D Avatars from Monocular video. The 3D prior uses a feedforward network for speed and the 2D single step diffusion prior uses a clean reference frame to correct imperfect novel view renderings.
๐ฝ๏ธ Project Page: https://t.co/UPPMdJvRGe
๏ฟฝ๏ฟฝ๏ฟฝ Paper: https://t.co/QtcYa91SBJ
๐ป Code coming soon!
Here's a nice "proof without words":
The sum of the squares of several positive values can never be bigger than the square of their sum.
This picture helps make sense of how โโ and โโ norms regularize and sparsify solutions (resp.). [1/n]
Atย #ICLR2025, we will present "NeuFace: A Large-Scale 3D Face Mesh Video Dataset via Neural Re-parameterized Optimization."
๐ Hall 3 + Hall 2B #69
๐ Thu, Apr 24, 3:00โ5:30 pm Singapore Time
I'd really like to meet & discuss with fellow researchers. Let's connect!
(1/3)
TL;DR: We introduce a novel optimization approach for obtaining reliable 3DMM labels from large-scale internet facial videos.
The tracking code has already been released, and we will release a demo dataset shortly. Stay tuned!
Page: https://t.co/2hH1rqTy07
(2/3)