Why create robot intelligence for just one hand, when we could have it learn from many?
GEN-1, our latest embodied foundation model, now supports a broad range of end effectors from 5-finger hands, to specialized tools, and everything in between.
From imitation to Spatial Reasoning
Learn2Fold is based on a simple idea: treat cloth folding as a robotics task, moving beyond imitation learning toward better generalization across deformable objects. We hope this is a real step toward reasoning.
Most of my creative energy these days goes into my work at Google, but I still often find myself opening Dreams just to sculpt and paint. I haven’t seen anything else that delivers that same sense of effortless 3D sculpting — it’s the perfect kind of input for generative AI workflows, where you want tools that let you express as much creative intent as possible in a way that feels natural, inspiring, and fun.
Working in Dreams feels less like operating software and more like painting or playing an instrument. And honestly, that’s exactly the kind of experience we should be aiming for with 3D creation in the age of generative AI
Pixels are the result. Strokes are the story.
We turn one image into vectorized Bézier strokes and differentiable smudges that could have painted it. Watch the canvas come alive: https://t.co/1JKW5ZBEzi
Big thanks to @_akhaliq for the shoutout!🥳
Our #SIGGRAPH2025 paper RenderFormer does end-to-end rendering from triangle meshes to images, with global illumination💡 —— using just a simple transformer!!!🪄
🔗Project page: https://t.co/vk6cA4FgOE
💻Code: https://t.co/08Goe9mI13
🎉Excited to share that ARM: Appearance Reconstruction Model for Relightable 3D Generation, has been accepted at CVPR (✨Highlight)! ARM enables 3D generation with PBR materials for relighting under novel views/lighting. Check https://t.co/yeerJwxdJx for more results and details!
Introducing NVIDIA Cosmos, an open-source, open-weight Video World Model. It's trained on 20M hours of videos and weighs from 4B to 14B. Cosmos offers two flavors: diffusion (continuous tokens) and autoregressive (discrete tokens); and two generation modes: text->video and text+video->video.
Physical AI has a big data problem. Synthetic data to the rescue! We apply Cosmos to large-scale synthetic data generation for robotics and autonomous driving, and now you can too! It's all yours to finetune.
Check it out: https://t.co/ZGCeWRd2vj
Introducing 🧞Genie 2 🧞 - our most capable large-scale foundation world model, which can generate a diverse array of consistent worlds, playable for up to a minute. We believe Genie 2 could unlock the next wave of capabilities for embodied agents 🧠.
GS^3: Efficient Relighting with Triple Gaussian Splatting
Abstract:
We present a spatial and angular Gaussian based representation and a triple splatting process, for real-time, high-quality novel lighting-and-view synthesis from multi-view point-lit input images.
To describe complex ap pearance, we employ a Lambertian plus a mixture of angular Gaussians as an effective reflectance function for each spatial Gaussian.
To generate self-shadow, we splat all spatial Gaussians towards the light source to obtain shadow values, which are further refined by a small multi-layer perceptron.
To compensate for other effects like global illumination, another network is trained to compute and add a per-spatial-Gaussian RGB tuple.
The effectiveness of our representation is demonstrated on 30 samples with a wide variation in geometry (from solid to fluffy) and appearance (from translucent to anisotropic), as well as using different forms of input data, including rendered images of synthetic/reconstructed objects, photographs captured with a handheld camera and a flash, or from a professional lightstage.
We achieve a training time of 40-70 minutes and a rendering speed of 90 fps on a single commodity GPU. Our results compare favorably with state-of-the-art techniques in terms of quality/performance.
We have updated the paper Gaussian Splashing with more results and experiments. Please visit the project page for more information!
Link: https://t.co/4C6ZaTs2ky
TL;DR: Gaussian Splashing is a unified framework combining 3D Gaussian Splatting and Position-Based Dynamics.
@Hachemhugo @FeiHuangFH I believe temporarily high citations do not necessarily equal high value. An important work might be based on an article that went unnoticed twenty years ago.
@Hachemhugo @FeiHuangFH Just some random thoughts. How do we define impact? Is it measured by follow-up citations? If so, does publishing in trending areas yield more impact, and should researchers focus on these directions more?
Our 2013 SIGGRAPH paper "Femto-Photography: Capturing and Visualizing the Propagation of Light" has received the 2024 Test of Time Award! It's given to papers "that have had a significant and lasting impact on computer graphics and interactive techniques over at least a decade".
Thanks @_akhaliq ! DiLightNet has been accepted to #SIGGRAPH and we can now acknowledge the authorship🎉🎉🎉 Check https://t.co/9SmFRhVXar for more results!
Thanks AK for sharing our work!
Neural Gaffer is an end-to-end 2D relighting diffusion model that accurately relights any object in a single image under various lighting conditions.
Moreover, by combining with other generative methods, our model enables many downstream 2D tasks, such as text-based relighting and object insertion.
Our model can also operate as a strong relighting prior for 3D tasks, such as relighting a radiance field directly in minutes without an inverse rendering construction process.
Please refer to our webpage for more results and details: https://t.co/zAxnXP2Rpw