Heading to #ICLR2026 in Rio 🇧🇷 this week to present DEAS!
📍 Poster session: Fri, April 24, 10:30 AM – 1:00 PM
📌 Location: Pavilion 3 P3-#1308
Happy to chat about robot learning, VLAs, and RL!
I’ll be on the job market soon and looking for exciting opportunities!
#ICLR
Excited to share Adaptive Low-Pass Guidance (ALG): a simple training-free, drop-in fix that brings dynamic motion back to Image-to-Video models! Demo videos, paper, & code below!
https://t.co/4NzYDfCFSb (🧵 1/7)
Excited to present FastTD3: a simple, fast, and capable off-policy RL algorithm for humanoid control -- with an open-source code to run your own humanoid RL experiments in no time!
Thread below 🧵
Cool work from @JitengMu on image editing! This looks like the future of Photoshop!
Just select a few patches and move it, and boom, you get the edited picture! This enables the new control ability of the diffusion model: Only change the part you want to change.
https://t.co/x2xeUy6GaW
Suppose that we train two INRs: One for a natural image, and another for its pixel-shuffled version.
Which INR would fit faster?
Expected: Natural Image
Reality: Pixel-permuted image 🤯
(under some conditions)
We look closer into when & why this happens in our #CVPR2024 oral.
Call for Papers: #INRV2024 Workshop on Implicit Neural Representation for Vision @ #CVPR2024! Topics: Compression, Representation using INR’s for images, audio, video & more! Ddl: 3/31. Submit now! @CVPR
Website: https://t.co/380CMq5Ouc
Submission Link: https://t.co/KwkGqyKjhy
🙌 At #NeurIPS to present two papers
Wed. 10AM: Efficient meta-learning / 5PM: Improving MAE with meta-learning.
I am working on continual learning/efficiency/AutoML for LLMs! DM me if you are interested in chatting
👀 Also looking for an internship, welcome any recommendations
Best Paper Award Honorable Mention at #3DV2019
Learned Multi-View Texture Super-resolution
Audrey Richard; Ian Cherabier; Martin R. Oswald; Vagia Tsiminaki; Marc Pollefeys @mapo1; Konrad Schindler
#TBThursday#3DV2024
I will present Multi-View Masked World Models (MV-MWM) for visual robotic manipulation.
Please visit #ICML2023 poster session at 7/25 (Tue) 2:00-3:30pm!
Imagine a 2D image serving as a window to a 3D world that you could reach into, manipulate objects, and see changes reflected in the image.
In our new OBJect 3DIT work, we edit images in this 3D-aware fashion while only operating in the pixel space!
🧵
Collaborative Score Distillation for Consistent Visual Synthesis
paper page: https://t.co/uMZnQUmgO2
Generative priors of large-scale text-to-image diffusion models enable a wide range of new generation and editing applications on diverse visual modalities. However, when adapting these priors to complex visual modalities, often represented as multiple images (e.g., video), achieving consistency across a set of images is challenging. In this paper, we address this challenge with a novel method, Collaborative Score Distillation (CSD). CSD is based on the Stein Variational Gradient Descent (SVGD). Specifically, we propose to consider multiple samples as "particles" in the SVGD update and combine their score functions to distill generative priors over a set of images synchronously. Thus, CSD facilitates seamless integration of information across 2D images, leading to a consistent visual synthesis across multiple samples. We show the effectiveness of CSD in a variety of tasks, encompassing the visual editing of panorama images, videos, and 3D scenes. Our results underline the competency of CSD as a versatile method for enhancing inter-sample consistency, thereby broadening the applicability of text-to-image diffusion models.
🤔 Learning large-scale neural fields (NFs) requires a huge amount of memory and time to train…
💡I’m excited to share “Learning Large-scale Neural Fields via Context Pruned Meta-Learning”.
Paper: https://t.co/usIAAwtVFh
"JPEG Compressed Images Can Bypass Protections Against AI Editing"
Since neural nets are vulnerable to adversarial perturbations, you might hope that you could modify your image to resist AI image editing. But... [1/2]
ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation code release is out
github: https://t.co/mhijcWUVc6
@Gradio demo: https://t.co/fHKEcMwNFU