Excited to share the last chapter of my PhD! Presenting SIGMA-Gen: Structure and Identity Guided Multi-subject Image Generation accepted at ICLR 2026.
Paper:https://t.co/G8nYBJOyYZ
Page:https://t.co/pmfJaLnpcL
Demo:https://t.co/M6vOv7cUJS
#ICLR2026#ComputerVision#GenerativeAI
Excited to share the last chapter of my PhD! Presenting SIGMA-Gen: Structure and Identity Guided Multi-subject Image Generation accepted at ICLR 2026.
Paper:https://t.co/G8nYBJOyYZ
Page:https://t.co/pmfJaLnpcL
Demo:https://t.co/M6vOv7cUJS
#ICLR2026#ComputerVision#GenerativeAI
We also introduce an automatic synthetic data generation pipeline which we use to generate SIGMA-Set27K providing identity, structure, and spatial information for over 100k unique subjects across 27k images.
🎯 GTA-CLIP
A transductive CLIP approach combining attribute expansion, image-language clustering, and fine-tuning. Improves over CLIP and TransCLIP across architectures.
Led by @oindrilla
🔗 https://t.co/pasfbFgs9S
📍 Oct 21 | 11:30–1:30 | Poster 1 (#120)
GTA-CLIP proposes a novel inference time strategy to improve zero shot classification with VLMs through joint transduction and adaptation in image and language spaces.
We are presenting YouDream today at #NeurIPS2024 at East Exhibit Hall #2706 in the morning session. Drop by and chat with both @Sandy_uta and me at the poster!
Project page: https://t.co/D8xmgnJEq1
Our new paper "Human-in-the-Loop Visual Re-ID for Population Size Estimation" co-authored with @MajiSubhransu, @Grant_Van_Horn, and Dan Sheldon, will be presented at #ECCV2024.
Visit our poster (#51) on Thursday!
We'll also be presenting at the @CV4E_ECCV workshop on Monday.
That's cool. Create posed and anatomic consistent animals with text-to-3D.
"YouDream: Generating Anatomically Controllable Consistent Text-to-3D Animals"
Project: https://t.co/ErX3smIIOZ
Paper: https://t.co/rhLW01lvin
I hope we will get some code soon to try it out!
8 papers submitted to daily papers so far today
If your a author with at least one indexed paper on HF, submit your paper directly to daily papers: https://t.co/VbXh61IzHE
Work done with @Sandy_uta (https://t.co/WZpEM7yuD8) and Prof. Alan Bovik(https://t.co/xiR7f791uo)
arXiv: https://t.co/LygtZoiynb
page: https://t.co/D8xmgnJEq1
A 2D TetraPose ControlNet guides the 3D generation, thus implicitly encoding both pose and camera angle in the control image. The ControlNet is trained to generate any tetrapod animal such as birds, reptiles, amphibians, and mammals given an input 2D pose image.
Excitingly, YouDream can generate imaginary animals never before seen based on an artist’s designed 3D pose. Previous methods do not faithfully follow the text in case of these low-represented scenarios, while YouDream is able to generate compelling creative 3D assets.
YouDream can solve 3D inconsistency issues of SDS based methods guided by T2I diffusion models, by utilizing a 3D pose prior. We solve anatomic and geometric inconsistencies such as the Janus head problem or multiple limbs without training on any 3D datasets.
🚨 New paper
YouDream: Generating Anatomically Controllable Consistent Text-to-3D Animals
YouDream achieves multi-view consistency without being trained on 3D datasets and generates imaginary assets impossible to be made using prior methods.🧵 below.
page: https://t.co/6GJEVTfuE2
It was a blast to be at #CVPR2024 meeting old friends and new. Thoroughly enjoyed presenting AdaptCLIPZS (https://t.co/imdgtl5nwt) with the help of @Grant_Van_Horn and @MajiSubhransu