Personal Update: Excited to share that I've joined DoorDash's newly minted AI Research team as a Member of Technical Staff, working alongside @aryanxshah and @andyfang. We're defining the next era of post-training for agentic commerce, memory, multimodal understanding, and robotics for last-mile delivery.
We're hiring! DM me if you want to help shape our research agenda, technical direction, and mission.
The last four years at Meta were an incredible adventure, and I'm deeply grateful to the colleagues who I had the opportunity to work with.
This also comes with a move to the San Francisco Bay Area. Looking forward to reconnecting with old friends and colleagues out here.
C-DiffGAN encourages knowledge retention by 1) generating reminiscences of previous low-resource speaker data, then 2) crossmodally aligning to them to mitigate catastrophic forgetting
5/n
Adapting Generative models for co-speech gestures to new speakers, but without forgetting the style of previous speakers? Let’s make it more challenging by only having 2-10 minutes of data for the new speakers.
Website: https://t.co/iJI9W5GXId
Find us at #ICCV2023
1/n
We propose an approach, C-DiffGAN, to learn a multi-speaker co-speech gesture generation model in a continual learning setting. i.e. By watching only 2 minutes of videos for every speaker in a sequence (without having access to previous speakers’ data).
4/n
Challenging and enlightening as it was, it wouldn't have been possible without my wonderful co-authors Simbarashe Nyatsanga, @SvitozarTaras, Gustav Henter, and Michael Neff
Last year we surveyed the body of co-speech gesture generation work with a heavier focus on the past decade of data-driven learning.
Link: https://t.co/ThcsCHSW5a
A Comprehensive Review of Data-Driven Co-Speech Gesture Generation
Simbarashe Nyatsanga,@SvitozarTaras , @chahuja, Gustav Eje Henter, Michael Neff
tl;dr: if you are starting to work in gesture recognition/generation, that is a good read to start.
https://t.co/PYFrueB6UU
We're hiring interns in the Autonomous Systems and Robotics Research group here at Microsoft.
If you are a PhD student interested in working at the intersection of fluids simulations and ML, please consider applying:
https://t.co/HeFXV8Uf9p
Learning to generate personalized co-speech gestures from spoken language, but do not have 5-10 hours of supervised training data for that person?
Check us out today at #CVPR2022 poster session 4.2 from 230-500pm at poster 180b
w/ @_dongwonlee@lpmorency
1/n
More co-speech gesture generation work
1 https://t.co/kIwx6e9d9P
2 https://t.co/Ahwu88IGe6
3 https://t.co/IC7exQhPoV
4 https://t.co/IBPnx3RSQy
5 https://t.co/iSAymZF3UX
6 https://t.co/DRoZe3kZhl
7 and many more
DM me if you are starting out in this area or just want to chat
4/n
Check out this workshop at #CVPR2022 (and upcoming survey paper). It beautifully builds a taxonomy encompassing key challenges of multimodal ML research
If recent models like DALL.E, Imagen, CLIP, and Flamingo have you excited, check out our upcoming #CVPR2022 tutorial on Multimodal Machine Learning - next monday 6/20 9am-1230pm
https://t.co/nnBbipg34G
slides, videos & a new survey paper will be posted soon after the tutorial!
Really excited to release the video of my guest lecture on Multimodal Deep Learning for CMU's Deep Learning class @rsalakhu@mldcmu
It covers 5 fundamental concepts in multimodal ML: representation, alignment, reasoning, translation & co-learning
Youtube: https://t.co/I9Ukv0r2NE
Tomorrow, we launch "The Art of the Paper", a course on the principles, mechanics, & culture of scientific writing. While core to our work, this material seldom gets a formal treatment. I'm excited, nervous, & grateful to @mldcmu for the creative freedom.
https://t.co/iHtTVV469X