‼️500h of 3D motion data released‼️
Our team at the Codec Avatars Lab just released a large scale dataset of 3D tracked human motion, including audio and text annotations. Check it out here: https://t.co/ZnRxNjtJx1
Looking for a research internship in 2025?
The Social AI Research group at Meta’s Codec Avatars Lab in Pittsburgh is offering topics such as neural rendering, body tracking, motion synthesis, and animation from multimodal sensor inputs.
Link:
https://t.co/iuNi6Lfr7Y
Working on your next big speech paper but still looking for a suitable dataset?
Check out EARS: 100h of full-band expressive, anechoic recordings of speech from 107 speakers with 22 different emotions, 7 different reading styles, and more.
https://t.co/MSdUp72Vb1
If you're at #CVPR2024 this week, don't miss talking to @AlexRichardCS we are hiring for postdoctoral and full time positions! https://t.co/duh7vsy1Mz https://t.co/1jlojI6oHu
📢Curious about the future of 3D scene understanding? Join us at the 1st Workshop on Multimodalities for 3D Scenes @CVPR! Learn about the latest research on using vision, audio, touch, and language to understand 3D scenes around us.
📅June 17, 1:30-5:20PM
https://t.co/mwzTFWF6HB
Real Acoustic Fields are here!
Check out or dataset of densely captured room impulse responses paired with multi-view images! See you all on CVPR :)
project page: https://t.co/4amfXqqTZF
arxiv: https://t.co/CGwmez9HwZ
@danveloper The restriction comes from our data collection: we only train on dyadic conversations, so the model has never seen 3+ people in the same environment.
Motion generation for photorealistic avatars? Say no more!
Check out how we animate full body avatars exclusively from audio input!
Paper: https://t.co/6w7M96FoGW
Project page: https://t.co/Xny77x54iE
Dataset + Code: https://t.co/yM0VPfv5wN
3D body models now have sound!
We demonstrate 3D spatial audio synthesis for 3D full body models. (NeurIPS 2023 Spotlight)
Paper: https://t.co/hWJZEyIs5G
Dataset + Code (to be released): https://t.co/0Uf1by98We
Work done in collaboration with @AlexRichardCS, @ovrdr, Vamsi Ithapu, @NataliaNeverova, Kristen Grauman, and Andrea Vedaldi @MetaAI@RealityLabs
The poster session is Tues-PM and I will also give a talk at the sight and sound workshop Monday 4:30pm.
A new generalized and universal audio-visual speech enhancement model, powered by SSL.
A single model for denoising, source sep, inpainting and lip-reading!
Checkout the paper and demo video!!
🗣️🤖🔊
More details below
with @mhnt1580
📢Excited to share our recent #ECCV2022 paper: "LiP-Flow: Learning Inference-Time Priors for Codec Avatars via Normalizing Flows in Latent Space" with Shugao Ma, @akcalakcal, Stanislav Pidhorskyi, @AlexRichardCS, Shih-En Wei, Jason Saragih and @OHilliges.
**Dataset Release!**
We released high-quality 3D face data of 13 persons captured with up to 150 cameras while performing 100+ facial expressions!
https://t.co/hwAP7zHzHq