Microsoft just dropped VASA-1.
This AI can make single image sing and talk from audio reference expressively. Similar to EMO from Alibaba
10 wild examples:
1. Mona Lisa rapping Paparazzi
🚨 𝘛𝘳𝘪𝘨𝘨𝘦𝘳 𝘞𝘢𝘳𝘯𝘪𝘯𝘨: 𝘓𝘢𝘯𝘨𝘶𝘢𝘨𝘦, 𝘷𝘪𝘰𝘭𝘦𝘯𝘤𝘦, & 𝘨𝘦𝘯𝘦𝘳𝘢𝘭 𝘢𝘸𝘧𝘶𝘭𝘯𝘦𝘴𝘴 𝘧𝘳𝘰𝘮 𝘵𝘩𝘦 𝘸𝘰𝘳𝘴𝘵 𝘰𝘧 𝘩𝘶𝘮𝘢𝘯𝘪𝘵𝘺.
As a social media manager, these past couple of weeks have been some of the most emotionally taxing of my entire career…