AI-powered human video generation
Text → Talking Video |
Image → Animated Persona |
Built by ByteDance AI Lab
DXkPp3Xx1aLg4g6GWbQfccsVXPLexY5rWHK5SPSpump
Official Launch! Our $OHUM token is now live on https://t.co/4gY2WtCMSG! Experience the synergy of cutting-edge AI and blockchain innovation with OmniHuman-1. Thank you for your incredible support as we take this exciting step forward
DXkPp3Xx1aLg4g6GWbQfccsVXPLexY5rWHK5SPSpump
❗️❗️SCAM ALERT: Please be advised that there is no token/coin launch by OmniHuman-1 at this time. Any claims or tweets suggesting otherwise are fraudulent. These coins do not belong to us. Always verify updates through our official channels. Stay safe!
More Halfbody Cases with Hands;
Here, we also provide additional examples specifically showcasing gesture movements. Some input images and audio come from TED, Pexels and AIGC.
In terms of input diversity, OmniHuman supports cartoons, artificial objects, animals, and challenging poses, ensuring motion characteristics match each style's unique features.
OmniHuman can support input of any aspect ratio in terms of speech. It significantly improves the handling of gestures, which is a challenge for existing methods, and produces highly realistic results.
OmniHuman supports various visual and audio styles. It can generate realistic human videos at any aspect ratio and body proportion (portrait, half-body, full-body all in one), with realism stemming from comprehensive aspects including motion, lighting, and texture details.
OmniHuman significantly outperforms existing methods, generating extremely realistic human videos based on weak signal inputs, especially audio. It supports image inputs of any aspect ratio, whether they are portraits, half-body, or full-body images.
We propose an end-to-end multimodality-conditioned human video generation framework named OmniHuman, which can generate human videos based on a single human image and motion signals (e.g., audio only, video only, or a combination of audio and video).
In OmniHuman, we introduce a multimodality motion conditioning mixed training strategy, allowing the model to benefit from data scaling up of mixed conditioning. This overcomes the issue that previous end-to-end approaches faced due to the scarcity of high-quality data.
One image. One prompt. Infinite possibilities. OmniHuman-1 transforms a single photo into a lifelike talking persona. No green screens, no CGI—just AI at work. Watch this! 👇
The future of human video generation is here.
Watch as [OmniHuman-1] turns text into a hyper-realistic talking human in seconds! No actors, no cameras—just pure AI magic.