Vox-style animation using Gemini Omni Flash on Scenema. A six-minute AI explainer on FIFA's 2026 World Cup finances.
rt + comment and we'll send you the full prompt.
The narration is one consistent voice from start to finish. The visual language, paper collage on aged newsprint with hot red accents, holds through every single shot.
Try it at https://t.co/glyemSfmyY
Vox-style animation using Gemini Omni Flash on Scenema. A six-minute AI explainer on FIFA's 2026 World Cup finances.
rt + comment and we'll send you the full prompt.
The same recurring worker cutout appears in roughly 60 shots and stays the same character every time he shows up. The stadium, the Aramco logo, and the category 1 ticket also keep their identity across every appearance.
Scenema Audio is now on ComfyUI. You can install through ComfyUI Manager, load the pre-wired workflow, and generate performance quality speech.
https://t.co/ppOanmdhpu
That’s was really the most fundamental understanding we had to arrive at at Scenema. Text to video is DOA because unless you’re generating one-off, 15s clips, they’re not usable. But for anything that requires visual and vocal consistency, model separation is key.
Keep your character consistent across every output with https://t.co/nRKttcpdoX.
Choose from 10+ built-in styles, from photorealistic to anime, or create custom styles from reference images.
🎙️ Scenema Audio is UNLIKE any TTS you have seen before
🎭 and it actually PERFORMS, not just pronounces
🔹 Built on LTX 2.3's 22B audiovisual model trained on real film audio
🔹 Emotional arcs that shift mid-generation using action tags
🔹 Zero-shot voice cloning, any emotion even if never recorded that way
🔹 10 languages demoed: Arabic, French, Spanish, German, Hindi, Thai, Portuguese, Polish, Swedish, Chinese
🔹 Runs locally on 16GB+ VRAM, full precision on 48GB
🔹 Free, open source, Docker deploy in one command
🔥 Watch the full video below 👇
https://t.co/hTvZx60wmO
DreamX + MoCam + Scenema — The video AI trio
3 quiet but powerful drops:
DreamX — open-world video generation from a single image
MoCam — motion-aware camera AI for better video capture
Scenema — AI that generates full scene audio from visuals
Video AI is shipping faster than anyone can keep up.
@Iancu_ai Would love for you to check us out. Scenema Audio (Zero-Shot Expressive Voice Cloning and Speech Generation) is open source. Whereas you could use both to prompt and direct on our platform https://t.co/nRKttcpdoX
Let us know what you think!
@jennykrakovsky We would love to support you on this exciting journey. Would love for you to check out https://t.co/hrEpdQf6Or. Please DM us or join our discord for further assistance.
A new audio model is live on ModelScope! Meet Scenema Audio, a 13B expressive speech generation model. Zero-shot voice cloning, emotional acting, scene-aware audio. 🤖 https://t.co/xZmg8ssVfm
Built on an audio diffusion transformer extracted from LTX 2.3's audiovisual model. It learned how people actually sound in real scenes: angry, laughing, whispering, crying, exhausted, terrified.
🎭 Action tags direct emotional shifts mid-generation. Voice can go from calm to breaking down within a single output.
🧒 Native child voices: six-year-olds, toddlers, teenagers — not pitch-shifted adults.
🌧️ Scene-aware audio: describe the environment and get rain, thunder, crowds alongside the voice.
🎙️ Zero-shot voice cloning: 10-20s reference audio, no fine-tuning needed. Any voice, any emotion.
🌍 13 languages. Long-form narration with automatic splitting and voice continuity.
New open-source speech generator, Scenema Audio
> Incredible emotion control
> Zero shot voice cloning
> Action tags for stage direction, multilingual, long-form narration
> Free & local
Give it a listen!
https://t.co/wLjBvuGUKK