💡Let me introduce Seedance 1.0 Video Foundation model🎬, a model that supports multi-shot video generation from both text and images.
🥇Seedance 1.0 ranked #1 in both text-to-video & image-to-video on Artificial Analysis’s third-party benchmarks. (Note: Veo 3 Preview’s audio was not available for fairness.
🍹It achieves breakthroughs in semantic understanding and prompt following, and can create 1080p videos with smooth motion, rich detail, and cinematic aesthetics.
Key Features:
🍭Smooth & Stable Motion: Seedance 1.0 has a wide dynamic range, generating fluid, large-scale movements. From subtle expressions to active scenes, it maintains a high level of stability and physical realism.
🍦Native Multi-Shot Storytelling:Natively supports the generation of narrative videos with multiple cohesive shots. It maintains consistency in the main subject, visual style, and atmosphere across shot transitions and spatio-temporal shifts.
🔥Diverse Stylistic Expression: From photorealism and cyberpunk to traditional Chinese animation and claymation, Seedance 1.0 can accurately interpret diverse stylistic prompts to support a wide range of creative needs.
👑Precise Semantic & Prompt Following: Accurately parses natural language prompts, enabling stable control over multi-agent interactions, complex action sequences, and a rich variety of camera movements to precisely translate your textual concepts into videos.
Technical Report: https://t.co/EG1MdT9Rs4
Official Website: https://t.co/rA80AuTRKe
#seed #seedance #bytedance #video-generation #LLM #ai
🌟 Let me introduce Seedream 3.0, the ultimate Image Generation Foundation Model brought to you by ByteDance’s Seed Team.
🚀 This game-changer landed on multiple platforms, like Doubao and Jimeng in early April 2025.
🙌Seedream3.0 has the capabilities of 2K resolution direct output, small text layout, and high generation efficiency, which greatly reduces the threshold for visual creativity in posters and covers.
Key Features:
🌟 Comprehensive capability upgrades—everything is better, faster, and more powerful.
✍️ Enhanced text rendering—perfect for complex Chinese characters and typography.
🎨 Aesthetic improvements—your visuals will look amazing next-level. 😍
📸 Native high-resolution output—up to 2K for stunning details.
💡 Efficient inference cost—speed and performance optimized.
Seedream 3.0 (formerly known as Mogao) entered the Artificial Analysis rankings and landed in the first tier, right alongside GPT-4o. 💪 It crushed other models like Recraft V3, HiDream, Reve Image, Imagen 3 (v002), FLUX1.1 Pro, and Midjourney v6.1. 🔥
📄 Technical Report: https://t.co/gkzGZZJwqY 🌐 Official Website: https://t.co/x9cxdR0pfk
#seedream_3_0 , #genAI , #text_to_image_model , #AI #image_generation
Our reward function during training is based not only on the final generation quality but also on the total number of denoising steps. By adjusting the attenuation factor γ, we can obtain models with varying statistical denoising steps.
So glad that TPDM has been accepted by #CVPR2025! TPDM is a series of diffusion models featuring an adaptive denoising scheduler, trained through reinforcement learning. Big thanks to @guojunq and @William74312006 for their incredible support!
paper link: https://t.co/oHcmK2x3qr
Compared to the original Stable Diffusion 3 model, we are able to achieve more accurate image generation and utilize fewer steps.
SD3-Medium (15steps, 28steps) TPDM-SD3-Medium (15steps)
The adaptive scheduler is implemented through a plug-and-play Time Prediction Module, which predicts the next noise level based on the current image denoising state.
We proposed a new RL training paradigm for large-scale text-to-image DiT. Based on PPO, an adaptive denoising scheduler is trained to achieve faster and better image generation without changing the original DiT parameters.