Excited to share Rho-1, a 19B omni model we trained from scratch! 🚀
Rho-1 can draw a scene, find objects in it, animate it into video, edit that video by prompt, reason across a multi-turn chat, and more! All in one network, with no pipelines and no tool calls.
Today, we are releasing a research preview of Rho-1, our 19B omni model that understands and generates text, images, video and robot actions, all inside a single neural network.
really proud to share what the team has been working on.
rho-1 is our 19B omni model. reasoning, real-time video simulation, and continuous robotic control inside a model.
we are very early on the compute curve and candid about current limits in the post, but the compounding gains across modalities are already clear.
Incredibly proud of what the team built! Here are some unedited screen recordings of Rho-1, exactly as it ran.
No agent calling specialist models behind the curtain. Just one model, end to end.
One reason this model is so interesting is that it brings together two things usually built separately: understanding and generation.
We pretrained them as one, from day one, on next-token prediction and flow matching.
Most AI systems are agentic pipelines that hand jobs off to specialist models, losing context at every step.
Rho-1 is one model. In a single conversation, it creates, edits, and reasons about text, images, and video, all reading from one shared world state.
World models are increasingly central to how agents learn and plan.
Today we're releasing WorldModelGym, a benchmark built around a single question: if an agent uses a world model to choose among actions, does it pick the right one?
We call this decision-based fidelity. 100+ tracks across Atari, Meta-World, DeepMind Control, and classic control. One frozen policy. Reality scores it.
Read the full post → https://t.co/OzVd1n6Vth
Most evaluations of world models ask, "Does this look right?" 🌎
We built WorldModelGym to ask a different question: if an agent actually uses this model to make decisions, does it still choose well? 🏋️🏃♂️🏁🦾
A simulation can look physically plausible and still lead an agent completely astray. 📉⚠️
Stay tuned. More tomorrow at https://t.co/gaf6iEgEiw. 👀
We are excited to join forces with @moonvalley and merge the teams!
Together, we are building multimodal models that reason about, simulate, and act in the physical world.
Read the full announcement: https://t.co/tPPFrgoO8H
Stay tuned for a few exciting updates about our work in the coming weeks.
🎬 Introducing Stable Cinemetrics, to be presented at NeurIPS 2025.
We present the first taxonomy of professional controls to systematically study and control video generative models through the lens of filmmaking.
Interactive webpage with paper link: https://t.co/Eh4Hw3hBZl
🧵
Stability AI released Stable Diffusion 3.5 yesterday. Below are comparisons of how Stabile Diffusion has improved in the past year since SDXL in July 2023
We have also added @StabilityAI's Stable Diffusion 3.5 & the Turbo variant to our Image Arena. Our Image Arena crowdsources preferences to understand & compare the quality of image models - currently we have >800k preferences submitted.
Link to Image Arena below 👇
Guess who's back? Back again! 🎵 @StabilityAI is back, tell a friend 🎤
Stable Diffusion 3.5 Large is here 🔥
- 🏋️ 8B parameters
- Full 💪 and 🏎️💨 4-step Turbo variant
- 🧾 🤝 commercial use (for orgs below 1M year/rev)
- 🧨 day-0 LoRA fine-tuning support
Our analysis shows that Stable Diffusion 3.5 Large leads the market in prompt adherence and rivals much larger models in image quality.
Stable Diffusion 3.5 Large Turbo offers some of the fastest inference times for its size, while remaining highly competitive in both image quality and prompt adherence, even when compared to non-distilled models of similar size. (3/4)
Introducing Stable Diffusion 3.5, our most powerful models yet.
This open release includes multiple variants that are highly customizable for their size, run on consumer hardware, and are free for both commercial and non-commercial use under the permissive Stability AI Community License.
You can download Stable Diffusion 3.5 Large and Stable Diffusion 3.5 Large Turbo from Hugging Face and the inference code on GitHub now. Stable Diffusion 3.5 Medium will be released on October 29th.
Read more here: https://t.co/1DH1xEqfJs (1/4)
Today, we are pleased to announce the availability of Stable Diffusion 3 and Stable Diffusion 3 Turbo on the Stability AI Developer Platform API.
We have partnered with @FireworksAI_HQ , the fastest and most reliable API platform in the market, to deliver these models.
In keeping with our commitment to open generative AI, we aim to make the model weights available for self-hosting with a Stability AI Membership in the near future.
You can get started and learn more here: https://t.co/r7xoprFQJO
Prompt: Awesome artwork of a wizard on the top of a mountain, he's creating the big text "Stable Diffusion 3 API" with magic, magic text, at dawn, sunrise.
Today, we are releasing Stable Video 3D, a generative model based on Stable Video Diffusion. This new model advances the field of 3D technology, delivering greatly improved quality and multi-view.
The model is available now for commercial and non-commercial use with a Stability AI Membership.
Learn more and read the research paper here: https://t.co/uXwZalGBky (1/3)
The @intel Gaudi2 chips are awesome & run the multimodal diffusion transformer arch that powers #SD3 faster than H100s (!) in scaled training pre fp8
Way cheaper TCO & Gaudi3 set to be 4x faster..
We also saw 673 tok/s inference on our upcoming StableBeluga 2.5 70b model (!)
(1/3) Today, we're publishing our research paper that dives into the underlying technology powering Stable Diffusion 3.
Prompt: A beautiful painting of flowing colors and styles forming the words “The SD3 research paper is here!”, the background is speckled with drops and splashes of paint.
Stabillity AI presents Stable Diffusion 3
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
Diffusion models create data from noise by inverting the forward paths of data towards noise and have emerged as a powerful generative modeling technique for high-dimensional, perceptual data such as images and videos. Rectified flow is a recent generative model formulation that connects data and noise in a straight line. Despite its better theoretical properties and conceptual simplicity, it is not yet decisively established as standard practice. In this work, we improve existing noise sampling techniques for training rectified flow models by biasing them towards perceptually relevant scales. Through a large-scale study, we demonstrate the superior performance of this approach compared to established diffusion formulations for high-resolution text-to-image synthesis. Additionally, we present a novel transformer-based architecture for text-to-image generation that uses separate weights for the two modalities and enables a bidirectional flow of information between image and text tokens, improving text comprehension, typography, and human preference ratings. We demonstrate that this architecture follows predictable scaling trends and correlates lower validation loss to improved text-to-image synthesis as measured by various metrics and human evaluations. Our largest models outperform state-of-theart models, and we will make our experimental data, code, and model weights publicly available.
Announcing Stable Diffusion 3, our most capable text-to-image model, utilizing a diffusion transformer architecture for greatly improved performance in multi-subject prompts, image quality, and spelling abilities.
Today, we are opening the waitlist for early preview. This phase is crucial for gathering insights to improve its performance and safety ahead of open release.
You can sign up to join the waitlist and learn more here: https://t.co/lX8z1UTc8s #stablediffusion3
Prompt: Epic anime artwork of a wizard atop a mountain at night casting a cosmic spell into the dark sky that says "Stable Diffusion 3" made out of colorful energy
Here's my take on the Sora technical report, with a good dose of speculation that could be totally off. First of all, really appreciate the team for sharing helpful insights and design decisions – Sora is incredible and is set to transform the video generation community.
What we have learned so far:
- Architecture: Sora is built on our diffusion transformer (DiT) model (published in ICCV 2023) — it's a diffusion model with a transformer backbone, in short:
DiT = [VAE encoder + ViT + DDPM + VAE decoder].
According to the report, it seems there are not much additional bells and whistles.
- "Video compressor network": Looks like it's just a VAE but trained on raw video data. Tokenization probably plays a significant role in getting good temporal consistency. By the way, VAE is a ConvNet, so DiT technically is a hybrid model ;) (1/n)
OLMo is here! And it’s 100% open.
It’s a state-of-the-art LLM and we are releasing it with all pre-training data and code. Let’s get to work on understanding the science behind LLMs. Learn more about the framework and how to access it here:
https://t.co/utvPpWwJIp