Introducing LTX-2.5, the world model the world builds on.
One of the biggest upgrades yet to the model already powering film, robotics, and real-time workflows. Higher pixel fidelity, multishot scenes that hold together across cuts, and a pretrained foundation built to be finetuned across domains.
Plus: Diffusion Fidelity Rendering, a new rendering approach built for pixel quality that holds up frame by frame, even on a cinema screen.
This isn't a tool you rent. It's a foundation you build on, open and yours.
1/7
New paper 🧵 ScenA generates a multi-speaker audio scene: overlapping speech, laughter, real room noise, from a text description and a few reference voices.
When trained in the obvious way, it ignores the text and decides who speaks on its own. We found out why and fixed it.
Our paper:
"LaMI: Augmenting Large Language Models via Late Multi-Image Fusion"
has been selected for an Oral Presentation at #ACL2026!
LaMI boosts LLM visual commonsense by generating complementary images from a text prompt and late-fusing their evidence into the prediction
🧵
We introduce 🌍GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens.🌍
Most feed-forward 3DGS methods still start from pixel, voxel, or dense view-aligned primitives.
We take a different route: align first, decode later. 🧵👇
Excited to announce that "Let it Snow!" has been accepted to #CVPR2026!🎉
We present a framework for scene-wide dynamic editing of static 3D Gaussian Splatting scenes with dynamic weather effects.
https://t.co/578C2u20hm
[1/5]
LTX-2.3 is here.
For decades, creative software has been defined by its interface. We think the next era gets defined by the engine underneath.
LTX-2.3 is a major engine upgrade:
→ Sharper detail
→ Stronger motion
→ Cleaner audio
→ Native vertical format
We present DyPE, a framework for ultra high resolution image generation.
DyPE adjusts positional embeddings to evolve dynamically with the spectral progression of diffusion.
This lets pre-trained DiTs create images with 16M+ pixels without retraining or extra inference cost.
🧵👇
Introducing LTX-2: the most complete open-source AI creative engine.
- Synchronized audio and video generation
- Native 4K fidelity, up to 50 fps and 10 s+ sequences
- API-first design for seamless integration into creative pipelines
- Runs efficiently on consumer GPUs
- Fully open and accessible, with weights releasing later this year
[1/10] 🤔 What if you wanted to generate a 3D model of a “Bolognese dog” 🐕 or a “Labubu doll” 🧸?
Try it with existing text-to-3D models → they collapse.
Why? These concepts are rare or new, and the model has never seen them.
🚀 Our solution: MV-RAG
See details below ⬇️
[1/6] 🎬 New paper: Story2Board
We guide diffusion models to generate consistent, expressive storyboards--no training needed.
By mixing attention-aligned tokens across panels, we reinforce character identity without hurting layout diversity.
🌐 https://t.co/aRG81nu5qK
[1/10]🚨 Introducing RewardSDS! 🚨
Standard SDS-based text-to-3D methods struggle with fine-grained alignment to user intent, often leading to artifacts or misaligned generations.
Our solution 👇