Tired of blind hyperparameter sweeps? Chimera offers a scientific recipe for scaling visual generation with MoE, KDA/MLA, NoPE & HeteroP. It finds model size matters more than data—and Chimera-11B-A2B matches strong baselines with far less compute.
Interested in how frontier labs pre-train image/video generation models?
We were too.
Since those recipes are rarely made public in full, we started from the most mature pretraining playbook available in the open: how modern LLMs are built.
Introducing Chimera: a visual generation model family that brings LLM-style hybrid linear attention and scaling co-design to visual generation.
In large-scale pretraining, nearly every design choice eventually shows up.
That means solving architecture design and scaling as one coupled problem: every architectural choice changes how the model scales, and scaling behavior determines which choices actually survive.
Chimera approaches both jointly, building a model family that remains predictable as model size, compute budget, and data distribution change.
Our key architectural observation is a simple division of labor: a single raster-ordered KDA stream carries long-range state, while modality-aware short convolutions preserve native local geometry.
Together, they form an effective and elegant linear-attention backbone for multidimensional visual data, with periodic MLA providing direct global interaction and sparse MoE expanding capacity at controlled activated compute.
This design comes with a useful effect: NoPE.
In Chimera, position is represented by the computation itself. KDA’s ordered recurrence and learned state decay encode order and recency, while the short convolutions encode local spatial-temporal structure.
Explicit positional embeddings are not needed in our design.
Because these mechanisms are not tied to a fixed training grid or sequence length, Chimera shows strong zero-shot extrapolation in both space and time.
Trained exclusively on 1K images and 5-second videos, it directly generates coherent 4K images with little visible quality degradation and 30-second videos with only 6.5% FID degradation over the final five seconds, all zero-shot, without resolution- or length-specific finetuning.
But architecture alone is not a pretraining recipe unless it scales predictably.
Thus, scaling should not be treated as an afterthought: fitting a curve over model sizes is easy; making that curve meaningful is much harder.
If every model size is differently under-tuned, your scaling law may simply be measuring optimization error.
We propose HeteroP to transfer proxy-tuned hyperparameters module by module across width and depth, giving us a consistently tuned model family. This allows us to fit Chinchilla-style laws over activated model size, training tokens, and the image-video data mixture.
The laws not only provide the reference for compute-optimal model and data size, but also suggest that visual generation may be more model-hungry than we tend to assume.
Under the same parametric loss-fitting method used in Chinchilla, compute-optimal model size grows as FLOPs^0.516 for images, compared with FLOPs^0.46 for language. Video is even more model-hungry, with the exponent rising from 0.516 for images to 0.544 for video.
Guided by these laws, we trained an 11B-parameter Chimera that activates only 2B parameters per token, using ~600 H100 days.
- It matches Wan-2.1 2B pretraining loss with 7.3x fewer FLOPs, and runs 2.14x faster than full attention at 255K tokens.
- It matches FLUX.1-dev and Z-Image-Turbo on GenEval and outperforms both on DPG-Bench, using roughly 20× less training compute than Z-Image-Turbo.
Ultimately, Chimera indicates that once you pretrain at scale, every decision shows up.
And if the Kimi K3 recipe caught your attention, Chimera may look oddly familiar, except the tokens are pixels and frames with diffusion models.
A team effort from team @ChongjianG30781 , me, @VisionSteve , Jiuxiang Gu, @Xu_Arthas , @chenziwee , @ShaotengLiu , @Jingorz , @YicongHong , @Zefan_Cai , @HaoTan5 ; supported by Hailin Jin and @kalyank_s at @Adobe@AdobeResearch
Introducing Long-LRM++ — for feed-forward, high-res, detail-preserving scene reconstruction
✨ Up to 64 960×540 inputs
🔍 Readable text
📉 4× fewer Gaussians
⚡ Real-time rendering
📷 End-to-end from unposed inputs w/ DA3 poses in 11s (w/⬆️ quality than DA3’s own GS predictor ;)
2⃣ Generative image editing
🤩The SOTA editing models today are great!
🤔BUT they lack precise spatial editing. Text alone is too ambiguous.
💡We let users specify their editing intent, then use a diffusion model to turn it into photorealistic images.
https://t.co/qBM07PZB8X
We are presenting two papers at SIGAsia!
1⃣ Dynamic 3D Recon.
🤩Learning-based recon is amazing!
🤔BUT, they cannot handle casual videos with both camera and object motion.
💡We explicitly learn *object-centric poses* for improved 3D motion modeling!
https://t.co/NB5OToz5T4
Our new work, coupled diffusion sampling, allows fast and diverse multi-view editing!
Not by training — but by having one diffusion model guide another. 🤝
As a bonus: it can construct video-editing datasets
and can make video models generate longer videos.
🧵👇
Excited to share our new work: Generative Video Motion Editing with 3D Point Tracks.
We propose a framework that uses 3D point tracks to precisely edit both camera and object motion in a video, unlocking a wide range of new editing applications.
Checkout our new work on reconstructing 3D objects using photos taken with extremely different lightings! This work was the fruits of my internship at @GoogleDeepMind this past summer!
How can we reconstruct 3D objects under *extreme lighting variations*? 🌤️🌥️🌆🌃🌉
How about ...
🤔 appearance embedding?
BUT it cannot capture view-dependent appearances.
🤔 inverse rendering?
BUT it suffers from ambiguities.
💡 Our idea: Relighting comes to the rescue!
3D illusions are fascinating! 🤩
But it takes exceptional artistic skills to make one.
We present Illusion3D - a simple method for creating 3D multiview illusions, where the interpretations change depending on your perspectives.
Let's play Where's Waldo, shall we? 😆
Excited to introduce our new paper, Generative Omnimatte: Learning to Decompose Video into Layers, with the amazing team at Google DeepMind!
Our method decomposes a video into complete layers, including objects and their associated effects (e.g., shadows, reflections).
I’ll present our paper, In-N-Out (https://t.co/5J3eOKaSrB) 5pm-6:30pm today at Arch 4A-E#236. #CVPR2024
The paper is about how to achieve faithful semantic editing for Out-of-Distribution input.
Stop by our poster and say hello! ☕️
I’ll be talking VideoGigaGAN at AIS workshop (https://t.co/SEEkOGgYDI) today at 3:40pm at Arch 3A. Feel free to come if you’re interested in high frame quality Video Super-Resolution. #CVPR2024
https://t.co/EPDZVarLzX
Happy to see you again, Seattle! #CVPR#CVPR2024
I’ll present one paper, In-N-Out, at 5pm Wed (Arch 4A-E Poster #236), and give a talk about my recent VideoGigaGAN at AIS Workshop at 3:40pm Monday (Arch 3A).
I'll also be in the job market next year. DM me if you want a chat!☕️
Attending CVPR next week?
Make sure to stop by our oral presentation of
*Seeing the World through Your Eyes*
📅 Wed 19 June 16:36 - 16:54 EDT (Orals 2C)
@HadiZayer and @kevinzhang25 will tell you how we reconstruct the Kirby!