💫Editing in 3D should feel as simple as moving a box.📦
Meet Thinking in Boxes: 3D Editing in Real Images Made Easy - fit 3D Boxes to any object in any image, drag/rotate/scale them or change the viewpoint, and get a photorealistic edit!
🧵👇
Demand more variety from your models!🚀
Dont Settle at the Mode!✨
Flow models suffer from diversity collapse--generating from the same mode across different seeds. Ask FLUX for "a photo of a bear" N times & you'll get similar looking images.
Our ECCV 2026 paper fixes this! 🧵👇
It also generalizes beyond FLUX models–Sana, Qwen-Image, and the timestep-distilled FLUX.2-Klein, all suffer from diversity collapse, and the same fix works!
@weikaih04 Thank you! We didn't use WildDet3D for training because our model requires paired training data (before/after edit), which isn't present in the dataset. However, we did adapt a small subset for evaluation, and those results are available in the paper. Thank you for the dataset!
💫Editing in 3D should feel as simple as moving a box.📦
Meet Thinking in Boxes: 3D Editing in Real Images Made Easy - fit 3D Boxes to any object in any image, drag/rotate/scale them or change the viewpoint, and get a photorealistic edit!
🧵👇
Trained with synthetic multi-object scenes🤖 and a small subset of real video frames🎬 — it generalizes to complex spatial edits for in-the-wild images🏞️!