Can a better harness, not better weights, close the gap to frontier image models?
On multi-reference image generation, a 4B model becomes competitive with Nano Banana Pro.
A really fun project co-led with @shimao0114 🚀
Our new preprint on harness optimization for image generation agents is out 🎉
A small open 4B model, competitive with Nano Banana Pro on multi-reference image generation — with no fine-tuning, just a better harness.
Harness optimization (letting a coding agent rewrite the code around a model) works well where answers can be checked, like code and math. Image generation relies on a noisy AI judge instead, and one lucky score can mislead the search.
AutoRef handles this with beam search: it keeps several promising harnesses instead of picking just one.
One discovered harness:
1️⃣ makes FLUX.2 [klein] 4B competitive with Nano Banana Pro
2️⃣ still works with a different number of reference images (searched with 4, works with 3–5)
3️⃣ works as-is on other image generation models, including a larger one (FLUX.2 [klein] 9B) and a different family (Qwen-Image)
Paper: https://t.co/DYZxg7PJct
Code: https://t.co/oiP8KO1XLT
w/ @Ku_Onoda@yusuke_iwasawa_@szk_masa@ymatsuo@frt03_
1/9 🎉 Our paper "Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation" was accepted at #WACV2027 in Round 1!
An RL post-training method for generative models: we train sets of images to cover multiple target modes, letting different images represent different possibilities.
8/9 See also very similar concurrent work by
@AntChen_ et al.: ROSA https://t.co/Mfa4ESn0P8
@RyanBoldi, @ishapuri101 et al.: VPO https://t.co/looZbxlNRr
Presenting our work today at #ICLR2026! 🚀🇧🇷
Better policy gradients for differentiable simulators. Come say hi!
Paper: https://t.co/Y05FYDYIUV
⏰ 10:30 AM – 1:00 PM
📍 Pavilion 4, Poster 4504