Diverse Video Generation with DPP-Guided Policy Optimization is accepted to #CVPR2026 ✨
We tackle diversity in T2V models by formulating diversity as a set-level optimization problem.
Huge thanks to my advisor @PINguAR and my collaborator Connor Dunlop
✨We introduce Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization: a framework for diverse video generation that combines Determinantal Point Processes and GRPO theories to enforce explicit reward on diverse generations.
#CVPR2026@TKazimi415 presented collaborative work w/Connor Dunlop, fellow Ph.D. student @SanghaniCtrVT, & advisor Pinar Yanardag on a framework that enables text-to-video models to produce diverse yet prompt-faithful videos from a single input prompt.
https://t.co/cWxWP5n30t
✨ We introduce VideoMLA, an MLA-style latent KV cache for autoregressive video diffusion.
Multi-Head Latent Attention (MLA) has been central to the efficiency of DeepSeek-V2 and V3. In this work, we study whether the same principle can be transferred from language modeling to long-horizon video generation.
Game engines let you spawn anything into a scene. World models don't. Once the camera moves, you're stuck with whatever the model dreamed up.
🧵 We fix that with SPAWN, a training-free method that injects a custom concept into the rollout from either an image or a text prompt.
Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization
Main proceedings
@TKazimi415 Connor Dunlop @PINguAR (all VT)
https://t.co/cWxWP5n30t
TL;DR: Enables text-to-video models to produce diverse yet prompt-faithful videos from a single input prompt.
GEMLAB is heading to #CVPR2026! We’re excited to share our papers and talks across controllable generative AI, long-horizon video modeling, and personalization! #CVPR2026#CVPR
Full schedule: https://t.co/To0i22tYog
Proud to share: GEMLAB has 3 main-track and 1 Findings accepted to #CVPR2026! More details soon. Huge congrats to my amazing students & collaborators. See you in Denver!
P.S. GEMLAB is keeping the tradition alive 3 years in a row with 3 papers at every CVPR 🚀 https://t.co/L8GfK7kAeE
So grateful for the incredible students and alumni of GEMLAB! Together, we published 4 main conference papers + 4 workshop papers at @NeurIPSConf. You’re all absolute gems 💎
Thanks to my collaborator Connor Dunlop, and our advisor @PINguAR for her invaluable insights.
Check out our project page!
🌐 Project Page: https://t.co/cwjQU8soGw
Paper: https://t.co/4easn51Bg2
Huggingface: https://t.co/T4R5BUrskU
✨We introduce Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization: a framework for diverse video generation that combines Determinantal Point Processes and GRPO theories to enforce explicit reward on diverse generations.
Our method improves video diversity on Wan2.1 and CogVideo, enhancing motion, temporal variation, and scene diversit, while keeping CLIP fidelity stable. It also raises video-quality scores, showing DPP-GRPO boosts diversity without sacrificing alignment.
Check out our work and let us know what you think!
Project Page: https://t.co/dSkaLHCxyN
Arxiv: https://t.co/XTxTudNx54
HuggingFace: https://t.co/lGyjgg9i9u
We introduce Audit & Repair, a collaborative multi-agent framework that autonomously identifies, corrects, and refines inconsistencies across multi-panel story visualizations.
This is a joint work with @akdemir_kiymet, supervised by @PINguAR