๐ New paper alert: On-Policy Self-Distillation without Any Supervision !!
๐ง Can a model truly on-policy โself-distill"?
On-policy self-distillation removes the need for a stronger teacher, but still relies on an external ground-truth solution to make the self-teacher more capable.
Our new work asks: can we remove supervision altogether?
Can an LLM improve its reasoning entirely from its own generationsโwithout ground-truth answers, verifiers, or a stronger teacher?
โป๏ธ We introduce U-OPSD, a simple approach toward genuine self-distillation:
For each problem, the model samples multiple solutions from itself:
๐ณ๏ธ Agreement identifies a likely solution through majority vote
๐ Disagreement reveals trajectories where the model may be going wrong
๐ The pseudo solution conditions the self-teacher, which provides dense token-level guidance along those disagreeing trajectories
U-OPSD turns self-consistency into supervision: agreement builds the teacher context, while disagreement determines where to distill.
๐ No external supervision at all
๐ Unlabeled problems only:
๐ซ GT solutions
๐ซ Verifier
๐ซ External strong teacher
๐ซ Environment feedback
๐ Yet, across 5 math reasoning benchmarks and 6 Qwen3 settings, U-OPSD consistently improves the base model and matches or surpasses supervised SFT, GRPO, and OPSD.
๐ Non-thinking mode:
๐ +8.5 / +10.7 over the 4B / 8B base models
๐ +3.2 / +2.3 over GT-supervised OPSD
๐ +7.0โ11.3 over label-free self-rewarding RL methods including TTRL, RENT, and Intuitor
๐ง Thinking mode:
๐ +2.2 / +1.9 over the 4B / 8B base models
๐ค On par with GT-supervised OPSD and GRPO
๐ +0.8โ1.4 over label-free self-rewarding RL methods including TTRL, RENT, and Intuitor
Meet U-OPSD ๐
๐ ArXiv Paper: https://t.co/sGBO2uT8DQ
๐ป GitHub: https://t.co/T5eZtbUlYJ
๐ Project Page: https://t.co/7oDlOLCPK2
๐ค Hugging Face Paper: https://t.co/hWWVJoISOw
(Plz upvote if you can !!
Amazing collaboration with @ Bingyang @JoLiang17@tian_yunjie@Di Fu and Nuno !
Grateful for all the insightful discussions and exploration together!
#LLM #OPD #OPSD #Distillation
๐Do we really need to give models more information to help them learn?
Recent on-policy self-distillation methods rely on answer hints, reasoning traces, or visual evidence to construct a better learning signal.
We wondered:
โWhat if none of them were necessary?
๐We introduce Visual Contrastive Self-Distillation (VCSD), a simpler form of OPSD driven purely by input conditioning.
Instead of adding supervision, we compare the same EMA teacher under two matched visual conditions:
๐ผ๏ธ Original image
โฌ Content-erased control
Their prediction difference reveals what the image actually contributes, and becomes the self-distillation signal. By suppressing shared language priors and emphasizing visually supported tokens, VCSD rebalances the distillation target toward the visual modality.
โ No external teacher
โ No answer hints
โ No reasoning traces
โ No visual evidence
โ No additional inference-time cost
๐ Across 6 model scales and 7 vision-language benchmarks, VCSD consistently outperforms matched OPSD, with gains of up to +4.5 points.
Meet VCSD ๐
๐ Project page: https://t.co/pIFQI0a1vL
๐ Huggingface Paper: https://t.co/MxaYrk1Lrb
Amazing collaboration with @tian_yunjie , @Williamiumli , @Yuqi_Jia7 , @zhoutianyi , @furongh And Di Fu!
Grateful for all the insightful discussions and exploration together!
@umdcs@ucsd_cse
Excited to share this work that I've been fortunate to collaborate on!
We propose a scalable way to collect training-time reward-hacking trajectories and show that they lead to substantially stronger monitors for detecting real inference-time reward hacking.
Big thanks to all my collaboratorsโthis was a pleasure to work on๐
Many monitors are trained & evaluated on prompt-elicited hacking trajectories, where models are explicitly asked to exploit the reward signal. But the real test is whether they catch the training-time hacks that naturally emerge during RL training without hacking instructions.
โ๏ธWe find a large mismatch:
๐ Detection accuracy: 97.1% on prompted hacks โ BUT 28.0% on training-time hacks.
๐ To study this, we introduce Trace-and-Amplify (TA), a framework for collecting training-time reward-hacking trajectories at scale without hacking instructions.
โ Monitors trained on TA-collected trajectories perform much better on real inference-time hacks, and generalize better to held-out hacking types:
Detection accuracy 59.98% (PE-trained) โ 90.16% (TA-trained).
๐ Project: https://t.co/unI5L32fRH
๐ Paper: https://t.co/Q7NXjg1ERU
Joint work with @LilichenLi3146, @hgzhou42, @JoLiang17, @zhoutianyi, and @cho_jui_hsieh at PKU, UCLA, and UMD
@PKU1898@umdcs@UCLACS
Grateful for everyoneโs contributions and support!
Many monitors are trained & evaluated on prompt-elicited hacking trajectories, where models are explicitly asked to exploit the reward signal. But the real test is whether they catch the training-time hacks that naturally emerge during RL training without hacking instructions.
โ๏ธWe find a large mismatch:
๐ Detection accuracy: 97.1% on prompted hacks โ BUT 28.0% on training-time hacks.
๐ To study this, we introduce Trace-and-Amplify (TA), a framework for collecting training-time reward-hacking trajectories at scale without hacking instructions.
โ Monitors trained on TA-collected trajectories perform much better on real inference-time hacks, and generalize better to held-out hacking types:
Detection accuracy 59.98% (PE-trained) โ 90.16% (TA-trained).
๐ Project: https://t.co/unI5L32fRH
๐ Paper: https://t.co/Q7NXjg1ERU
Joint work with @LilichenLi3146, @hgzhou42, @JoLiang17, @zhoutianyi, and @cho_jui_hsieh at PKU, UCLA, and UMD
@PKU1898@umdcs@UCLACS
Grateful for everyoneโs contributions and support!
๐ Before an AI can explore the world, it must learn what questions are worth asking.
โ โWhat color is the fireplace?โ
โก๏ธ โIs the object on top of the fireplace shown in the mirror?โ
๐ผ๏ธ Same image. Completely different supervision.
Only one forces a model to inspect, ground, reason, and explore.
Today's Vision-Language Models (VLMs) are trained primarily as answerers. But intelligent systems should also know how to ask good questions.
Most VLMs learn from fixed question distributions. We let the question distribution evolve.
We introduce Self-Evolving Visual Questioner (SeeQ), a framework that enables a VLM to continuously improve itself as a visual questioner using its own generated QA data.
โ No human annotations
โ No external teacher models
โ No reward models
โ ๏ธ A key challenge is collapse: without careful control, self-generated questions quickly converge toward repetitive and low-information patterns.
๐ Our self-evolving proposal โ rewrite โ filter loop maintains exploration while continuously pushing the model toward broader, harder, and more visually grounded questions.
๐ To measure this evolution, we built an evaluation framework for questions that demand:
๐ Harder visual search
๐ฏ Tighter grounding
๐ง Deeper contextual and spatial reasoning
๐ Greater diversity
๐ After just two rounds of self-evolution:
โฌ๏ธ Question quality improves by 82%
๐ช QA performance is preserved on VStar, CVBench, and RealWorldQA
Meet SeeQ ๐
๐ Project page:
https://t.co/a0HIK2Qwep
๐ Arxiv:
https://t.co/RiEoo0cQI1
๐ป Codebase:
https://t.co/8cvbxqzKpJ
Many thanks to the collaborators Hengguang @hgzhou42 , @Ming_Liiii , Lichen Li, @cho_jui_hsieh and @zhoutianyi at UMD and UCLA @umdcs@UCLACS
Hot take: robots should not dream in pixels.
Pixels are too low-level.
Latents are too opaque.
ฮผโ predicts a third thing:
3D motion traces.
On real robots, it beats ฯโ.โ โ with ~1/100 the data scale and no action labels for world-model pretraining. ๐งต
https://t.co/UfmrqNlBtw
(this video features voiceover narration)
๐ Presenting our work on 21st Oct at #ICCV2025 Honolulu! ๐๐๏ธ ๐บ Find us at Poster Exhibit Hall-1 #151 or DM to chat
Diffusion Curriculum: Synthetic-to-Real Data Curriculum via Image-Guided Diffusion
(1/3)
๐ Website: https://t.co/AAOyZHtYJe
๐ Paper: https://t.co/I5AdMA6AC5
๐Real-world data is noisy, imbalanced & scarce โ brittle models
๐จT2I synthesis boosts diversity but suffers from a large syn-to-real gap
โจDisCL leverages a spectrum of syn-to-real data with image guidance and curriculum to progressively bridge this gap
(2/3)
๐จColor plays a critical role in human perception and reasoning in daily applications, e.g., medical test kit readings, shopping, wild animal study, and agriculture.
But is it so for existing #AI and #VLMs? We propose๐ฅColorBench, which exposes weaknesses of existing๐คVLMs๐