💡Which Heads Matter for Reasoning? RL-Guided KV Cache Compression
We introduce RLKV, a novel reasoning-critical head identification method by reinforcement learning, to guide KV cache compression for reasoning LLMs.
Paper: https://t.co/Ng3rhkjBzU
Page: https://t.co/27FNGLdw8i
Heading to #ICML2026 in Seoul 🇰🇷with 3 students. See you soon my friends chatting efficient AI chewing kbbq 😃
Check the collection of our 4 papers here: https://t.co/OAc1xNWqg2
✈️Heading to #ICLR2026 in Rio with 4 students. Gonna present 4 poster papers and 2 workshop papers from our lab. See you there, my friends, in Rio! 👋🏻🍺
1 poster in the 4/23 morning 10:30AM, 3 in the afternoon 3:15PM.
Paper keywords: pruning, diffusion models, fine-grained visual reasoning, mixup, MLLMs, autoregressive image gen, parallel decoding, kernel generation on mobile devices, reasoning, KV cache compression.
1. OBS-Diff: Accurate Pruning For Diffusion Models in One-Shot
🌟 TLDR: We present an OBS-based structured pruning for DiT models, good results, no finetuning.
🌐 Webpage: https://t.co/vQ7D1ck5Dn
📝 Arxiv: https://t.co/iuaIGBRSiT
💻 Code: https://t.co/sYyHn64ovx
🤗 HF ckpts: https://t.co/sGrZYAyQds, https://t.co/4DaV6vjEqP
���️📍Presentation: 4/23, 3:15PM, P4, #3014 (https://t.co/BzgkMk1Mlv)
2. RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
🌟 TLDR: We introduce a multi-stage RL method designed for fine-grained visual reasoning; tested on our newly minted map dataset, ReasonMap.
🌐 Webpage: https://t.co/Dw4fkRFzhr
📝 Arxiv: https://t.co/mR2wyAqOAH
💻 Code: https://t.co/90CtXw0WMr
🤗 HF dataset: https://t.co/wm06kp2vHD
🗓️📍Presentation: 4/23, 3:15PM, P4, #3409 (https://t.co/iO4fO2TNll)
3. MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
🌟 TLDR: We propose MergeMix, a unified paradigm that bridges SFT and RL with an efficient token-merge-based mixup augmentation.
🌐 Webpage: https://t.co/oeAmKBfJJD
📝 Arxiv: https://t.co/tr8qAfvIGM
💻 Code: https://t.co/4LoDuK0bc6
🗓️📍Presentation: 4/23, 10:30AM, P4, #3401 (https://t.co/oTqkPCwGX0)
4. Autoregressive Image Generation with Randomized Parallel Decoding
🌟 TLDR: (the title tells pretty much :))
🌐 Webpage: https://t.co/ZSr2CJhYeD
📝 Arxiv: https://t.co/02g0B8ofnJ
💻 Code: https://t.co/J93ROnPj3y
🗓️📍Presentation: 4/23, 3:15PM, P4, #3010 (https://t.co/R28Z5uhrcD)
5. (Workshop) MobileKernelBench: Can LLMs Write Efficient Kernels for Mobile Devices?
🌟 TLDR: Can LLMs Write Efficient Kernels for Mobile Devices? We built a dataset, benchmark, and made the very first attempt to answer the question. An agentic pipeline MoKA is proposed.
🌐 Webpage: https://t.co/ygcCD0eOKE
📝 Arxiv: https://t.co/L3nyefhxtC
💻 Code: https://t.co/jyPh0JK70S
🗓️📍Presentation: 4/26, Room 203 A+B, 9AM-5PM, #52 (https://t.co/yJcZqcpyxu)
6. (Workshop) Which Heads Matter for Reasoning? RL-Guided KV Cache Compression
🌟 TLDR: Some heads matter more for reasoning tasks (math & code). Based on this discovery, we developed a KV cache compression method via RL training.
🌐 Webpage: https://t.co/IgwNKsGgP4
📝 Arxiv: https://t.co/6X6a7IoXYB
💻 Code: https://t.co/x7zQScnfft
🗓️📍Presentation: 4/27, Room 101A, 9AM-5PM, #33 (https://t.co/LTzSoIJaGA)
Thanks to the lead authors @Alright_lone (@Westlake_Uni ), @si_feng32704 (NUS) & Kaiwen Tuo (Tongji U), @Xander_K1ng (Westlake), @haopengl33 (HKUST), Xingze Zou & Jing Wang (ZJU), @wjdu24 (NTU).
Sicheng, Kaiwen, Xin, and Haopeng will also be at Rio to present the papers. Welcome to shoot them with questions at the poster sessions :)
Unfortunately, @Alright_lone cannot attend due to exams... :(
#AI #ICLR #LLM #MLLM #PhD #Research
AReaL v1.0 released: Effortless #RL to make your #OpenClaw self-evolve 🚀:
•🛠️ One-click agentic RL for any existing agent
•📈 Open-source SOTA on tau2-bench
•💎 A new PyTorch-native 5D-Parallel Engine Archon
•🤖A full #opencode recipe
GitHub: https://t.co/bHaE6lRzXt
Just watched The Thinking Game — an incredible documentary.
Feeling truly fortunate to be part of this rising wave of the “thinking game” in my twenties.
Highly recommended. https://t.co/3CkGm4YMwq via @YouTube
💡Which Heads Matter for Reasoning? RL-Guided KV Cache Compression
We introduce RLKV, a novel reasoning-critical head identification method by reinforcement learning, to guide KV cache compression for reasoning LLMs.
Paper: https://t.co/Ng3rhkjBzU
Page: https://t.co/27FNGLdw8i