🎉Glad to announce our survey on Mixup are now released on arXiv🌠
-Nearly 220 papers with lots of training paradigms, data modalities
-Summarize papers based on improvement
-Reformulated mixup as the framework
-Appeal designing a mixup unified framework to reduce the gap
#Mixup
Heading to #ICML2026 in Seoul 🇰🇷with 3 students. See you soon my friends chatting efficient AI chewing kbbq 😃
Check the collection of our 4 papers here: https://t.co/OAc1xNWqg2
@Shiwei_Liu66@LupinLSY The SGG (https://t.co/HQsAp029Pj) also employs a Layer-wise LR adaptive technique. This looks like a useful technique for LLMs & MLLMs beyond vanilla AdamW or Muon.
Thrilled to share our #ICML2026 work! We rethink optimizer design by leveraging the row-block diagonal dominance of the Transformer Hessian. We show that simple row-normalization (RMNP) is equivalent to Muon's orthogonalization.
Full story in the thread! 👇
1/n Please stop by👋. This is not just another ICML 2026 optimizer paper. We have rich intuition to share on why simple preconditioners like orthogonalization and row-normalization specifically benefit NNs optimization. Quick overview below 🧵
Liner is partnering with @spoticlr at #ICLR2026 — supporting Best Paper and Travel Awards for LLM research.
And to celebrate, we're giving away:
✈️ Round-trip flights + hotel to #ICML2026 in Seoul
🎁 $300 Liner Credits
Follow @search_liner + repost to enter by 4/27.
Liner is built for research workflows. Find papers, verify sources, and write with citations in one place.
See you in 🇧🇷 and 🇰🇷!
@iclr_conf@icmlconf
One static model does not fit all😭
We just dropped our latest work: Functional Neural Memory. Instead of static models, we generate custom "parameters" for every single input.
✅Prompt your model anytime
✅Instant personalization
✅Better instruction following
✅Flexible & dynamic memory (w/o memory bank✌️)
(🧵1/6)
🚀 dVoting
We introduce dVoting, a fast voting strategy that boosts reasoning capability in dLLMs without any training.
- Training-free & Simple
- +7.66% on GSM8K, +7.20% on MATH500
- First efficient test-time scaling strategy for dLLMs
Page: https://t.co/QkysyekJn2
💡Which Heads Matter for Reasoning? RL-Guided KV Cache Compression
We introduce RLKV, a novel reasoning-critical head identification method by reinforcement learning, to guide KV cache compression for reasoning LLMs.
Paper: https://t.co/Ng3rhkjBzU
Page: https://t.co/27FNGLdw8i