New preprint! Introducing MaTS - a new framework for merging individual task models into a multitask model by matching them in their task subspace
Work done w/ @mohitban47 @colinraffel
📄 https://t.co/VwA1DT85KF
💾 https://t.co/HnuOFVL2gi
🧵 ⬇️
Excited to announce that I will be presenting Model Merging via Data-Free Covariance Estimation at COLM '26!
Thanks again to my amazing co-authors @dtredsox13@PTikeng Colin Raffel and Guillaume Rabusseau
New preprint! Introducing ACTMat: Model Merging via Data-Free Covariance Estimation
Work done w/ @dtredsox13, @PTikeng, Colin Raffel and Guillaume Rabusseau
TL;DR: It's RegMean, but without needing data for covariance estimation (C ≈ Δᵀ Δ)
📄 https://t.co/qB0oafSxvq
🧵(1/6)
Prior to Biden taking office in 2020, less than 1% of @NSF grants were related to DEI. Now that number exceeds 10%. These grants prioritize aspects other than scientific merit, and subsidize woke ideology with taxpayer dollars.
@rupasubramanya reports ⬇️
https://t.co/nRYOQp7Ngu
🚨 Model Merging competition @NeurIPSConf!🚀
Can you revolutionize model selection and merging?Let's create the best LLMs!🧠✨
💻Come for science
💰Stay for $8K
💬Discord: https://t.co/eGgyBifqeq
🔗Sign up: https://t.co/afTxLA1jvi
Sponsors: @huggingface@SakanaAILabs@arcee_ai
Is Kevin onto something? We found that LLMs can struggle to understand compressed text, unless you do some specific tricks. Check out https://t.co/DRO2IbTFCg and help @hoonkp, @alemi, Jeffrey Pennington, @ada_rob, @jaschasd, @noahconst and I make Kevin’s dream a reality.
We bring you assorted methods for merging LoRA adapters in 🤗 PEFT.
These methods, for now, support both text and image generation use cases. Methods: TIES, DARE, and more! This is now supported for both Transformers and Diffusers 💡
Details 👉
https://t.co/7Q7mPUUbtI
1/4
🙋♂️Want to prune your large LVLM effectively and efficiently?
We are excited to share our #ICLR2024 paper and introduce ECoFLaP, a coarse-to-fine pruning method to effectively prune large models (both multi-modal & uni-modal)!
https://t.co/bsKOBQsmlt
@jaeh0ng_yoon@mohitban47
🧵
Feels like a good time to promote Tam et al., 2019 "Optimal Transport-based Alignment of Learned Character Representations for String Similarity".
This is one of multiple papers that encode string comparisons, e.g. edit distance, into neural nets.
https://t.co/C3atixCxP3
New preprint! Introducing MaTS - a new framework for merging individual task models into a multitask model by matching them in their task subspace
Work done w/ @mohitban47 @colinraffel
📄 https://t.co/VwA1DT85KF
💾 https://t.co/HnuOFVL2gi
🧵 ⬇️
MaTS outperforms closed-form solutions to the linear system and achieves state-of-the-art results for multitask merging. More details in our paper (https://t.co/VwA1DT85KF) and code (https://t.co/HnuOFVL2gi)!
Check out the camera-ready version of TIES-Merging to be presented at @NeurIPSConf 2023!
We have added more experiments on
1. Merging for robustness on a single task.
2. Merging for better Initialization and Finetuning
3. We show that interference exists even when merging models of the same task.
4. TIES-Merging is more robust to variability in the merging coefficient.
5. Updated our GitHub to provide a minimal example to run TIES-Merging on your models.
We thank the #NeurIPS2023 Reviewers & AC for their hard work!
https://t.co/0lAXG3118n
@dtredsox13@LChoshen @colinraffel @mohitban47
Thanks for the shoutout + covering our #ACL2023nlp work on MeetingQA, @JayAlammar@cohere@forai_ml! It was great interacting with you 😃
PS: For those interested, details at 👉 https://t.co/t706r2ZxLL
cc/ Trung, @david_s_yoon, Hanieh, @FranckDernoncou and @mohitban47
Our work on Data Augmentation for Learning from Limited Data has been accepted to #TACL! We are presenting it at #ACL2023 on Wed 11:00-12:30 in Session 7.
Paper: https://t.co/wK3bplmCp4
Poster + Video: https://t.co/AyixNzrANk
@jiaao_chen @colinraffel @mohitban47@Diyi_Yang
Data augmentation has been one of the most common approaches for mitigating the need for labeled data&improving data efficiency. We provide an empirical*survey of data augm for limited data learning in NLP: https://t.co/y4jTDOSgTg
w/ Derek Tam @colinraffel @mohitban47@Diyi_Yang
Correction: My #ACL2023 poster on (un)faithfulness of extractive summarization was moved to Poster Session 3 at 09:00-10:30 AM EDT in Frontenac Ballroom and Queen's Quay. See you tomorrow!
@byryuer@mohitban47 @uncnlp