Just last week @leokarlin and @Eli_Schwartz presented our BrAD paper at @CVPR, and now we are releasing our first demo for Unsupervised Cross-Domain Image Retrieval on @Huggingface ๐ค with @Gradio!
Come check it out! โค๏ธ
https://t.co/P0OZgKJ5d9
@IBMResearch
๐จ Rough luck with your #ICCV2025 submission?
Weโre organizing the 4th Workshop on Whatโs Next in Multimodal Foundation Models at @ICCVConference in Honolulu ๐บ๐ด
Send us your work on vision, language, audio & more!
๐๏ธ Deadline: July 1, 2025 ๐ https://t.co/t2DmcZAlWM
IPLOC accepted to ICCV25 โบ๏ธ
Thanks to all the people that were part of it ๐ฉท
The idea for this paper came by a lake during a visit to Graz for a talk. It has traveled with me through too many countries and too many wars, and itโs now a complete piece of work.
IBM released Granite-Vision-3.1-2B, a small vision LM with impressive performance on different tasks ๐ฎ๐ฅ
it comes with transformers and vLLM support from the get-go ๐
you can run it in Colab T4, so I built a notebook to put it to test, find it on the next one โคต๏ธ
๐ฏ Introducing Sparse Attention Vectors (SAVs): A breakthrough method for extracting powerful multimodal features from Large Multimodal Models (LMMs). SAVs enable SOTA performance on discriminative vision-language tasks (classification, safety alignment, etc.)!
Links in replies!
๐Using just ~20 attention heads & only few-shot examples, SAVs:
- Outperform both LoRA and few-shot baselines
- Work with image, text, & interleaved inputs
- Extract features without finetuning - ready to go at test time!
This project was a cross-collaborative effort between researchers from UC Berkeley, Carnegie Mellon University, and MIT-IBM Research (@berkeley_ai, @CMU_Robotics, @MITIBMLab). Many thanks to all of the collaborators and co-authors on this work:
Brandon Huang, Tianning (Ray) Chai, @ZhiqiuLin, @ArbelleAssaf, @RogerioFeris, @leokarlin, @trevordarrell, @RamananDeva, @roeiherzig
There are some approaches to improve few-shot performance on VLMs: from simple fine-tuning on more images to using newer techniques such as Task Vectors - compact implicit representations of in-context examples compressed in the modelโs attention heads (https://t.co/aVfwb091qy).
๐จ๐๐ฎ๐ฅ๐ญ๐ข๐ฆ๐จ๐๐๐ฅ ๐๐๐ฌ๐ค ๐๐๐๐ญ๐จ๐ซ๐ฌ (MTV) is accepted to @NeurIPSConf!๐ฅณ
We show how to perform many-shot ICL in LMMs using MTV representation---a compact implicit representation of many-shot examples stored in the modelโs attention head.
Kudos to all collab.&authors
With the recent widespread discussions on X about OpenAI's newest o1 model๐arithmetic skills, this feels particularly timely. Excited to share that NumeroLogic, our work on improving LLMs' numerical reasoning capabilities, has been accepted to @emnlpmeeting! See you in Miami! ๐ฅณ
Tokenization algorithms like BPE are the standard practice in NLP ๐
However, they have been mostly overlooked for discrete acoustic units (DAUs) ๐๏ธ
We show that tokenizing DAUs yields significant task performance and speed improvements. ๐
Accepted to #Interspeech2024 ๐ฌ๐ท
This fantastic work was done by the outstanding two undergrads, Brandon Huang and Chancharik Mitra, as well as @ArbelleAssaf, @leokarlin, and @trevordarrell.
and..Happy 4th of July!๐บ๐ธ
Trans-LoRA
towards data-free Transferable Parameter Efficient Finetuning
Low-rank adapters (LoRA) and their variants are popular parameter-efficient fine-tuning (PEFT) techniques that closely match full model fine-tune performance while requiring only
Thanks for the highlight @_akhaliq!
We offer a simple and nearly-data-free way to move (large quantities) of custom PEFT models within or across LLM families or even across PEFT configurations. Useful for LLM cloud hosting when old base models need to be deprecated & upgraded
๐จ @CVPR Challenge alert!๐จ
Think you have a good MML for documents, infographics and text-rich images?
Challenge โWhat is Next in Multimodal Foundation Models?โ extended the deadline and opened for all:
https://t.co/e5wt5EKVPC
$10K in prizes
Deadline: 5th June 2024 (11 days)!
In case you haven't noticed, we have $10K winner prizes for the top teams!
๐ค๐ค๐ค
Also, we extended the deadlines for Phase 1 and Phase 2 to June 5th.
Challenge: https://t.co/1PCRHsP6Cc
Git: https://t.co/j5rONVxu4r
@CVPR