๐ Generative video grounding decodes the spatial tube box by box, in sequence. A 100-frame tube costs hundreds of decoding rounds, and attention drifts off the video as it goes.
Parallel Tube Decoding does the whole tube in 2 rounds. 79ร faster, better accuracy, 4B params.
1/2
@peter_richtarik Maybe AI conferences need to allocate some budget to incentivize good reviewers (not just top reviewers), and something more than a free registration.
Knock knock! ๐ช Who's there? ๐ค It's GLaMM! ๐
Excited to share that our paper, GLaMM: Pixel Grounding Large Multimodal Model, will be presented at CVPR 2024! ๐ Tomorrow at 10:30 AM, poster #326! ๐ฅ
GitHub: https://t.co/IXnOR72Oxa
Meet the team at Arch4A-E, poster #326! ๐ฉโ๐ป๐จโ๐ป
@CVPR Hi Everyone attending @CVPR
While asking for the poster, try asking the poster by giving the author name who submitted the poster, not always the first author name.
Good Luck #CVPR2024
Hi Everyone attending @CVPR
While asking for the poster, try asking the poster by giving the author name who submitted the poster, not always the first author name.
Good Luck #CVPR2024
โกHow can complimentary strengths of image & video encoders help Video-LMMs? We present new results + diverse instruction data + benchmark!
๐ VideoGPT+ : https://t.co/EM5ns1uKn0
LLaVA++ based on Phi-3: https://t.co/shFo2aqWmf
LLaVA++ based on Llama 3: https://t.co/wK8ILGhVO9
So many fine-tuned variants in this org: https://t.co/AJaJz9Wqwb
๐Exciting updates to our recent effort to extend LLaMA3 and Phi3 for *visual* understanding. Enjoy!
๐ปOnline demo: https://t.co/yHcQ8niYmd
๐ Chat in Google Colab: https://t.co/C4ULZzSSnd
๐LoRA, fully FT and S2 FT models added! https://t.co/o4AVEn0AYF
@skalskip92@KhanSalmanH@abacaj@FahadShahbazkh3@hanoonaRasheed@mbzuai Hi, currently the demo is not available however all the codes and models are available. You can run the demo offline in case if you have resources. We are working on arranging resources to host an online demo.
@baptistejamin@abacaj When it comes to fine-tuning on a downstream task, LLaMA-3 8B I see more stable than Phi-3. However low rank adaptation, I notice Phi-3 is still the winner in most of the cases.
๐ Hey everyone! We are excited to announce our new project: โLLaVA++: Enhancing Visual Capabilities with #LLaMA-3 and #Phi-3, now available on GitHub.
GitHub: https://t.co/blJCdj3nnG
Please do not forget to STAR (โญ) our GitHub Repo (LLaVA-pp). Thanks a ton! :)