🎓 We’re thrilled to announce that Xijun Wang @xijunwang_cs, Tianrui Guan @terryguan97, and Pooja Guhan @GuhanPooja have successfully defended their PhD theses! Huge congratulations to all three. Your hard work and dedication have truly paid off. We’re so proud of you! 👏🥳
Attention sinks in LLMs are weird. There’s ~20% of heads that don’t seem to do anything.
Do these heads matter? Turns out that if we get rid of them, benchmark scores don’t change.
.@xijunwang_cs will be presenting “SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition”, which utilizes prompt learning for aerial video action recognition. ✈️
Thurs Oct 17
17:30 - 17:45
Room 6
#ECCV2024 Visit our ViLA poster and Say Hi to Prof.
@MingCLinCS !
Time: October 3, 4:30pm - 6:30pm
Location: Poster session 6, ID 274
Arxiv: https://t.co/LjIpvuBBLy
Code: https://t.co/yohW1Rp0AL
Big shout-out to @shanyangmie!
@gammaumd
Xijun Wang @xijunwang_cs will be presenting "ViLA: Efficient Video-Language Alignment for Video Question Answering", which addresses both efficient frame sampling and effective cross-modal alignment! 🗣️
Time: October 3, 4:30pm - 6:30pm
Location: Poster session 6, ID 274
Join us at #93 of the Poster session on Thursday, February 22! #AAAI24@gammaumd
Say hi to @shanyangmie there!
Our ICAR can recommend items with good visual compatibility including similarity (color, geometry, texture, etc.) and complementarity (like table vs chair).
Join us at Naupaka #74, Saturday 17:15-19:15! #WACV2024
We proposed PMI Sampler for aerial footage, it efficiently selects key frames using patch-wise mutual information. Ideal for moving camera footage with dynamic backgrounds.
@xijunwang_cs @DivyaKRaman1 @dmanocha@gammaumd
See you Saturday 17:15-19:17 at Naupaka #51 #WACV2024
For aerial videos, we proposed MITFAS to focus on the regions corresponding to salient motions and find the more informative frames by using mutual information.
@RuiqiXian@dmanocha@gammaumd
📢 Sharing a couple of interesting observations and results of #GeminiAI Pro Vision on our #HallusionBench (Big thanks to @GoogleDeepMind for making those API available! 🙏):
1. Language hallucination is a significant issue, often resulting in outputs that include irrelevant information not found in the question prompt or the accompanying image. A few examples are provided.
2. In terms of accuracy on #HallusionBench, Gemini Pro Vision demonstrates inferior performance compared to #gpt4V and #LLaVA, as detailed in the Table 2 included.
3. Our benchmark results indicate that Gemini Pro Vision exhibits the least bias between 'Yes' and 'No' responses compared to all other models tested, including #gpt4V. See the Table 3 included.
4. Accessibility -- All of the results are obtained using Gemini API. We are facing some "internal error" issue while using Google AI Studio on the web page.
I can't wait to see the release of Gemini Ultra and evaluate whether those problems are alleviated or fixed. Keep an eye on our GitHub page (https://t.co/q2FFw3g42Q) for updates on more results, and stay tuned for the arXiv update as those models are released!
🔥CoBig shoutout to @Google for launching their groundbreaking large multimodal model #Gemini.
🚩However, there are still obvious hallucinations with Gemini. Here are a few examples with Gemini Pro.
Want to learn more? Look at our HallusionBench Paper: https://t.co/3qjsR3Drh7
How to strengthen both CNNs and Transformers?
Check our work “SCSC: Spatial Cross-scale Convolution Module to Strengthen both CNNs and Transformers” on Oct. 2nd (1:30pm-6pm) at #ICCV2023 Workshop on New Ideas in Vision Transformers
@room P01
ArXiv: https://t.co/Ozc5PGk7zB
How to localize a robot(equipped with a LiDAR) in a Large-scale point cloud map?
We present "CrossLoc3D - Aerial-Ground Cross-Source 3D Place Recognition"
Code: https://t.co/nuDdbW9UWj
PDF: https://t.co/jYZQzRTg6B
#ICCV2023
🧵