Top Tweets for #HallusionBench
🔥Excited to announce #AutoHallusion, our most recent automatic #hallucination approach that automates and scales up the generation of (image, prompt) pairs to induce hallucinations of #VLMs. Now you can build your own #HallusionBench to evaluate and diagnose VLM hallucinations!
Vision-Language Models (VLMs) suffer from hallucinations👀🙈: their reasoning may ignore some objects or create non-existing ones. This is usually caused by their language model bias dominating the facts in the input visual signals, just like the human brain relying on memory🧠 and past experiences🌎 (called "schema" in cognitive science).
But building a benchmark to evaluate and diagnose VLM hallucinations need a lot of human efforts #HallusionBench.
Can we automatically generate (image, prompt) pairs that can induce hallucinations of VLMs? 🤔
The answer is Yes.
We introduce #AutoHallusion🚀, an automatic hallucination pipeline that scales up the benchmark generation process. It probes LLMs to determine image content that contradicts with the LLM prior, and then generate the image by manipulating objects in the image using image-editing models. (1/n)👇
🔥🔥🔥Excited to present #HallusionBench at the @MMFMWorkshop and @CVPR! 🎉 Catch me 👀 at poster #97 on Tuesday, June 18, from 9:40-11:00 AM and at poster session #457 on Thursday, June 20, from 10:30 AM to noon.
Come by and let's chat! 🥳
🔥🔥🔥Thrilled to announce that our #HallusionBench was accepted by #CVPR2024! See you in Seattle!
@CVPR @gammaumd @UMDResearch @UMDscience
#cvpr #VLMs #hallucination #gpt4v #Gemini #LLaVa
Congrats to the comprehensive large multi-modal model evaluation framework by LMMs-Eval team👇! And thanks for including our #hallusionbench @CVPR
Accelerating the Development of Large Multimoal Models with LMMs-Eval
Repo: https://t.co/fqz953JIPj
Blog: https://t.co/Bgc8jPuTof
We are offering a one command evaluation API for fast and thorough evaluation of LMMs over 39 datasets (increasingly).

🎉🎉🎉 Our paper #HallusionBench was accepted by #CVPR2024! We will be updating the results on Gemini Pro and Ultra as soon as API becomes available!
@CVPR @gammaumd @UMDResearch @UMDscience
#cvpr #VLMs #hallucination #gpt4v #Gemini #LLaVa
📢 Sharing a couple of interesting observations and results of #GeminiAI Pro Vision on our #HallusionBench (Big thanks to @GoogleDeepMind for making those API available! 🙏):
1. Language hallucination is a significant issue, often resulting in outputs that include irrelevant information not found in the question prompt or the accompanying image. A few examples are provided.
2. In terms of accuracy on #HallusionBench, Gemini Pro Vision demonstrates inferior performance compared to #gpt4V and #LLaVA, as detailed in the Table 2 included.
3. Our benchmark results indicate that Gemini Pro Vision exhibits the least bias between 'Yes' and 'No' responses compared to all other models tested, including #gpt4V. See the Table 3 included.
4. Accessibility -- All of the results are obtained using Gemini API. We are facing some "internal error" issue while using Google AI Studio on the web page.
I can't wait to see the release of Gemini Ultra and evaluate whether those problems are alleviated or fixed. Keep an eye on our GitHub page (https://t.co/q2FFw3g42Q) for updates on more results, and stay tuned for the arXiv update as those models are released!

🔥🔥🔥Thrilled to announce that our #HallusionBench was accepted by #CVPR2024! See you in Seattle!
@CVPR @gammaumd @UMDResearch @UMDscience
#cvpr #VLMs #hallucination #gpt4v #Gemini #LLaVa
📣 Check out #HALLUSIONBENCH🚀: a cutting-edge image reasoning benchmark on which the SOTA Large Vision-Language Models (#LVLMs) #GPT4-V & #LLaVA-1.5 still fail! We show how to deceive these LVLMs by simple edits of existing images. https://t.co/CMppRmPKr8

📢 Sharing a couple of interesting observations and results of #GeminiAI Pro Vision on our #HallusionBench (Big thanks to @GoogleDeepMind for making those API available! 🙏):
1. Language hallucination is a significant issue, often resulting in outputs that include irrelevant information not found in the question prompt or the accompanying image. A few examples are provided.
2. In terms of accuracy on #HallusionBench, Gemini Pro Vision demonstrates inferior performance compared to #gpt4V and #LLaVA, as detailed in the Table 2 included.
3. Our benchmark results indicate that Gemini Pro Vision exhibits the least bias between 'Yes' and 'No' responses compared to all other models tested, including #gpt4V. See the Table 3 included.
4. Accessibility -- All of the results are obtained using Gemini API. We are facing some "internal error" issue while using Google AI Studio on the web page.
I can't wait to see the release of Gemini Ultra and evaluate whether those problems are alleviated or fixed. Keep an eye on our GitHub page (https://t.co/q2FFw3g42Q) for updates on more results, and stay tuned for the arXiv update as those models are released!

🚀 Exciting insights from our #HallusionBench on #GeminiAI Pro Vision!
🧠 Hallucination challenges, comparison with #gpt4V & #LLaVA, unique bias insights, and more!
Stay tuned for Gemini Ultra updates on our GitHub👉 https://t.co/KfTAT42HUM
#AIResearch #LLMs
@@GoogleDeepMind
📢 Sharing a couple of interesting observations and results of #GeminiAI Pro Vision on our #HallusionBench (Big thanks to @GoogleDeepMind for making those API available! 🙏):
1. Language hallucination is a significant issue, often resulting in outputs that include irrelevant information not found in the question prompt or the accompanying image. A few examples are provided.
2. In terms of accuracy on #HallusionBench, Gemini Pro Vision demonstrates inferior performance compared to #gpt4V and #LLaVA, as detailed in the Table 2 included.
3. Our benchmark results indicate that Gemini Pro Vision exhibits the least bias between 'Yes' and 'No' responses compared to all other models tested, including #gpt4V. See the Table 3 included.
4. Accessibility -- All of the results are obtained using Gemini API. We are facing some "internal error" issue while using Google AI Studio on the web page.
I can't wait to see the release of Gemini Ultra and evaluate whether those problems are alleviated or fixed. Keep an eye on our GitHub page (https://t.co/q2FFw3g42Q) for updates on more results, and stay tuned for the arXiv update as those models are released!

📢 Sharing a couple of interesting observations and results of #GeminiAI Pro Vision on our #HallusionBench (Big thanks to @GoogleDeepMind for making those API available! 🙏):
1. Language hallucination is a significant issue, often resulting in outputs that include irrelevant information not found in the question prompt or the accompanying image. A few examples are provided.
2. In terms of accuracy on #HallusionBench, Gemini Pro Vision demonstrates inferior performance compared to #gpt4V and #LLaVA, as detailed in the Table 2 included.
3. Our benchmark results indicate that Gemini Pro Vision exhibits the least bias between 'Yes' and 'No' responses compared to all other models tested, including #gpt4V. See the Table 3 included.
4. Accessibility -- All of the results are obtained using Gemini API. We are facing some "internal error" issue while using Google AI Studio on the web page.
I can't wait to see the release of Gemini Ultra and evaluate whether those problems are alleviated or fixed. Keep an eye on our GitHub page (https://t.co/q2FFw3g42Q) for updates on more results, and stay tuned for the arXiv update as those models are released!

Check out our #HallusionBench, specifically design to test and evaluate image/video related reasoning ability of multi-modal #LLMs.
#Gemini Pro still 🤔 struggles on our #HallusionBench 📢Check our latest updates at https://t.co/CMppRmPKr8 with more evaluation instances and more recent VLMs (15 models on leaderboard so far).
#Gemini Pro still 🤔 struggles on our #HallusionBench 📢Check our latest updates at https://t.co/CMppRmPKr8 with more evaluation instances and more recent VLMs (15 models on leaderboard so far).
HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models
paper page: https://t.co/T4wpT0XbEg
Large language models (LLMs), after being aligned with vision models and integrated into vision-language models (VLMs), can bring impressive improvement in image reasoning tasks. This was shown by the recently released GPT-4V(ison), LLaVA-1.5, etc. However, the strong language prior in these SOTA LVLMs can be a double-edged sword: they may ignore the image context and solely rely on the (even contradictory) language prior for reasoning. In contrast, the vision modules in VLMs are weaker than LLMs and may result in misleading visual representations, which are then translated to confident mistakes by LLMs. To study these two types of VLM mistakes, i.e., language hallucination and visual illusion, we curated HallusionBench, an image-context reasoning benchmark that is still challenging to even GPT-4V and LLaVA-1.5. We provide a detailed analysis of examples in HallusionBench, which sheds novel insights on the illusion or hallucination of VLMs and how to improve them in the future.

📣 Check out #HALLUSIONBENCH🚀: a cutting-edge image reasoning benchmark on which the SOTA Large Vision-Language Models (#LVLMs) #GPT4-V & #LLaVA-1.5 still fail! We show how to deceive these LVLMs by simple edits of existing images. https://t.co/CMppRmPKr8

Last Seen Hashtags on Sotwe
Most Popular Users

Elon Musk 
@elonmusk
241.5M followers

Barack Obama 
@barackobama
119.1M followers

Cristiano Ronaldo 
@cristiano
113.5M followers

Donald J. Trump 
@realdonaldtrump
111.8M followers

Narendra Modi 
@narendramodi
107.2M followers

Rihanna 
@rihanna
98.5M followers

NASA 
@nasa
92.3M followers

Justin Bieber 
@justinbieber
91.7M followers

KATY PERRY 
@katyperry
89.5M followers

Taylor Swift 
@taylorswift13
83.4M followers

Lady Gaga 
@ladygaga
74.9M followers

Virat Kohli 
@imvkohli
72.5M followers

Kim Kardashian 
@kimkardashian
70.7M followers

YouTube 
@youtube
68.8M followers

Neymar Jr 
@neymarjr
65.5M followers

Bill Gates 
@billgates
64.8M followers

Selena Gomez 
@selenagomez
62.5M followers

The Ellen Show
@theellenshow
62.3M followers

CNN 
@cnn
61.8M followers

X 
@x
60.8M followers








