👉Key observations:
1️⃣Current models often struggle to balance timely responses with coherent global understanding.
2️⃣They may stay silent for much of the video, produce sparse captions, or fail to decide when a response should be triggered.
🧵5
🥳🙌Excited to introduce 🔥#Omni-DuplexEval🔥, a benchmark for evaluating real-time duplex omni-modal interaction.🚀
Models must not only understand streaming video/audio, but also decide when to respond.
Paper: https://t.co/k5Vq2NgCkf
Code/Data: https://t.co/YDgbMwFkPT
🧵1
Our experiments reveal a substantial gap.
The best-performing model achieves only 39.6% overall, compared with 81.8% for real-time human performance.
On Proactive Reminder, the best model reaches only 20.0%.
🧵4
Mitigating racial bias from LLMs is a lot easier than removing it from humans!
Can’t believe this happened at the best AI conference @NeurIPSConf
We have ethical reviews for authors, but missed it for invited speakers? 😡
I am now in Bangkok🇹🇭 for #ACL2024 🥳
I will present our work:🔍UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs
arXiv: https://t.co/Mc9ssPFp0E
GitHub: https://t.co/uBdV6D77th
📍Poster Session 3
⏰August 12 16:00 - 17:30
See you🥳
🧵6
The key results are as follows:
1️⃣ #GPT4V only achieves 17.23%. #GPT4 gets 29.50% on text-only tasks. #OlympiadBench is more challenging.
2️⃣A huge gap between closed- and open-source models.
3️⃣ The challenge lies more on question-withimages, Physics and none-English text.
🥳🙌Excited to release 🔥#OlympiadBench🔥, an Olympiad-level bilingual multimodal scientific benchmark. The best-performing model, #GPT4V, attains an average score of 17.23%. Such a challenging benchmark🚀
🤖GitHub: https://t.co/qmDbWhNDHT
📎Arxiv: https://t.co/a4eU5TAex3
🧵1
🧵8
Mistake Analysis of #GPT4V
We manually sample and check 97 maths (55 for English and 42 for Chinese) and 67 physics Olympiad-level open-ended problems that #GPT4V fails, and analyze the type of mistakes, the overall results are as follows