🤩Next-Generation Video Understanding Benchmark🤩
📽️Video-MME-v2📽️ provides a robust and faithful evaluation of video models with two highlights:
* Progressive Multi-Level Evaluation Dimensions
* Grouped Non-Linear Evaluation Mechanism
- Leaderboard: https://t.co/pc1f7J55Pf
@khalide_f1@_akhaliq Exactly — that’s also what we observed.
Current reasoning models rely heavily on textual cues.
With subtitles, “thinking” tends to help.
But without text, adding “thinking” can even hurt performance.
It suggests that visual reasoning is still far from mature.
🔥🔥 Sharing our work: Video-MME-v2!
A team of 60+ amazing colleagues spent nearly a year building Video-MME-v2.
🤔 Due the existing saturation problem!
🚀 3,300+ human-hours
👉 Human: 90.7 vs the best Gemini-3-Pro: 49.4
❗A substantial gap!
Project: https://t.co/Pm5lkeDHXB
📣🔥 Video-MME-v2 is here!
🎯 Tackling the saturation of video understanding benchmarks
🚀 Built with 3,300+ human-hours over nearly a year
🔍 Progressive tri-level hierarchy & group-based nonlinear scoring
👉 Human: 90.7 vs the best Gemini-3-Pro: 49.4
Project: https://t.co/7Ev4QiR5Fb
Paper: https://t.co/oldRq23ToS
🔥 Excited to share Video-MME-v2! 🔥
We built it to tackle a growing issue: video understanding benchmarks are getting saturated.
🏃🏻 Over 3,300 human-hours, nearly a year of effort
🌟 A new design with a progressive hierarchy + group-based nonlinear evaluation
What we found:
👉 Human: 90.7 vs 👉 Gemini-3-Pro: 49.4
The gap is still huge.
Explore More at:
Page: https://t.co/U9cykLC9ZX
Paper: https://t.co/FsqBcNGt84
Video-MME-v2
A new benchmark for video understanding featuring a progressive tri-level hierarchy and grouped non-linear scoring. Built with 3,300 human-hours across 800 videos to expose gaps between leaderboard scores and true model capabilities.
📢ICLR2026 Acceptance Prediction is out!
🚀Find the acceptance of your paper in advance (predicted):
https://t.co/WweNFNJf93
🛠️Code of Multi-Agent Framework and Benchmark is available:
https://t.co/SgYY8CydTQ
🎯Our goal is to understand the how and why behind paper decisions.
VITA: Towards Open-Source Interactive Omni Multimodal LLM
abs: https://t.co/0f0RiNZPIb
project page: https://t.co/zplb8zxS5i
Model and code coming soon
Omnimodal (video, image, text, audio), multilingual model with GPT-4o-like experience
VITA
Towards Open-Source Interactive Omni Multimodal LLM
discuss: https://t.co/hQp2guRjZa
The remarkable multimodal capabilities and interactive experience of GPT-4o underscore their necessity in practical applications, yet open-source models rarely excel in both areas. In this paper, we introduce VITA, the first-ever open-source Multimodal Large Language Model (MLLM) adept at simultaneous processing and analysis of Video, Image, Text, and Audio modalities, and meanwhile has an advanced multimodal interactive experience. Starting from Mixtral 8x7B as a language foundation, we expand its Chinese vocabulary followed by bilingual instruction tuning. We further endow the language model with visual and audio capabilities through two-stage multi-task learning of multimodal alignment and instruction tuning. VITA demonstrates robust foundational capabilities of multilingual, vision, and audio understanding, as evidenced by its strong performance across a range of both unimodal and multimodal benchmarks. Beyond foundational capabilities, we have made considerable progress in enhancing the natural multimodal human-computer interaction experience. To the best of our knowledge, we are the first to exploit non-awakening interaction and audio interrupt in MLLM. VITA is the first step for the open-source community to explore the seamless integration of multimodal understanding and interaction. While there is still lots of work to be done on VITA to get close to close-source counterparts, we hope that its role as a pioneer can serve as a cornerstone for subsequent research.
Video-MME
The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Gemini 1.5 Pro far outperforms other models, including GPT4o
proj: https://t.co/MAbHO7dyi6
abs: https://t.co/xTR7iuYLWr