Excited to share several of our recent works:
1. MMBench (ECCV'24 Oral@6C, Oct 3, 13:30): A comprehensive mutli-modal evaluation benchmark adopted by hundreds of teams working on LMMs.
https://t.co/nSkbH43t9d
2. Prism (NeurIPS'24): A framework that can disentangle and assess the perception and reasoning abilities of VLMs, and a potential cost-effective solution for vision-language tasks. https://t.co/9TVRFwE185
3. MMBench-Video (NeurIPS'24 Dataset): A Long-Form Multi-Shot Benchmark for Holistic Video Understanding. https://t.co/kBSa2PgoUU
4. VLMEvalKit (MM'24 OpenSource): An open-source evaluation toolkit of LMMs, supporting 100+ different LMMs and ~50 multi-modal benchmarks. https://t.co/I7dnLWdTmk
If you want to learn more about our work or talk about LMM or other topics, I'm always happy to have a chat. Besides, our team also has openings for intern/full-time researchers and engineers. Feel free to DM me if you are interested.
Thrilled to see myself in the #3 spot on HuggingFace’s most influential users for July!
I look forward to doing more impactful works to give back to the community in the future.
New SoTA VLM: InternLM XComposer 2.5 🐐
> Beats GPT-4V, Gemini Pro across myriads of benchmarks.
> 7B params, 96K context window (w/ RoPE ext)
> Trained w/ 24K high quality image-text pairs
> InternLM 7B text backbone
> Supports high resolution (4K) image understanding tasks
> Video understanding and multi-turn, multi-image chat supported too
> Bonus: Capable of generating web pages (w/ prompt) and high quality text-image articles :O
GG InternLM team, first a SoTA 7B LLM followed by a SoTA 7B VLM! ⚡
Demo and model checkpoints below!
InternLM-XComposer-2.5
A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
We present InternLM-XComposer-2.5 (IXC-2.5), a versatile large-vision language model that supports long-contextual input and output. IXC-2.5 excels in various text-image comprehension and composition applications, achieving GPT-4V level capabilities with merely 7B LLM backend. Trained with 24K interleaved image-text contexts, it can seamlessly extend to 96K long contexts via RoPE extrapolation. This long-context capability allows IXC-2.5 to excel in tasks requiring extensive input and output contexts. Compared to its previous 2.0 version, InternLM-XComposer-2.5 features three major upgrades in vision-language comprehension: (1) Ultra-High Resolution Understanding, (2) Fine-Grained Video Understanding, and (3) Multi-Turn Multi-Image Dialogue. In addition to comprehension, IXC-2.5 extends to two compelling applications using extra LoRA parameters for text-image composition: (1) Crafting Webpages and (2) Composing High-Quality Text-Image Articles. IXC-2.5 has been evaluated on 28 benchmarks, outperforming existing open-source state-of-the-art models on 16 benchmarks. It also surpasses or competes closely with GPT-4V and Gemini Pro on 16 key tasks.
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
- Excels in various text-image tasks w/ GPT-4V level capabilities with merely 7B LLM backend
- Opensourced
https://t.co/fto4phT4Cn
Built our @Gradio app and deployed ShareCaptioner-Video on @huggingface Spaces with ZeroGPU. Now, you can try to generate detailed caption for your own video. Have fun!
https://t.co/Xnm8b1ar99
ShareGPT4Video
Improving Video Understanding and Generation with Better Captions
We present the ShareGPT4Video series, aiming to facilitate the video understanding of large video-language models (LVLMs) and the video generation of text-to-video models (T2VMs)
ShareGPT4Video
Improving Video Understanding and Generation with Better Captions
We present the ShareGPT4Video series, aiming to facilitate the video understanding of large video-language models (LVLMs) and the video generation of text-to-video models (T2VMs)
📣📣📣We are excited to announce the release of Open-Sora Plan v1.1.0.
🙌Thanks to ShareGPT4Video's capability to annotate long videos, we can generate higher quality and longer videos.
🔥🔥🔥We continue to open-source all data, code, and models!
https://t.co/C28gHbiPrU