1/7 We're launching Tongyi DeepResearch, the first fully open-source Web Agent to achieve performance on par with OpenAI's Deep Research with only 30B (Activated 3B) parameters! Tongyi DeepResearch agent demonstrates state-of-the-art results, scoring 32.9 on Humanity's Last Exam, 45.3 on BrowseComp, and 75.0 on the xbench-DeepSearch benchmark.
7/7🚀 Try it out!
Tongyi DeepResearch agent marks a significant step towards AI that can autonomously turn information into insight.
We're open-sourcing the model, framework, and complete solutions to empower the community. Dive in and build with us!
🔗 Homepage: https://t.co/jvs2lLCihB
🔗 Blog: https://t.co/fJv2sf8LCj
🔗 Model HuggingFace: https://t.co/97Smx2EB0i
🔗 Model ModelScope: https://t.co/jCwb0OaLcq
🔗 GitHub Repo: https://t.co/3pn2ZkzIYD
📝 New from FAIR: An Introduction to Vision-Language Modeling.
Vision-language models (VLMs) are an area of research that holds a lot of potential to change our interactions with technology, however there are many challenges in building these types of models. Together with a set of collaborators across academia, we’re releasing ‘An Introduction to Vision-Language Modeling’ — we hope that this new resource will help anyone who would like to enter this field to better understand the mechanics behind mapping vision to language.
Full paper ➡️ https://t.co/lid1qBuT0N
This guide covers how VLMs work, how to train them and approaches to evaluation — and while it primarily covers mapping image to language, it also discusses how to extend VLMs to videos.
We hope that releasing this guide will inspire and enable more work in this space.