Introducing MultiUI, 7.3M general multimodal instructions synthesized from webUIs using text-based LLMs. Models trained on MultiUI not only excel in UI tasks but also generalize surprisingly well to Doc/OCR/chart understanding. https://t.co/0S3FDEcQbl
Working on multimodal instruction tuning and finding it hard to scale? Building Web/GUI agents but data is too narrow?
Introducing 🚀MultiUI: 7.3M multimodal instructions from 1M webpage UIs, offering diverse data to boost text-rich visual understanding.
Key takeaways:
🌟WebUI-trained models show major gains in visual web understanding and agent tasks. 💻
🌟Models also generalize well to non-UI tasks like DocVQA/OCR. 📄
How it works:
We generate multimodal instructions with a text LLM using structured text from webpage accessibility trees. We then pair them with UI screenshots, to train multimodal models.
Homepage: https://t.co/pxEU9iMkqc
Paper: https://t.co/2yLqAlhNxA
Dataset: https://t.co/e2Kir53W8K
Model: https://t.co/2WBgGnU2Xc
Congrats to the student lead @jeepliu1212 and the team @tianyue_01@99Solaris@QuYuxiao@XiongChenyan@WenhuChen@gneubig !
More details are in the following threads ⬇️
Introducing MultiUI, 7.3M general multimodal instructions synthesized from webUIs using text-based LLMs. Models trained on MultiUI not only excel in UI tasks but also generalize surprisingly well to Doc/OCR/chart understanding. https://t.co/0S3FDEcQbl
Working on multimodal instruction tuning and finding it hard to scale? Building Web/GUI agents but data is too narrow?
Introducing 🚀MultiUI: 7.3M multimodal instructions from 1M webpage UIs, offering diverse data to boost text-rich visual understanding.
Key takeaways:
🌟WebUI-trained models show major gains in visual web understanding and agent tasks. 💻
🌟Models also generalize well to non-UI tasks like DocVQA/OCR. 📄
How it works:
We generate multimodal instructions with a text LLM using structured text from webpage accessibility trees. We then pair them with UI screenshots, to train multimodal models.
Homepage: https://t.co/pxEU9iMkqc
Paper: https://t.co/2yLqAlhNxA
Dataset: https://t.co/e2Kir53W8K
Model: https://t.co/2WBgGnU2Xc
Congrats to the student lead @jeepliu1212 and the team @tianyue_01@99Solaris@QuYuxiao@XiongChenyan@WenhuChen@gneubig !
More details are in the following threads ⬇️
(4/8)📊Results:
Disparity between Open-source and Proprietary MLLMs: GPT-4V and Claude outperform open-source MLLMs including GUI agent MLLMs by a large margin, highlighting a discernible gap in the capabilities of current open-source MLLMs compared to proprietary ones.
(3/8)📊Results:
Challenging Nature of Web Understanding Tasks: Even the most powerful MLLMs, GPT-4V and Claude Sonnet achieve average scores of 64.6 and 65.8, respectively, leaving ample room for improvement.
Weak Grounding Ability for most MLLMs.
(2/8)🧐Highlights:
Comprehensiveness: Spanning 139 websites with 1.5K samples, encompassing 12 domains and 87 sub-domains.
Multi-granularity: website, element, and action.
Multi-tasks: understanding, OCR, grounding, and reasoning.
High quality: human verification and curation.
(1/8)🚀We introduce VisualWebBench, a multimodal benchmark designed to assess the understanding and grounding capabilities of MLLMs in web scenarios. Encompassing seven tasks, and comprises 1.5K human-curated instances from 139 real websites, covering 87 sub-domains.