How well can AI agents find bugs in virtual 3D worlds? 🔍
Introducing WorldAuditBench, a benchmark for auditing interactive 3D worlds with static and interactive anomalies, such as invisible walls, floating chairs, and objects inconsistent with the surrounding scene.
Across 213 tasks in 13 environments, the strongest agent we evaluated achieves just 42.3% success, compared with 83.4% for humans.
Finding these anomalies requires agents to explore, interact, and collect visual evidence to show what’s wrong.
Here’s what we test and what we found 🧵
How well can AI agents find bugs in virtual 3D worlds? 🔍
Introducing WorldAuditBench, a benchmark for auditing interactive 3D worlds with static and interactive anomalies, such as invisible walls, floating chairs, and objects inconsistent with the surrounding scene.
Across 213 tasks in 13 environments, the strongest agent we evaluated achieves just 42.3% success, compared with 83.4% for humans.
Finding these anomalies requires agents to explore, interact, and collect visual evidence to show what’s wrong.
Here’s what we test and what we found 🧵
(5/N)
Coding agents can now create fantastic 3D worlds and games, but these worlds are still imperfect. We hope WorldAuditBench helps build better auditors to find anomalies and improve the worlds these agents create.
Explore WorldAuditBench 👇
Paper: https://t.co/fnDvgxdCCi
Project: https://t.co/opVcAwfOhd
1/ We keep adding tools, skills, and specialist agents to powerful agentic systems whenever they seem useful. Each addition feels like an upgrade. But what if an evolving harness quietly makes the agent worse at tasks it could already solve?
Our results show that it can. More surprisingly, even today’s self-evolving methods cannot reliably adapt to new capabilities while retaining earlier competence.
Introducing 🎉EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?
Congratulations to the team @Weixin_WeChat ! It’s exciting to see researchers beginning to evaluate their models on our MMEB-V3—and to see further advances in multimodal embedding technology being translated into real-world applications.
MMEB-V3 builds on our MMEB/VLM2Vec series, expanding universal embedding evaluation across a broader range of modalities and scenarios. The paper has been accepted and will be presented at COLM 2026.
Some of the AI in Weixin works quietly behind the scenes.
Meet WeMM-Embedding, a multimodal embedding model developed by our Vision team.
It helps power search and recommendations across text, images and videos — already deployed across Channels, Official Accounts, Moments and e-commerce.
Its 9B model ranked #1 on both MMEB-v2 and MMEB-v3.
Now it’s open source ↓
GitHub: https://t.co/FztiyBFtkV
Hugging Face: https://t.co/KnVhAoIZAG
Great to see Qwen3-VL-Embedding achieving state-of-the-art results on our MMEB series!
Meanwhile, our team has several major feature releases planned for the MMEB/VLM2Vec series in the coming months. Stay tuned!
🚀 Introducing Qwen3-VL-Embedding and Qwen3-VL-Reranker – advancing the state of the art in multimodal retrieval and cross-modal understanding!
✨ Highlights:
✅ Built upon the robust Qwen3-VL foundation model
✅ Processes text, images, screenshots, videos, and mixed modality inputs
✅ Supports 30+ languages
✅ Achieves state-of-the-art performance on multimodal retrieval benchmarks
✅ Open source and available on Hugging Face, GitHub, and ModelScope
✅ API deployment on Alibaba Cloud coming soon!
🎯 Two-stage retrieval architecture:
📊 Embedding Model – generates semantically rich vector representations in a unified embedding space
🎯 Reranker Model – computes fine-grained relevance scores for enhanced retrieval accuracy
🔍 Key application scenarios:
Image-text retrieval, video search, multimodal RAG, visual question answering, multimodal content clustering, multilingual visual search, and more!
🌟 Developer-friendly capabilities:
• Configurable embedding dimensions
• Task-specific instruction customization
• Embedding quantization support for efficient and cost-effective downstream deployment
Hugging Face:
https://t.co/QBTP0XEVmk
https://t.co/c0DB96xxP5
ModelScope:
https://t.co/MhPATTPjYJ
https://t.co/JyNIYQqLAm
Github: https://t.co/qG5Khxk7o9
Blog: https://t.co/K52IC34oNV
Tech Report:https://t.co/dtqAOkMprp
🚨 Introducing VLM2Vec-V2 & MMEB-V2 🚨
At @Salesforce, we're advancing multimodal embeddings beyond natural images to unify videos, visual documents, and images in a single 2B parameter model.
📄 Paper: https://t.co/Sck0sBTs9v
💻 Code: https://t.co/bgJxxHWkn9
🤗 Model: https://t.co/vLdbrSFYZN
🔍 Key innovations:
➡️ MMEB-V2: Comprehensive benchmark with 78 datasets spanning video retrieval, moment retrieval, video classification, video QA, and visual document retrieval
➡️ Unified training framework with instruction-guided contrastive learning that supports temporal understanding and structured document reasoning
📊 Results highlight:
➡️ Outperforms baselines across all modalities (58.0 overall score)
➡️ Strong cross-modal performance from videos to documents while maintaining excellent image task capabilities
🌎 Real-world applications: AI agents, multi-modal search engines, RAG systems, and cross-modal information retrieval
#MultimodalAI #FutureOfAI #EnterpriseAI
#NewPaperAlert
Since we released VLM2Vec, we have received surprising amount of attention from the community. The most common request is to expand VLM2Vec to more modalities like docs, videos, screenshots.
Today, we are so excited to introduce VLM2Vec & MMEB-v2! 🚀 We're advancing multimodal embeddings to unify videos, images, and visual documents with a single powerful 2B model! Check out our work!
📷 Paper: https://t.co/CQD3ne6K6i
📷 Code: https://t.co/uTNXE9SqHH
📷 Model: https://t.co/nZSaL9KeWT
🎉 We’re excited to announce the release of the VLM2Vec-V2 tech report!
It’s inspiring to see the VLM2Vec/MMEB community continuing to grow.
We warmly welcome all feedback, suggestions, and contributions from the community. Let's build this together!
Introducing VLM2Vec & MMEB-v2! 🚀
We're advancing multimodal embeddings to unify videos, images, and visual documents with a single powerful 2B model!
Check out our work!
📜Paper: https://t.co/jPCo4DV0NQ
💻Code: https://t.co/ml0XaxGHOG
🤖Model: https://t.co/Mj6QmvIAe6
🚀 New benchmark alert!
StructEval evaluates how well LLMs generate structured outputs - from JSON APIs to React components to scientific diagrams.
We tested 12 models and found some interesting results… 🧵
Why structured outputs matter (and why they're hard):
- Real apps break when JSON is malformed or HTML doesn't render
- Current benchmarks focus on "does it make sense?" not "does it actually work?"
- No unified way to evaluate both code syntax AND visual correctness
What we built:
- 2,035 examples across 44 task types including both generation (text→structural format) and conversion (structural format 1→ structural format 2) tasks.
- Automated evaluation: Syntax score + Keyword matching + VQA Scores (e.g. ask VLM as a judge about whether a element in the rendered image is there/placed correctly), combined in a weighted average way.
The results? Even SOTA models struggle:
- o1-mini (best overall): 75.58%
- GPT-4o: 76.02%
- Best <10B open-source (Qwen3-4B): 67.04%
- Some challenging sub-tasks like Text→Mermaid: <50% across all models and >30% gaps between the best and the worst 😬
Website: https://t.co/6nXKoramEv
Paper: https://t.co/Ns08oujLfN
Code: https://t.co/U1I8kzfiJX
Details👇(0/5)
Just realized that there are already 30 models on our MMEB (multimodal embedding benchmark) leaderboard. See https://t.co/LyXiVowVXX.
The best models are already 10 points above VLM2Vec now. Excited to see more work from the community to build more and more powerful embedding models that can encode all the modalities seamlessly!
🌟 Hey Singapore! Our #ICLR25 week kicks off tomorrow! 🇸🇬
Up first we present: 🔥 "VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks"
https://t.co/cZ5OfxbTvS
Poster Session | Thursday, April 24 | 3:00pm - 5:30pm
Our approach:
1️⃣ Uses VLM backbone to deeply integrate image & text
2️⃣ Employs contrastive learning across modalities
3️⃣ Enables fixed-dimensional vectors from any text/image combo
🧑🔬 In Singapore? Visit our poster session to learn how we achieved 10-20% improvement over existing models!
🌏 Can't make it? Follow @SFResearch for daily updates & insights throughout the conference! #AIResearch #MachineLearning
👁️ Looking for VLMs that go beyond generators to transform multimodal embeddings? Meet "VLM2VEC: Training Vision-Language Models for Massive Multimodal Embedding Tasks"
📎 Paper: https://t.co/QfIN9r9y53
💻 Website: https://t.co/k2WDCgMUdJ
Our #ICLR25-featured paper shows how vision language models transform into powerful embedders for classification, VQA, retrieval, and visual grounding. We unlock strong emergent capabilities by deeply fusing vision and language rather than shallow combinations.
🇸🇬 Visit us in Singapore to see how we're redefining multimodal representation learning! #MultimodalAI #VLMs
📣 From efficient key caches and multimodal embeddings to self-improving reasoning and faithful context adherence... we're thrilled to present a broad range of powerful new research at #ICLR2025! 🎉
Bookmark our accepted papers below, and we'll see you in Singapore, @iclr_conf !
🔖 REGENESIS: LLMs can grow into reasoning generalists via self improvement
👉https://t.co/4ygBcgs2Kv
🧠Becky Xiangyu Peng Congying Xia Xinyi Yang Caiming Xiong Jason Wu Chen Xing
🔖SiReRAG: Indexing Similar and Related Information for Multihop Reasoning
👉https://t.co/XsZaruRQAM
🧠 Nan Zhang, Prafulla Choubey, Alexander. Fabbri, Gabriel Bernadett-Shapiro, Jason Wu
🔖FaithEval: Can Your Language Model Stay Faithful to Context, Even If “The Moon is Made of Marshmallows''
👉https://t.co/ixb2zTWoZe
🧠 Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke, Xuan Phi Nguyen, Caiming Xiong, Shafiq Joty
🔖Preference Optimization for Reasoning with Pseudo Feedback
👉https://t.co/EHSjB7yl11
🧠Fangkai Jiao, Geyang Guo, Xingxing Zhang, Nancy F. Chen, Shafiq Joty, Furu Wei
🔖ThinK: Thinner Key Cache by Query-Driven Pruning
👉https://t.co/ddxjZ1hVXk
🧠Yuhui Xu; Zhanming Jie; Hanze Dong; Lei Wang; Xudong Lu; Aojun Zhou; Amrita Saha; Caiming Xiong; Doyen Sahoo
🔖Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
👉https://t.co/pfQxnKv5Bc
🧠Zirui Zhao, Hanze Dong, Amrita Saha, Caiming Xiong, Doyen Sahoo
🔖VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
👉https://t.co/cZ5OfxbTvS
🧠Ziyan Jiang, Rui Meng, Xinyi Yang, Semih Yavuz, Yingbo Zhou, Wenhu Chen
🔖Integrating Expertise of Software Engineering Agents
👉https://t.co/ldRXNV3qkj
🧠Kexun Zhang, Weiran Yao, Zuxin Liu, Yihao Feng, Zhiwei Liu, Rithesh Murthy, Tian Lan, Lei Li, Renze Lou, Jiacheng Xu, Bo Pang, Yingbo Zhou, Shelby Heinecke, Silvio Savarese, Huan Wang, Caiming Xiong
Congrats to our researchers for the incredible body of work! #MachineLearning #AIResearch