🏆 Best Resource Paper Award at #ACL2026@aclmeeting
🚀Thrilled to share that our deep research benchmark paper "HSCodeComp: A Realistic and Expert-level Agent Benchmark for Hierarchical Rule Application" reveived a Best Resource Paper Award #ACL2026!
🎯 Hoping this sparks more focus on reliable AI Agents in the real world.
https://t.co/ToFSPPaYb3
Multimodal and Industrial AI team🚀🚀🚀
LLM's answer is only useful if it survives strict standards checks. Partial correctness can mask safety-critical contradictions.
Data: https://t.co/F2x2phop7j…
Code: https://t.co/HdJdjNie6I…
Paper: https://t.co/i2IdQnq15W
1/6) Excited to share our latest work from the Multimodal and Industrial AI team at Alibaba: IndustryBench! 🚀⚙️
In industrial procurement, an LLM's answer is only useful if it survives strict standards checks. Partial correctness can mask safety-critical contradictions.
Check out the full paper for deep dives into capability dimensions and model comparisons! Feedback and PRs are highly welcome. 👇
Data: https://t.co/8ZflFcHw5W
Code: https://t.co/iTRMcJQhDr
Paper: https://t.co/reTXgWdDrf
#Alibaba #Gemini #Qwen #GPT #Claude #Kimi #GLM #Mimimax
🌺GPT-4o’s image generation is stunning — but how well does it handle complex scenarios? 🤔
We introduce 🚀CIGEVAL🚀, a novel method to evaluate models' capabilities in Conditional Image Generation 🖼️➕🖼️🟰🖼️. Find out how top models perform when conditions get truly challenging! 🔥
#ImageGeneration #AutoEvaluation #Multimodal #GPT4O
🚀 Introducing 𝗦𝗲𝗮𝗿𝗰𝗵-𝗥𝟭 – the first 𝗿𝗲𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻 𝗼𝗳 𝗗𝗲𝗲𝗽𝘀𝗲𝗲𝗸-𝗥𝟭 (𝘇𝗲𝗿𝗼) for training reasoning and search-augmented LLM agents with reinforcement learning!
This is a step towards training an 𝗼𝗽𝗲𝗻-𝘀𝗼𝘂𝗿𝗰𝗲 𝗢𝗽𝗲𝗻𝗔𝗜 “𝗗𝗲𝗲𝗽 𝗿𝗲𝘀𝗲𝗮𝗿𝗰𝗵” via RL.
Our 𝟯𝗕 𝗯𝗮𝘀𝗲 𝗟𝗟𝗠𝘀—including not just 𝗤𝘄𝗲𝗻 𝟮.𝟱 but also 𝗟𝗹𝗮𝗺𝗮 𝟯.𝟮—learn to 𝗿𝗲𝗮𝘀𝗼𝗻 and 𝗰𝗮𝗹𝗹 𝘀𝗲𝗮𝗿𝗰𝗵 𝗲𝗻𝗴𝗶𝗻𝗲𝘀 all on their own!
Everything will be 𝗳𝘂𝗹𝗹𝘆 𝗼𝗽𝗲𝗻 𝘀𝗼𝘂𝗿𝗰𝗲. Stay tuned!
Code: https://t.co/oWVQke0t4H
Experimental logs: https://t.co/W1zD0EVDNc
#R1 #deepresearch #deepseek
🚀Exciting to see how recent advancements like OpenAI’s O1/O3 & DeepSeek’s R1 are pushing the boundaries!
Check out our latest survey on Complex Reasoning with LLMs. Analyzed over 300 papers to explore the progress.
Paper: https://t.co/k1HGQTA2kN
Github: https://t.co/VpcNVcEBSg
🔍 Vision language models are getting better - but how do we evaluate them reliably? Introducing AutoConverter: transforming open-ended VQA into challenging multiple-choice questions!
Key findings:
1️⃣ Current open-ended VQA eval methods are flawed: rule-based metrics correlate poorly with true performance (0.09 on VQAv2), while model-based eval has reproducibility issues (updates in GPT-4o versions constantly increase scores by 6% on MMVet).
2️⃣ To address this challenge, we propose AutoConverter, an agentic framework that automatically converts open-ended VQA to multiple-choice questions. It generates distractors matching/exceeding human difficulty, with only 3% of generated questions incorrect.
3️⃣ Using AutoConverter, we built VMCBench: 9,018 multiple-choice questions from 20 datasets testing 33 VLMs in a unified format!
🎯 Our goal: Make VLM evaluation more reliable, efficient & scalable
https://t.co/BSnOW8QmYr
Joint work with a really fantastic team: @hhhhh2033528 (co-lead) @leoliuym@XiaohanWang96@jmhb0@elaine__sui@ChenyuW64562111@AkliluJosiah2@Ale9806_@anjiangw advised by @lschmidt3@yeung_levy!
Sharing the slides of my talk at Princeton yesterday--"A holistic and critical look at language agents":
https://t.co/7ljTsAJwnU
LLM-based language agents are exciting, but it's also undeniably a quite chaotic space: are agents the next big thing, or are they just thin wrappers around LLMs?
I have been giving this talk 10+ times this year (at CMU/Stanford/Apple/Amazon/etc.), hoping to bring some scientific rigor to this emerging topic. I also learned and sharpened my thinking in this process. Finally, I feel comfortable sharing a close-to-final version with everyone. Comments are welcome!
In this 76-page deck, I talk about
1. the definition of language agents (and why that's the best name)
2. the evolution of AI agents
3. the power of language in agents, demonstrated through our latest work on memory (HippoRAG), world models and model-based planning (WebDreamer), grounding (UGround), and tool use (STE)
4. Exciting future directions (planning, synthetic data, multimodal perception, continual learning, and safety)
🧵
[1/5] Super excited to share our paper "Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language Models" which has been accepted to EMNLP2024! #NLProc#EMNLP2024
📄https://t.co/1bzmYAS0st
🚨Know Where You’re Uncertain When Planning with Multimodal Foundation Models: A Formal Framework
🚀𝐀𝐛𝐬: https://t.co/pUoVNKG0f2
A new framework for handling uncertainty in multimodal foundation models, enhancing robot planning reliability! 🤖🚗
💡The Challenge:
Current models struggle with unpredictable environments, as they can’t accurately separate perception and decision uncertainties. This limits their effectiveness in real-world robotics and autonomous driving.
🔍The Approach:
• Uncertainty Disentanglement: Isolates perception uncertainty (visual recognition) and decision uncertainty (planning reliability).
• Targeted Quantification: Uses conformal prediction for perception and Formal-Methods-Driven Prediction (FMDP) for decision-making.
• Active Sensing: Dynamically re-observes high-uncertainty scenes to improve visual input.
• Automated Refinement: Fine-tunes model with high-certainty data, boosting consistency.
🔧Results:
Reduces output variability by up to 40% and enhances task success rates by 5%, showcasing how uncertainty disentanglement can significantly improve model robustness.
#AI #Robotics #MachineLearning #Uncertainty #AutonomousSystems #UTAustin #TAMU #ML #AutonomousDriving
Excited to share our new EMNLP paper! 📄 We uncover that different LLM layers naturally play distinct roles in multitask learning pipelines. 🧠 Layers contribute as both shared and task-specific components—without explicit parameter assignment! - https://t.co/ccL4XwxTO1
Researchers often have to ask for recommendation letters for visa/job applications, etc.
I wrote a script that allows you to find who cites your papers frequently to create a list of potential letter writers: https://t.co/96nDCt0epQ
Hope it's helpful, improvements are welcome!
🚨Vision-Language Models (VLMs) are truly amazing. Ever wonder if their visual and textual "brains" always agree? I am excited to share our latest paper, where we tackle a critical challenge in VLMs, dubbed the 𝐜𝐫𝐨𝐬𝐬-𝐦𝐨𝐝𝐚𝐥𝐢𝐭𝐲 𝐩𝐚𝐫𝐚𝐦𝐞𝐭𝐫𝐢𝐜 𝐤𝐧𝐨𝐰𝐥𝐞𝐝𝐠𝐞 𝐜𝐨𝐧𝐟𝐥𝐢𝐜𝐭𝐬.
💡 The Issue: VLMs like LLaVA and InstructBlip often contain conflicts between their visual and language components, even as they scale. Our research finds that increasing model size alone doesn’t eliminate these conflicts. This indicates a deeper alignment issue in current multimodal architectures.
🔍 How We Detect Conflicts
We developed a detection pipeline that evaluates whether the model’s visual and textual components provide consistent answers for the same question. We apply a contrastive metric to identify and isolate conflicting samples, allowing us to analyze patterns and gauge the severity of these conflicts.
🔧 Our Solution: Dynamic Contrastive Decoding (DCD)
To address this, we developed DCD, which selectively filters less reliable answers based on confidence. This method boosts accuracy by over 2% on standard datasets (ViQuAE, InfoSeek). For models without logits, we also designed prompt-based strategies that enhance reliability, especially in larger models.
💡Key Takeaways:
• Cross-modality conflicts are common and compromise model outputs.
• DCD and prompt-based strategies provide a scalable way to improve VLM reliability.
• Larger models are better at understanding and processing the information gap in our designed prompting-based strategies.
As VLMs become integral to many real-world applications such as autonomy, robotics, and healthcare, ensuring coherence across modalities is critical. Dive deeper into our findings and code here:
🌟𝐏𝐫𝐨𝐣: https://t.co/WSSBZf7Y6y
🚀𝐀𝐛𝐬: https://t.co/ojDIGqabDm
#AI #MachineLearning #VLM #FoundationModels #DeepLearning #MultimodalAI #KnowledgeConflict
I am hiring a research intern, working LLM (Llama 3+) safety. The internship is expected to start in Summer/Spring 2025, based in New York City. Please drop me an email at [email protected] (Subject starts with "[2025 Intern]")
Learn more here: https://t.co/hpeYKwP0nG
Visual Chain-of-Thought with ✏️Sketchpad
Happy to share ✏️Visual Sketchpad accepted to #NeurIPS2024. Sketchpad thinks🤔by creating visual reasoning chains for multimodal LMs, enhancing GPT-4o's reasoning on math and vision tasks
We’ve open-sourced code: https://t.co/izplXLRkDt