Our multimodal fact-checking work was selected as the BEST PAPER AWARD HONORABLE MENTION in #SIGIR2023. Huge thanks to the reviewers and committee members for recognizing our contributions. The dataset, code, checkpoints, and paper are at https://t.co/U16DQQAu5W.
@zhuohaol I'm Barry Yao, a final-year CS PhD at UC Davis working on data-efficient post-training of LLM/VLM and long-horizon agents. I'll be at COLM and would love to connect and chat if you're around! I have publications at ACL, EMNLP, EACL, SIGIR (Best Paper Honorable Mention), and AAAI.
(2/2) Our findings suggest several directions for future circuit-aware continual pre-training: adaptively allocating training effort based on learning and forgetting signals, scheduling data to reduce interference and exploit positive transfer between concept circuits.
(1/2) Our paper, “How Do Large Language Models Learn Concepts During Continual Pre-Training?” (https://t.co/onCjmgx1Va), was accepted to #EMNLP2026 Main!
We find that LLMs internal circuit provide consistent signals of concept learning, forgetting and interference.
@songyoupeng@CVPR I am a final-year PhD student at UC Davis, working on data-efficient adaptation of large multimodal models. Looking forward to having a coffee chat with you in CVPR.
I’ll be at #CVPR2026 in Denver from June 3–7 and would be happy to connect with researchers, practitioners, and friends in the community. I am currently seeking Fall 2026 research internship opportunities and 2027 full-time roles.
I’ll be joining Oracle as an Applied Scientist Intern at the Redwood City office from June to September, where I’ll be mentored by Dr. Avi Sil. If you will also be in the Bay Area this summer, I would be happy to catch up!
(2/2) We introduce a new lens on knowledge learning that goes beyond isolated facts, instead emphasizing the structured relationships among knowledge within and across concepts, how these relationships drive interference and synergy, motivating interference-aware data scheduling.
New preprint "How Do Large Language Models Learn Concepts During Continual Pre-Training?"(https://t.co/hFhwp55kQL): we study how individual concepts are acquired and forgotten, and how multiple concepts interact through interference and synergy during continual pretraining
(1/2) We present the first systematic study of concept acquisition & forgetting in LLMs during continual pretraining. We find links between internal mechanisms and learning dynamics, offering signals to guide training effort allocation.
Excited to share that our paper "Error-driven Data-efficient Large Multimodal Model Tuning" has been accepted to #ACL2025 !🎉
We propose an error-driven tuning framework for efficiently adapting large multimodal models (LMMs) to newly emerging tasks without requiring extensive task-specific training data.
In our approach, a generic LMM, acting as a student model, is first evaluated on a small validation set of the target task, and then a more powerful model, acting as a teacher model, identifies the erroneous steps within the student model's reasoning steps and analyzes its capability gaps from fully addressing the target task. Based on these gaps, targeted training samples are further retrieved from existing task-agnostic datasets to tune the student model and tailor it to the target task.
Experiments across three data scales and seven tasks demonstrate a 7.01% average improvement in the downstream performance of LMMs.
📄 Read the paper: https://t.co/qomp2FHfUj
Big thanks to my co-authors Dr. Qifan Wang and @lifu_huang!
🚨 New paper alert!
We introduce InterleavedBench📚, the first comprehensive evaluation benchmark for interleaved text-and-image generation, as well as InterleavedEval🔍, a powerful GPT-based evaluator that supports multi-aspect assessment.
arXiv: https://t.co/9FC15IMoxo
(1/n)
🚀 Excited to introduce my internship work at @Apple MLR : Many-to-many Image Generation with Auto-regressive Diffusion Models (https://t.co/YKTexvRdLM). Exploring the paradigm for domain-general multi-image to multi-image generation.
Vision-Flan
Scaling Human-Labeled Tasks in Visual Instruction Tuning
Despite vision-language models' (VLMs) remarkable capabilities as versatile visual assistants, two substantial challenges persist within the existing VLM frameworks: (1) lacking task diversity in pretraining and visual instruction tuning, and (2) annotation error and bias in GPT-4 synthesized instruction tuning data. Both challenges lead to issues such as poor generalizability, hallucination, and catastrophic forgetting. To address these challenges, we construct Vision-Flan, the most diverse publicly available visual instruction tuning dataset to date, comprising 187 diverse tasks and 1,664,261 instances sourced from academic datasets, and each task is accompanied by an expert-written instruction. In addition, we propose a two-stage instruction tuning framework, in which VLMs are firstly finetuned on Vision-Flan and further tuned on GPT-4 synthesized data. We find this two-stage tuning framework significantly outperforms the traditional single-stage visual instruction tuning framework and achieves the state-of-the-art performance across a wide range of multi-modal evaluation benchmarks. Finally, we conduct in-depth analyses to understand visual instruction tuning and our findings reveal that: (1) GPT-4 synthesized data does not substantially enhance VLMs' capabilities but rather modulates the model's responses to human-preferred formats; (2) A minimal quantity (e.g., 1,000) of GPT-4 synthesized data can effectively align VLM responses with human-preference; (3) Visual instruction tuning mainly helps large-language models (LLMs) to understand visual features.
Our entity linking work has been accepted by #EACL2024. Check our work: Ameli: Enhancing Multimodal Entity Linking with Fine-Grained Attributes (https://t.co/PJlCXil3l2). Congratulations to all collaborators! The dataset, code, and checkpoints will be released soon.
Our new work ✨The Art of SOCRATIC QUESTIONING: Recursive Thinking with Large Language Models✨ is accepted to #EMNLP2023. Inspired by the human cognitive process, we propose SOCRATIC QUESTIONING, a divide-and-conquer style algorithm that mimics the 🤔recursive thinking process.
Today we officially release ✨Vision-Flan✨, the largest human-annotated visual-instruction tuning dataset with 💥200+💥 diverse tasks.
🚩Our dataset is available on Huggingface https://t.co/XqFrpudysl
🚀 For more details, please refer to our blog https://t.co/HY1D6xrCPm