Super excited to TA this new AI Agents course at CMU this fall, and to give a guest lecture on agent training!
Looking forward to exploring how to build, evaluate, and train agentic LLMs with students 🔥
🚀Excited to release PACE: A Proxy for Agentic Capability Evaluation!
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA is expensive, slow, and infrastructure-heavy, often costing $$$ and taking hours or days per model.
❓But do we always need to run full agentic evaluations?
In PACE, we show that agentic benchmark performance can be accurately predicted from a small, carefully selected set of cheap non-agentic benchmark instances.
PACE automatically selects proxy instances from existing benchmarks covering skills like instruction following, planning, tool use, reasoning, coding, retrieval, and multimodal understanding.
Across 14 models, 4 agentic benchmarks, and 19 non-agentic benchmarks, PACE-BENCH achieves:
✅ 3.80% MAE for absolute score prediction
✅ 0.81 Spearman correlation for model ranking
✅ ~84% pairwise preference accuracy
✅ ~100× lower cost than target benchmark sampling
Beyond prediction, PACE also reveals what capabilities different agentic benchmarks actually require, e.g., planning, verification, long-context aggregation, and instruction following.
We hope PACE makes agentic evaluation cheaper, faster, and more accessible for model development, model selection, and routing :)
📃 Paper: https://t.co/I6iYUCADov
💻 Code: https://t.co/tLd7HrSgEA
I'm incredibly grateful to have worked with @lintangsutawika, @Jiarui_Liu_, @lltjuatja, @JiayiiGeng, @lrzneedresearch, @daniel_js_lee, @Aditya_Soni_8, @Vincent92965015, @xiangyue96, and @gneubig .
Updates:
Excited to share that Agent Data Protocol (ADP) is accepted to ICLR 2026 Oral! 🎉
We also added support for 3 new datasets: SWE-Play, MiniCoder, and Toucan, bringing us to 3M trajectories supported.
If you're training agentic LMs, try ADP + tell us what dataset/agent format you want next. PRs & requests welcome. Let's make this the open standard for agent training data 🔥
🚀Original post: https://t.co/ikWWie9Qk6
📄Read our paper: https://t.co/OlCTvhrXQ7
🌐Check our project website: https://t.co/A2cwXkvIam
New result on our VisualPuzzles benchmark 🧩
📃Gemini 3 Pro scores 52.7%, slightly below o3 (54.0) and o4-mini (57.0), and still under the lower 5th-percentile human performance (57.5%).
Still a long way to go on multimodal reasoning! 🧠
Original Post: https://t.co/2LUzksSGiS
Join us this Saturday at 10 PM EST for a talk by Yueqi Song @yueqi_song on Agent Data Protocol! 🤖
🎟️ Register here: https://t.co/0xBa9lzmXv
Learn about a comprehensive framework that enables: ✨ Reproducible research in agent behaviors
✨ Seamless dataset integration
✨ 1.3M training trajectories in the public ADP Dataset V1
✨ An average of 20% improvement over base models
📄 Paper: https://t.co/kbQuRV6PLC
We use LLMs for everyday tasks—research, writing, coding, decision-making. They remember our conversations, adapt to our needs and preferences. Naturally, we trust them more with repeated use.
But this growing trust might be masking a hidden risk: what if their beliefs are shifting and we don't notice?
We study the question "Do LM assistants change their beliefs as context accumulates?" in our new preprint: 👇
(1/n)
We just built and released the largest dataset for supervised fine-tuning of agentic LMs, 1.27M trajectories (~36B tokens)!
Up until now, large-scale SFT for agents is rare - not for lack of data, but because of fragmentation across heterogeneous formats, tools, and interfaces.
To solve this, we introduce the Agent Data Protocol, a new “interlingua” between a broad variety of heterogeneous agent datasets - coding, browsing, API/tool use - and unified agent training pipelines downstream.
We unified 13 datasets into ADP, converted them to be compatible with multiple agent frameworks, and observed ~20% average gains, reaching SOTA/near-SOTA without domain-specific tuning.
📄 Read our paper: https://t.co/OlCTvhrXQ7
🌐 Check our project website: https://t.co/wBggu0hQ2i
And this is just getting started, we can add more datasets, further expand the resources, and make training agent LMs easy for all. We’d love to have you join the shared effort and help to make ADP the open standard for the community 🚀
Can multimodal models really reason?
There are other benchmarks that purport to test multimodal reasoning, but they also require lots of domain specific knowledge.
We created a benchmark that is knowledge-light and reasoning-heavy, and found SOTA MLLMs lag far behind humans.
Humans can perform complex reasoning without relying on specific domain knowledge, but can multimodal models truly do that as well?
Short answer: No. Even the best models perform below the 5th-percentile human on our VisualPuzzles tasks.
🚀 Introducing VisualPuzzles🧩: a new benchmark designed to disentangle multimodal reasoning from domain knowledge.
Why VisualPuzzles?
VisualPuzzles challenges models with logic-based puzzles that require minimal prior knowledge. A key source: manually translated visual questions from the Chinese Civil Service Exam (中国国家公务员考试行测) 🔑
🤔Key Findings
- The best model models score under 57.5% (5th percentile human baseline)
- Knowledge ≠ Reasoning: On knowledge-heavy benchmarks like MMMU, reasoning strongly correlates with knowledge — but not on VisualPuzzles
- Larger models = better knowledge, but not necessarily better reasoning
- 🧠"Thinking" modes don't always help. More tokens = better knowledge recall, but more tokens ≠ better reasoning 🤷♀️
👉 Explore the benchmark:
🌐 Project Website + Leaderboard: https://t.co/eLhNr7B4MI
📄 arXiv: https://t.co/52foc8vpn9
🙌 With
@yueqi_song@tianyue_01
@Yibo_Kong
@Zecheng_Li@xiangyue96@gneubig
& supported by CMU NeuLab 💛
🙏 Special thanks to the LTI community at CMU for the insightful feedback, generous support, and always inspiring conversations.
🚨Can subtle changes in prompts promote target concepts (e.g., brands) without user suspicion?
We have a paper (https://t.co/N861d6psuo) accepted by #CHI2025.
🚀We reveal critical risks in LLM recommendations and motivate users to rethink reliance on external prompt providers.
The talk for our @NDSSSymposium
2024 paper "Group-based Robustness: A General Framework for Customized Robustness in the Real World" can be found at https://t.co/06nb9CbDmU. Thanks to the collaborators
@keanelucas
@lujobauer@mahmoods01@mk_reiter
Our paper "Group-based Robustness: A General Framework for Customized Robustness in the Real World" was presented at @NDSSSymposium 2024. Thanks to the collaborators @keanelucas @lujobauer@mahmoods01@mk_reiter
paper: https://t.co/C1G0i0YaoG
code: https://t.co/RO4fOd7hf7
I am glad to share our paper at #ICML2022 . It is the first time that I have attended a conference in person. I am really glad to make many new friends!
Paper: https://t.co/I1RrrsBPIM
Video: https://t.co/4DcHymV3w2