Excited to share our new preprint, ScholarEvolve!
How can agents keep learning over time automatically? We use new research to propose, implement, and evaluate their harnesses.
I’m at COLM this week. Would love to chat about agent evolution and any topics related to LLM agents!
https://t.co/tofei58uqF
#COLM2026
How well can AI agents find bugs in virtual 3D worlds? 🔍
Introducing WorldAuditBench, a benchmark for auditing interactive 3D worlds with static and interactive anomalies, such as invisible walls, floating chairs, and objects inconsistent with the surrounding scene.
Across 213 tasks in 13 environments, the strongest agent we evaluated achieves just 42.3% success, compared with 83.4% for humans.
Finding these anomalies requires agents to explore, interact, and collect visual evidence to show what’s wrong.
Here’s what we test and what we found 🧵
Excited to share our paper at #ICML2025! We've developed MELON🍉, a robust defense method against indirect prompt injection attacks on LLM agents that achieves near 0 ASR! Hope you enjoy🍉!
Grateful to my incredible collaborators! @WilliamWangNLP@WenboGuo4 @jd92wang @xianjun_agi
📜🚨 Check out our latest work on "Self-Resource Allocation in Multi-Agent LLM Systems" where we explore how LLMs can be used to optimize task allocation in multi-agent systems 🤖
🧵(1/3)
1/ Long chain-of-thought (CoT) reasoning boosts LLM performance—but with a computational overhead.
Checkout our new paper, ThinkPrune, where we explore a simple question: To what extent can we cut the reasoning length while keep the quality?
We show that by simply adding a hard length limit during RL training, the length can be reduced by ~40%-60% with only 3% decrease in accuracy on math benchmarks.
📄 Paper: https://t.co/ERQep1onK4
💻 Code: https://t.co/MyrD0W9ZFd
This work wouldn't have been possible without the incredible contributions of our amazing co-authors:
@CodeTerminator@YangZha26484161@jacobandreas@MITIBMLab@nlp_mit@liu_yujian@JiabaoJi@KaizhiQian
4/ This work was only made possible by the remarkable efforts of our amazing co-authors—@hou_bairu, @yujia_bao, @CodeTerminator. Your expertise, dedication, and collaboration made all the difference! 🚀
In RAG applications, LLMs often reprocess the same database chunks for different queries—leading to high latency and cost from handling massive input tokens.
We introduce KVLink, an efficient approach to reuse pre-computed KV caches of retrieved documents, drastically cutting redundant encoding while maintaining performance (see figure below ⬇️).
Highlights:
🥇 Comparable performance to standard decoding across 17 datasets.
⚡ ~90% reduction in time to first token with contexts of ~5K tokens.
🔧 Easily adaptable to current LLM training and inference frameworks by modifying the attention mask and positional embeddings. #AI #LLMs #RAG #NLP #GenAI #MachineLearning
📄 Paper: https://t.co/OmuhI0J5q0
💻 Code: https://t.co/uuL1emEJpn
3/ How does it work?
We address this challenge with four key components:
1️⃣ Position Consistency: Encode documents individually, without storing position embeddings. Restore them only at inference.
2️⃣ Trainable link tokens "stitch" separate KV caches, ensuring smooth cross-attention.
3️⃣ Mixed data fine-tuning.
4⃣All components are fully compatible with current Transformer framework, requiring only minimal changes to the attention mask and positional embeddings.
2/ Problem setting
To avoid repetitive KV cache encoding for overlapping retrieved documents across different queries, we consider:
Precompute KV caches for each document separately
During inference, concatenate precomputed caches for reuse.
🚨But here’s the challenge: LLMs have never seen separately encoded KV caches during training. This discrepancy between training and inference can hurt performance! Directly reusing the precomputed caches leads to an average ~30% accuracy drop on QA tasks.