So grateful for the chance to learn from @lintool and his great students, always thankful to both @lintool and @amirhkarimi_ for their support, guidance, and great advice along the way ππ
Our paper got an oral at the CTB workshop @ ICMLπ°π·
How stable are LLM leaderboards? π€
Less than you'd hope. We built a unified framework for auditing how small changes to pairwise votes ripple through leaderboard rankings
π https://t.co/tICCrQHYLB
π https://t.co/7Rmyb2WaRG
[1/3] Our agentic search work is at #ACL2026! π I can't make it to San Diego in person, but luckily @zijian42chen is presenting the poster β go say hi!
The one-line takeaway: πΉπππππ ππππππ πππ ππππππ. π§΅
BrowseComp-Plus has been accepted to ACL 2026!
Glad to see that within a year, BrowseComp-Plus has supported 100+ research works exploring the agentic search space.
Unfortunately, I am not able to attend ACL in person.
@zijian42chen will present the paper at ACL 2026, Sunday, July 5th, 14:00β15:30, Harbor G Oral Session. Do chat with him!
This is a great project, backed by tremendous effort from the collaborators. Congrats to the team!
Our paper got an oral at the CTB workshop @ ICMLπ°π·
How stable are LLM leaderboards? π€
Less than you'd hope. We built a unified framework for auditing how small changes to pairwise votes ripple through leaderboard rankings
π https://t.co/tICCrQHYLB
π https://t.co/7Rmyb2WaRG
Two more findings π
β In constructive mode, influence-guided Add selects matchups that shrink a target modelβs CI faster than the current active sampling method.
π£ Deprecating a model is not a free action: on Arena 55K, in the worst case, removal changed 5 Top-10 spots.
π€¨ Is your agent confused about what to build because it says there arenβt any guidelines?
Now your agent has no more excuses - track guidelines for TREC RAG 2026 are out π₯
And yes, theyβre available via SKILLz π
Tell your agents to showcase your agentic search system!
π Introducing AgentIR, a retriever that reads your agentβs mind (literally!)
π§ Unlike humans, agents explicitly expose thoughts in reasoning tokens. Put them to use!
π Simple, substantial gains for agents on BrowseComp-Plus, 35% (BM25) β‘οΈ 50% (Qwen3-Embed) β‘οΈ 67% (AgentIR)
π§΅
Retrievers, rerankers, and reward models are all language models that score text, and they're all vulnerable to the same adversarial attacks.
We propose unifying the study of how to make these models more adversarially robust. π§΅
https://t.co/rPLUSFLSax