Hey everyone! I am super excited to share that our new research report is live on ArXiv! ๐
IRPAPERS: A Visual Document Benchmark for Scientific Retrieval and Question Answering!
Thread with more details! ๐งต(1/11)
What happens if you skip first-stage retrieval and let a state-of-the-art cross-encoder score all 10,000 documents itself? ๐ค
In Drowning in Documents, BM25 beat it.
In this chapter, @mat_jacob1002 walks through the full-scoring experiments: golds plus 10,000 randomly sampled documents, every one scored by the reranker, and old-school keyword search still came out ahead.
Watch Mathew break down the experiment ๐
https://t.co/yhvQMooe0m
Check out this podcast with @mat_jacob1002 if you are interested in reranking / retrieval related research!
Mat also shared a bit about our work on pushing the pareto frontier of retrieval agents ๐
Scaling Test-Time Compute has been one of the clearest wins for improving AI systems. I am SUPER EXCITED to share that weโre bringing it to Search with new effort tiers in the Query Agentโs Search Mode!
Medium, high, and ultrahigh ๐
๐ฌWe benchmarked all three tiers against Hybrid Search across 8 reasoning-intensive and domain-specific retrieval benchmarks: 5 subsets from BRIGHT, IRPAPERS, WixQA, and OBLIQ-Bench Congress.
We find that every effort tier outperforms Hybrid Search, often by a wide margin. Ultrahigh effort roughly quadruples Hybrid Searchโs nDCG@10 on BRIGHT Biology, and delivers a ~7x improvement in Success@1 on OBLIQ-Bench Congress. More details can be found in the image below!
I hope you find this interesting!
YouTube: https://t.co/NZTSrxO0xr
Blog: https://t.co/Zlm45ZvBQS
Docs: https://t.co/NJfpEpCGzw
Retrieve more documents, rerank them, get better results. That's everyone's mental model of search. But push a cross-encoder past around 100 documents and recall@10 starts plummeting instead. ๐ค
In this chapter, @mat_jacob1002 traces how Drowning in Documents started: a drop so counterintuitive it looked like a bug, and the realization that rerankers act less like stronger retrievers and more like boosting, fitting the first stage's errors. ๐
https://t.co/ZJ0wSnSE4C
Retrieve more, rerank, get better results. That's what most of us expect from search, but it turns out to be wrong! ๐คฏ
I'm SUPER EXCITED to publish the 141st episode of the Weaviate Podcast with Mathew Jacob (@mat_jacob1002)! Mathew led the work behind "Drowning in Documents" during his time at Databricks and is now a Ph.D. student at the University of Washington working on ML systems!
This episode dives deep into "Drowning in Documents". This has been one of the most influential papers for us @weaviate_io as we are exploring scaling reranked retrieval. I think it is a must read for those working in Search and Information Retrieval. ๐
We begin with an overview of the paper, and then dive into full scoring with cross encoders and phantom hits. We then cover Listwise Rerankers, what next generation cross encoders might look like, and ranking cascades.
On the topic of ranking cascades, I loved learning about Mathewโs work with Melissa Pan (@melissapan), Negar Arabzadeh (@NegarEmpr) and collaborators on โNatural Language Query to Configuration for Retrieval Agentsโ. Per-query โeffortโ prediction is certainly going to be a huge component on the future of these search systems! (Congratulations to OpenRouter ๐)
We then dove into Mathewโs work on TraceLab, a super exciting effort to understand coding agents such as Claude Code and Codex. โจ๏ธ
This was a super fun conversation, and I really hope you find it useful!
YouTube: https://t.co/3lPicYhk5R
Spotify: https://t.co/sx0o06kilI
Retrieve more, rerank, get better results. That's what most of us expect from search, but it turns out to be wrong! ๐คฏ
I'm SUPER EXCITED to publish the 141st episode of the Weaviate Podcast with Mathew Jacob (@mat_jacob1002)! Mathew led the work behind "Drowning in Documents" during his time at Databricks and is now a Ph.D. student at the University of Washington working on ML systems!
This episode dives deep into "Drowning in Documents". This has been one of the most influential papers for us @weaviate_io as we are exploring scaling reranked retrieval. I think it is a must read for those working in Search and Information Retrieval. ๐
We begin with an overview of the paper, and then dive into full scoring with cross encoders and phantom hits. We then cover Listwise Rerankers, what next generation cross encoders might look like, and ranking cascades.
On the topic of ranking cascades, I loved learning about Mathewโs work with Melissa Pan (@melissapan), Negar Arabzadeh (@NegarEmpr) and collaborators on โNatural Language Query to Configuration for Retrieval Agentsโ. Per-query โeffortโ prediction is certainly going to be a huge component on the future of these search systems! (Congratulations to OpenRouter ๐)
We then dove into Mathewโs work on TraceLab, a super exciting effort to understand coding agents such as Claude Code and Codex. โจ๏ธ
This was a super fun conversation, and I really hope you find it useful!
YouTube: https://t.co/3lPicYhk5R
Spotify: https://t.co/sx0o06kilI
Finally, have to shoutout @KanZhu854772 for his amazing work with TraceLab. I would stay tuned for some of the work he's going to drop soon, going to be ๐ฅ.
Should also follow @UWSyFi while you're here to keep up with the latest from our group at UW!
@mat_jacob1002 ๐ Thank you so much Mathew! Learned so much from the conversation!
Yes! Super excited to build on this work and scale reranked retrieval! ๐
Had so much fun talking to @CShorten30 about our Drowning in Documents work! Was fun going down memory lane about this work :D! Stay tuned for the research that is going to be coming out of Weaviate๐