We are releasing the most significant technology transfer effort by the CIIR in the last decade. Please share with your network.
In this thread you will learn what Nage enables you to do.
๐งต๐
DSPy 3.4.0 out, with:
(1) native support for Jev and System One models - in the timeless DSPy syntax.
(2) a brand new optimizer, ReAnchor, specifically for calibrating outputs with confidence.
(3) lightning-fast import speeds for the LLM abstraction, via sibling library LM15)
SmallReason-ColBERT: An Ultra-Small Late-Interaction Retriever for Reasoning Intensive Retrieval
@perdactor et al. present a 32M retriever for reasoning-heavy queries, adding a learned token-weighting head.
๐https://t.co/IX9ZrP1UiB
๐จ๐ฝโ๐ปhttps://t.co/roDgX8pkPG
OBLIQ-IR: Training a Dense Retriever for Oblique Queries
@perdactor et al. train a dense retriever for oblique queries, where relevance depends on latent traits like stance or style rather than topic overlap.
๐ https://t.co/LAyyASU6VO
๐จ๐ฝโ๐ป https://t.co/oLyMwaPLa4
EXCISE: Query-Side Exclusion for Late-Interaction Retrieval
Presents a query-time fix for "exclusion inversion" in ColBERT retrievers, where excluded topics get ranked higher, not lower.
๐ https://t.co/lGDJHGqKCv
imo a significant part of the problem is the insistence of the community on misnaming the paradigm as โmulti-vector retrievalโ, we called it late interaction for a reason
the problem isnโt in having one vector; the problem is in the scoring function!
https://t.co/QupzaHoiSd
Completely agree; we already have these in our limitations.
What we do fits a similarity definition. It doesn't let the query reshape the document representation, and as you say, late interaction is finer but still the same fixed space.
We say a version of this in the paper's Limitations: we narrow the first-stage bottleneck rather than close it. That's where our follow-up work is going โ making the latent attribute explicit instead of hoping one fixed vector space encodes it. Promptriever is the right reference too.
We're still working on a new architecture for exactly that.
Genuinely useful question, thanks. Happy to compare notes.
most interesting thing happening rn is that you are allowed to use ai for everything. everything except writing. if you use ai for writing, they will unalive you in public
๐ Our lab has 10 papers accepted at #EMNLP2026 โ 9 Main Conference + 1 Findings!
This is ~1.5 years of work, and almost none of it landed on the first try. Several of these papers went through 4โ6 submission cycles before they were accepted. Rejection was part of the process, not the end of it.
Main Conference
1๏ธโฃ OBLIQ-IR: Training a Dense Retriever for Oblique Queries
Mahmoud Abdalla, ๐๐ฏ๐ฑ๐ฒ๐น๐ฟ๐ฎ๐ต๐บ๐ฎ๐ป ๐๐ฏ๐ฑ๐ฎ๐น๐น๐ฎ๐ต, Shaimaa Sedek, Adam Jatowt
๐ Coming soon ยท ๐ป Coming soon
2๏ธโฃ MEMORA: A Memory-Enhanced Multimodal Committee Reranking Agent for Reasoning-Intensive Retrieval
Mohamed Mahmoud, Mostafa Farouk Senussi, Mahmoud Abdalla, Mahmoud SalahEldin Kasem, ๐๐ฏ๐ฑ๐ฒ๐น๐ฟ๐ฎ๐ต๐บ๐ฎ๐ป ๐๐ฏ๐ฑ๐ฎ๐น๐น๐ฎ๐ต, Hyun Soo Kang
๐ Coming soon ยท ๐ป Coming soon
3๏ธโฃ SustainableQA: A Comprehensive Question Answering Dataset for Corporate Sustainability and EU Taxonomy Reporting
Mohammed Ali, ๐๐ฏ๐ฑ๐ฒ๐น๐ฟ๐ฎ๐ต๐บ๐ฎ๐ป ๐๐ฏ๐ฑ๐ฎ๐น๐น๐ฎ๐ต, Adam Jatowt
๐ https://t.co/HvJQC0tPLZ ยท ๐ป https://t.co/16pUx221nE
4๏ธโฃ TEMPO: Realistic Multi-Domain Benchmark for Temporal Reasoning-Intensive Retrieval
๐๐ฏ๐ฑ๐ฒ๐น๐ฟ๐ฎ๐ต๐บ๐ฎ๐ป ๐๐ฏ๐ฑ๐ฎ๐น๐น๐ฎ๐ต, Mohammed Ali, Muhammad Abdul-Mageed, Adam Jatowt
๐ https://t.co/rOt5FCopkl ยท ๐ https://t.co/kbE3zqsOyn
5๏ธโฃ Argus-Retriever: Vision-LLM Late-Interaction Retrieval with Region-Aware Query-Conditioned MoE for Visual Document Retrieval
๐๐ฏ๐ฑ๐ฒ๐น๐ฟ๐ฎ๐ต๐บ๐ฎ๐ป ๐๐ฏ๐ฑ๐ฎ๐น๐น๐ฎ๐ต, Mahmoud Abdalla, Mohammed Ali, Adam Jatowt
๐ https://t.co/HsaSXT0OVS ยท ๐ป https://t.co/KsdsVIR3Ys
6๏ธโฃ Large Language Models Systematically Favor Popular Options: Evidence and Mitigation Across MCQs
๐๐ฏ๐ฑ๐ฒ๐น๐ฟ๐ฎ๐ต๐บ๐ฎ๐ป ๐๐ฏ๐ฑ๐ฎ๐น๐น๐ฎ๐ต, Mohammed Ali, Bhawna Piryani, Mahmoud Abdalla, Adam Jatowt
๐ Coming soon ยท ๐ป Coming soon
7๏ธโฃ SmallReason-ColBERT: An Ultra-Small Late-Interaction Retriever for Reasoning-Intensive Retrieval
๐๐ฏ๐ฑ๐ฒ๐น๐ฟ๐ฎ๐ต๐บ๐ฎ๐ป ๐๐ฏ๐ฑ๐ฎ๐น๐น๐ฎ๐ต, Mohammed Ali, Adam Jatowt
๐ Coming soon ยท ๐ป Coming soon
8๏ธโฃ MoCA-Agent: A Market-of-Claims Code Agent for Financial and Numerical Reasoning
๐๐ฏ๐ฑ๐ฒ๐น๐ฟ๐ฎ๐ต๐บ๐ฎ๐ป ๐๐ฏ๐ฑ๐ฎ๐น๐น๐ฎ๐ต, AbdelRahim A. Elmadany, Sameh Al Natour, Hasan Cavusoglu, Adam Jatowt, Muhammad Abdul-Mageed
๐ https://t.co/SOsbIiAN20 ยท ๐ป https://t.co/RuQorzl2LR
9๏ธโฃ The Magnitude Mirage: Rethinking Confidence for Reasoning-Intensive Retrieval
Jamie Holdcroft, ๐๐ฏ๐ฑ๐ฒ๐น๐ฟ๐ฎ๐ต๐บ๐ฎ๐ป ๐๐ฏ๐ฑ๐ฎ๐น๐น๐ฎ๐ต, Adam Jatowt
๐ Coming soon ยท ๐ป Coming soon
Findings
๐ REGREACT: Self-Correcting Multi-Agent Pipelines for Structured Regulatory Information Extraction
Mohammed Ali, ๐๐ฏ๐ฑ๐ฒ๐น๐ฟ๐ฎ๐ต๐บ๐ฎ๐ป ๐๐ฏ๐ฑ๐ฎ๐น๐น๐ฎ๐ต, Adam Jatowt
๐ https://t.co/6YAjLOzH1a ยท ๐ป Coming soon
Huge congratulations to everyone in the lab, and thanks to our supervisor @adammo . To anyone sitting on a pile of rejections right now: keep resubmitting. It works. ๐
2/ The idea: a region-aware Mixture-of-Experts inside the document encoder.
A router reads each region's content, its 2D position, and a pooled query context z_q, then mixes 4 latent experts (+1 shared).
The page is now encoded differently per query โ D(q). Still MaxSim.
1/๐๏ธ Meet Argus-Retriever: the first late-interaction visual doc retriever where the document representation adapts to the query: D(q).
- 86.0 NDCG@5 on ViDoRe V1+V2 (SOTA open model)
- ๐ฆ 1024-dim head, 4.5ร smaller index