Party is over, time to regularize ColBERT models to fix efficient ANN
MUVERA and SMVE promised to simplify multi-vector retrieval infrastructure but broke on modern ColBERT models
We found a fix, and it does the exact opposite of what we expected
๐We launch Evaluation Cards (beta): a centralized public record of AI evaluation results ๐
Not another leaderboard. Every score comes with who ran it, the settings they used, what the benchmark tests and the other results reported for the same model, side by side. ๐งต๐
Tired of burning GPUs on fine-tuning and evaluation runs just to see if your retriever's training data is actually any good?
Worry Not! Introducing ECIsem (Semantic Residual Effective Contrastive Information), a pre fine-tuning metric that scores your triplets...
1/6
Can we just use a simpler alternative? We ran a stress test against basic variants like raw hardness or unweighted residual diversity.
We find that they fail to reproduce the true ranking. Every component of ECI is necessary to get data selection right.
5/6
Going to be joining @LightOnIO for the summer!! Will be working on some really cool search tools ๐ต๏ธโโ๏ธ๐
And thanks @AmelieTabatta & @raphaelsrty for the opportunity :)
Search Models that will be trained by meโฌ๏ธโฌ๏ธ
in case you missed it, OBLIQ-Bench is now on arXiv: https://t.co/RbL5aBOq4B
my hope is that this reduces the frequency of IR or search agents papers that I discard immediately as a reader because in 2026 theyโre still evaluating on long-expired MS MARCO, NQ, HotPotQA, BEIR, etc