๐๐๐ MIRACL is here!
The corpora + training/dev splits for all 16 known languages are now on HuggingFace!
๐ฅ53k queries, with 470k human-labeled relevance judgments!
๐ฅFind the data here: https://t.co/B5iNhW6Xlv
๐Link to paper: https://t.co/aavuJiXQvY
How was MIRACL made? Check out the methods from our winners๐ฅ
NOT CIIR: https://t.co/xGyVmvkYwX
Bott: https://t.co/GZ99nNcnUn
NLE: https://t.co/apw8n66hQP
๐ฅWSDM Cup Day tomorrow!
Location: Level 4, Esplanade 2
(UTC+8)
8:30-10:00 Keynote and Overview
13:30-15:00 Talks from our winners!
NOT CIIR:๐ฅin Known track @huang_zhiqi@pxyumass@jallan_umass
Bott:๐ฅin Surprise track๐ฅin Known track: Qi Zhang
NLE:๐ฅin BOTH tracks @cadurosar
๐ข๐ข MIRACL Competition is now over!
We would like to thank all the participants. We received submissions from over 40 different teams participating for #WSDM2023 Cup! ๐๐
Our final leaderboard results can be viewed publicly below๐
https://t.co/yhCIuTt4Mr
Recently, @CohereAI boasted "3X better performance" in multilingual text understanding. We tested that claim by evaluating Cohere embeddings on MIRACL:
tl;dr - We weren't able to replicate the 3X claim, but we did observe a 38% improvement over BM25. https://t.co/hfYSdL5C22
๐ข๐ข๐ข Extra 4๏ธโฃ days for submission!
MIRACL competition deadline is now extended to Jan. 16th (AoE)
Time to submit your models and results on the MIRACL Leaderboard! โฐ
Participate and win those amazing cash prizes in the two tasks! ๐
Submit here๐
https://t.co/ENYHwcBpSL
๐ฅSurprise languages and Test-B topics released๐ฅ
How good your model perform on German and Yoruba in MIRACL?
๐ Data: https://t.co/8vVoVzjjSq
๐ Baselines (for surprise language dev set): https://t.co/YFeUOu6s2H
(Reproducible on Pyserini)
๐ One week left to win prizes๐ฅ๐ฅ๐ฅ!
Now that the semester is over and those GPUs will be idle during the holidays... what better time to work on MIRACL? #wsdm2023 cup challenge on search across 18 different languages https://t.co/fQtMDDCxxP Leaderboard: https://t.co/hfxBnNUyr4
๐ฅTopics of Test-A released!
๐ซย Available on HuggingFace Datasets
๐ https://t.co/8vVoVziM2S
โก๏ธSubmit to leaderboard to view the Test-A scores!
๐ฅ๐ฅ๐ฅ A MIRACL is coming tomorrow!
๐๐๐ 77k queries, 700k judgments on Wikipedia in 16 languages!
๐ train/dev sets drop tomorrow
Join our mailing list and Slack to stay up to date!
๐ฌ Mailing List: https://t.co/oaDzp4CG1w
๐ฌ Slack: https://t.co/UXAuSnqWH2
Our mailing list is out now! ๐
If you're interested in participating in Project MIRACL ๐๐๐ - come join our mailing list! ๐ฌ
For more details, checkout below ๐
If you're interested in participating in Project MIRACL ๐๐๐ - come join our mailing list! ๐ฌ
https://t.co/oaDzp4CG1w
- Prizes worth over 6k USD to be won!
- Opportunity to try out the largest multi-monolingual retrieval dataset!
Introducing MIRACL ๐๐๐(Multilingual Information Retrieval Across a Continuum of Languages), an WSDM 2023 Cup challenge that focuses on search across 18 languages, which collectively encompass over three billion native speakers around the world. https://t.co/IexzhUNyYQ