🧵We just released a large-scale reproducibility effort on unbiased learning-to-rank (ULTR). It turned out that ULTR did not work as intended ... or at least that its success depends on how your measure success.
Many thanks to Philipp Hager, @JMRenders, Onno Zoeter, @mdr
Along with the SharpeRatio metric, we publicized open source-software called “SCOPE-RL”!
This software aims to streamline offline RL and OPE procedures and also the evaluation-of-OPE in the RL setting 🚀
https://t.co/qOtQ2bS60T
This work was first and foremost an effort to make ULTR more accessible, so we release:
- Jax-based implementations for all methods, including BERT-like training for document features: https://t.co/9r3RECRJ64
- Models and pre-processed datasets: https://t.co/4hSMzeug8t
🧵We just released a large-scale reproducibility effort on unbiased learning-to-rank (ULTR). It turned out that ULTR did not work as intended ... or at least that its success depends on how your measure success.
Many thanks to Philipp Hager, @JMRenders, Onno Zoeter, @mdr
We should strive to understand the connections and differences between annotations and clicks, and when to prioritize one metric over the other. For now, we discuss additional observations and potential explanations for these results in our #SIGIR24 paper: https://t.co/QoeQO8zNwH
😀We're looking for a talented researcher to join our team at Naver Labs Europe (@naverlabseurope) , working on LLMs and Retrieval!😃
Please apply here: https://t.co/0lea7ABHld !
A few weeks ago, @OtmaneSakhi gave a talk at @sharechatapp on "Fast Off-Policy Optimization for Billion Scale Catalogues".
I'm happy I get to share it, because it's absolutely great.
Some really cool contributions in his thesis, have a look!
https://t.co/Dmwy5sHeQg
@haitao_mao_@mdr Thank you very much Haitao for the kind words, and of course for releasing the dataset in the first place 🙏 There are still many things to understand on this dataset (and perhaps for new Baidu datasets in the future 👀 ? ). Let's make ULTR work in real-world !
Agents equipped with neural networks gradually lose the ability to learn from new experiences.
In the #NeurIPS2023 spotlight done with @GoogleDeepMind, we carefully analyze the phenomenon in RL and offer a way to make agents more compute-efficient.
https://t.co/3dCkkpjH46
🧵
Incredibly, the placebo effect is (mostly) not real.
It is a result of statistical confusion. Whenever you have a group with extreme values, they tend to exhibit regression to the mean. Eg. on average, sick people tend to become more healthy over time.
Thus if you give one group medicine, and one group placebo, the placebo group will also tend to get better over time, because of regression to the mean.
People have then misinterpreted this to think that it is the placebo pill that actively does this.
If you want to demonstrate a placebo effect, you have to construct a study where there are three groups:
• A. treatment
• B. placebo
• C. no treatment, no placebo
If B and C get different outcomes, that would demonstrate a placebo effect.
When this has been tried, mainly there has been no provable placebo effect. See the paper in the screenshot. (There is some evidence for an effect for pain, but this get's into a slightly different debate.)
The fact that the placebo effect is mainly not real, fortunately frees us from having to come up with convoluted explanations, as to why the placebo effect would work even when we tell the patient that it is a placebo, as in the quoted tweet.
Excited to share our KDD’23 paper 📚🎉:
“Off-Policy Evaluation of Ranking Policies under Diverse User Behavior”
https://t.co/UHPWvQop2J
We studied how to enable an accurate OPE when users’ browsing behavior is diverse and context-dependent.
Many thanks to my collaborators!
🎉 Introducing our new work “Safety-Aware Unsupervised Skill Discovery”: a scalable method that can discover diverse set of skills that satisfies any user-defined safety constraints. #ICRA2023@NAVERLABSHQ 🧵
website: https://t.co/zVty8Zb1hg
slides: https://t.co/iHYFzQo0pr