Very proud to present our latest work: 7 guiding questions to avoid data leakage in biological machine learning applications ✨🔍 We hope that reflecting on these questions helps researchers to identify issues or shortcuts leading to overly optimistic performance estimates. 📈🧑🔬
A Perspective from @itisalist @judith_bernett@RomanJoeres@ok55991@FloHasee@dg_grimm @bit_tumcs & @dbblumenthal discusses the issue of data leakage in machine learning models and presents 7 questions to identify and avoid problems as a result. https://t.co/R5E9vO23qA
So happy to announce that my paper with @itisalist and @dbblumenthal "Cracking the black box of deep sequence-based protein–protein interaction prediction" is finally published at Briefings in Bioinformatics https://t.co/nU53pCgnBk ! So what is it about? 1/13 🧵
@lipido @itisalist @dbblumenthal@hlfernandez Thank you, that's great to hear! In our tests, Topsy-Turvy was the method with the highest performance on our gold standard dataset. Since then, some models have been published that beat its performance, e.g., 10.1101/2023.11.09.566187 or TUnA (10.1101/2024.02.19.581072, 65% Acc)
What is the takeaway?
📈High acc. can be reached with simple methods for known proteins -> Know your prediction task and try baselines first!
🔮Current seq.-based methods aren't made for predicting the "dark interactome"
✅We made a leakage-free dataset for future development
12/13 🧵Because this strategy rendered most datasets too small for proper DL, we designed a larger gold standard training (163,192)/val (59,260)/test (52,048) dataset using the same partitioning strategy. The best method achieved 56% accuracy on it. https://t.co/7vN103qAN6
Happy and excited to finally share this project!
We show conclusively that high accuracies of deep learning-based PPI prediction models are exclusively due to data leakage via sequence similarities and node degree information.
https://t.co/nd8WcME747 @itisalist @dbblumenthal
Network-based disease module mining tools often yield non-robust modules and are prone to random bias. To address this problem, we’ve designed a new method using enumeration of diverse Steiner trees: https://t.co/6gOhqyFKQB @janbaumbach @KacprowskiTim @judith_bernett @itisalist