🚀 Exciting news! 🚀 The code for DataFix, our latest NeurIPS paper focusing on shift detection and correction, is now available on GitHub!
https://t.co/xRIJfDY1Ei
@_danielmas @ritageleta @DocXavi@alexGioannidis
Lots of standardized benchmarks.
Phase error correction for LAI, the only scalable method with this.
Insights...including that Native American dogs 🐕 did not vanish; living on (in part) in the Mexican Xoloitzcuintli. Also, Chihuahuas have a bit of coyote in them.
Credit also to contributor,
@Rich_Rast
And maestro,
@_danielmas
Use it with our toolkit,
https://t.co/bRBktspSzP
Gnomix is published!
https://t.co/8wq4951VCF
Still the most accurate & fastest local ancestry inference (LAI) method, bests Flare and others (Recomb-Mix & Orchestra) released since our preprint.
Easy use
https://t.co/RoAzHobn7e
Leads @HelgiHilmars@arvind0422@BarrabesMiriam
snputils (https://t.co/pyf7FyY1dF): A High-Performance Python Library for Genetic Variation and Population Structure https://t.co/wEi2mHYFbA 🧬💻🧪 https://t.co/bJALQDZpV2
Excited to share our preprint on iLTM: an Integrated Large Tabular Model!
arxiv: https://t.co/ZpT5ZiKJs9
No single technique consistently excels across all tabular tasks. iLTM addresses this by integrating distinct paradigms in a single architecture:
- Gradient Boosted Decision Trees (GBDT)-based embeddings
- Robust preprocessing with random feature projections
- A meta-trained hypernetwork
- Retrieval augmentation with Soft Nearest Neighbors
@d_bonet@marcal_cc@alexGioannidis
(1/N)
We’re excited to be part of #PMWC2025 —Our session: “Powering Discoveries in Medicine by Pairing Genetic Diversity with AI-enabled Data Analytics” is part of Track B and will be delivered by our CEO, Carlos D. Bustamante on February 7th at 10:15am PT.
https://t.co/J6Q2N7cBCD
Excited to share our latest PRS work! Our @GalateaBio and @genomelink team performed a comprehensive analysis of published @PGSCatalog models along with locally trained models using LDPred2, PRS-CSx, and SNPnet, across diverse populations using @UKBIOBANK and our own data
Introducing "HyperFast: Instant Classification for Tabular Data" at @RealAAAI, which received the Best Paper Award at @NeurIPSConf Table rep. workshop @TrlWorkshop!
We provide easy-to-use sklearn-like code: https://t.co/qrMF6XStAA
Some insights of the work below 👇🧵(1/N)
Loving the annual reunion at the Deep Learning Barcelona Symposium! Grateful to everyone who expressed interest in our work on homogenizing genomic databases 🧬. And a special shoutout to @ritageleta for delivering an excellent talk on our DataFix tool (NeurIPS) #DLBCN#Catalonia
Also, congrats🙌 to another amazing student @BarrabesMiriam together w/ @_danielmas, @ritageleta & @DocXavi, whose DataFix system for feature shift detection, inspired by merging genomic datasets, but much more general, was just presented at #NeurIPS2023.
https://t.co/Lkktywsfca
We presented our paper paper "Adversarial Learning for Feature Shift Detection and Correction" at @NeurIPSConf
We want to thank @ykilcher for comming by our poster!
https://t.co/hidnpTpTy4
#NeurIPS#NeurIPS2023#AI
Combining data from many sources to increase the sample size is a common step in biomedical studies.
Merging datasets typically involves domain-specific heuristics. But, how can you evaluate the quality of the final dataset?
DataFix can localize and correct potential errors!
The work combines iterative heuristics with Random Forest and Boosting Trees in order to localize and correct the features that suffer a distribution shift providing state-of-the-art performance.
@NeurIPSConf#NeurIPS#NeurIPS2023#AI
(3/3)
Our method DataFix for shift detection and correction is now available on GitHub:
https://t.co/T7DclMZymC
DataFix is presented in our paper "Adversarial Learning for Feature Shift Detection and Correction" at @NeurIPSConf#NeurIPS#NeurIPS2023#AI
(1/3)
The paper "Adversarial Learning for Feature Shift Detection and Correction" was co-lead by @BarrabesMiriam with the excellent help of @ritageleta @DocXavi@alexGioannidis
DataFix provides an effective way to localize and correct corrupted features within a dataset.
(2/3)