Excited to announce a new preprint that investigates the ability of sequence-to-function models (S2F; e.g., AlphaGenome) to identify causal expression modifying variants. (1\n)
https://t.co/1iktWUhfNI
@StevenSalzberg1@TuXinming@zhanxw Shameless plug, but maybe you’ll find this paper helps reconcile your finding/experience with our claim that alphagenome is not slop and can be quite useful (but also often is not. noncoding variant interpretation is very hard)
https://t.co/1iktWUhfNI
This suggests issues for fine-mapping and variant prioritization in most loci. But I also find that the very strongest predictions are highly reliable and supported by the data, suggesting AlphaGenome “hits” can often be trusted. (6/n)
I hope this work helps shed light on (esp. for less advanced users) important S2F failure modes -- and their prevalence causes -- and when variant effect predictions can and can’t be trusted (11/11)
AI holds huge promise for finding disease-driving genetic mutations, but top models frequently underestimate their impact.
New research exposes this underlying flaw and provides a roadmap to build smarter models for personalized medicine.
https://t.co/9WnSZqtDrD
In these cases AlphaGenome doesn't prominently upweight any other candidate SNV that could explain the eQTL association. So this isn't disagreement with fine-mapping; it underestimates all candidate causal eQTLs relative to nearby non-eQTLs. (5/n)
Excited to announce a new preprint that investigates the ability of sequence-to-function models (S2F; e.g., AlphaGenome) to identify causal expression modifying variants. (1\n)
https://t.co/1iktWUhfNI
It seems that yes, most causal eQTLs are underestimated by AlphaGenome (and likely other models), not just in absolute magnitude but also relative to nearby ones with no detectable expression association. Often severely enough that the causal eQTL is effectively ignored. (4/n)
Are S2F models ignoring most causal SNVs? Or maybe predicted effects are small, but are still stronger than nearby non-causal SNVs? Maybe fine-mapping is wrong and they aren’t causal, and models upweight other variant(s) instead that explain the observed eQTL association?
This paper was in part motivated by figures like this, found in most S2F papers. Predicted vs observed effect sizes for putatively causal fine-mapped eQTLs. Overall correlation looks good, but most points sit near a predicted effect of 0. What’s going on there? (2/n)