Just want to give a shout-out to David Kelley @drklly who I think often does not get the credit he deserves (outside our core community).
I want to highlight why I think he is such a fantastic scientist and leader in regulatory genomics. 1/
A common (MAF=9%) loss of function variant in CD36, encoding a fatty acid transporter, is found to be the major genetic risk factor of dilated cardiomyopathy (DCM) in African populations. 8% of the DCM cases in African populations could be explained by this single variant. Mind blowing!
Early days of the genome-wide association study (GWAS) era was fun with scientists competing each other to get their hands on the "low-hanging fruits" of genetic risk factors in human diseases.
Many discoveries were made in the early days such as
- CFH Y402H missense variant associated with AMD in Europeans
- PNPLA3 I148M missense variant associated with fatty liver disease in Europeans
- APOL1 risk variants associated with chronic kidney disease in Africans
Most of such discoveries were made in European populations as early GWAS studies involved only them, with few exceptions such as the APOL1 discovery (in which case the variant frequency and effect sizes were too big to not find).
With GWAS sample sizes increasing to hundreds of thousands, discoveries of low-hanging fruits got soon saturated with no major common variant genetic risk factors left to discover. But that was obviously true only for European populations.
We have still a lot of major genetic risk factors left to discover in non-European populations. With increasing sample sizes of Non-European populations in major biobanks, such discoveries are beginning to surface. Below is a great example.
A GWAS of DCM involving 1,802 cases and 93,804 controls of African ancestry in the Million Veterans Biobank identified a single GWAS locus driven by a nonsense variant in CD36 with a relatively larger effect size (het OR=1.25, hom OR=2.74) and high allele frequency (maf=9%).
The authors estimate the population attributable fraction for this variant to be 8.1%, which is comparable to some of the major clinical risk factors like BMI and type 2 diabetes.
The biology side of this discovery is also fascinating. Heart uses free fatty acids as the primary fuel for energy and the protein encoded by CD36 is a fatty acid transporter acting as the gateway for fatty acids to enter the heart cells. Loss of this transporter deprives the heart muscle of proper energy supply and makes it weak, which somehow is leading to dilated cardiomyopathy.
Around ~1% of African populations (this number will vary across sub populations) are homozygotes for this variant, who will now become target population for a CD36-focussed therapeutics (similar to APOL1. and PNPLA3 story).
Love this discovery, and easily one of my favorites of this year.
Huffman, Gaziano, Al Sayed, et al. Nat Gen 2025
https://t.co/Xna6b1ra5r
Introducing scE2G: a new model to link enhancers to target genes using single-cell data.
Excited that scE2G will enable building enhancer maps in hundreds of cell types in the human body!
Wonderful collaboration with @robin_andersson@613weilin@mayayayas and others
👇
Hi! I'm hiring a Research Engineer to join my team at Google DeepMind for the year. You'd be working with a great, interdisciplinary team on AI evals. Please share if you know anyone who might be interested!
https://t.co/g9ezvWTGI5
Note: this is a fixed-term, 12 month position
I´m thrilled to see our #LRGASP paper published in @naturemethods today. This is the work of many to benchmark long-read methods for transcriptomics, and a must-read paper for those using lrRNA-seq. Special thnks to @FJPardoPalacios and @TheBrooksLab (1/6)
https://t.co/xdLwq3LvLU
Our paper is out in @nature! This is from a @cshlbanbury meeting where a group of scientists got together to ask, can we ever identify the complete set of human genes? And how do we do that? with @elapertea@av_sparrow@carninci and many others https://t.co/YWXUcCAfFy
Interested in pursuing a PhD in Bioinformatics and Genomics? Our Penn State interdisciplinary program is now accepting applications! https://t.co/rlMBSexTRo
kmers are powerful in sequence comparison, but do not perform as good when the error rate is high; we have a solution: use subsequence. For details check: https://t.co/iDo6pLdxll. I’ll present it at #ISMBECCB2023#HiTSeq. Congrats to co-authors @xianglipsu@QianShi15@kanatos92!
Wonder why the BUSCO completeness of T2T-CHM13 is only 95.7%? Tricky protein-to-genome alignment. Try miniBUSCO. It is 1) much faster, 2) more accurate for well annotated lineages, 3) robust to frameshift errors, and 4) more lightweight. Fine work by @csuhuangneng
With plenty of colleagues, we are happy to introduce bioconvert, a software to convert bioinformatics files from one format to another. bioconvert currently includes 50 formats and 100 converters.