Excited to share the first preprint of my PhD @sangerinstitute looking into the prevalence and genetic mechanisms of gene misexpression in 4,568 whole blood RNA-seq samples from the INTERVAL cohort. https://t.co/5Ia0Z1YVPR
Great to get my last PhD project out and fully published! Focusing on high-throughput as well as accuracy led to interesting and different analyses of variant effect predictions
The last PhD project from @Ally_Dunham is now out in published form. A collaboration with @MoAlQuraishi where Ally developed a very fast protein missense variant effect predictor using a convolutional neural network model. https://t.co/9EnNhGQSTx
Along the way we also made a python package for working with ProteinNet structure/sequence data. It gives a convenient interface to read, process, write and train from ProteinNet files https://t.co/3rXhxUpKbM
The last PhD project from @Ally_Dunham is now out in published form. A collaboration with @MoAlQuraishi where Ally developed a very fast protein missense variant effect predictor using a convolutional neural network model. https://t.co/9EnNhGQSTx
Our story on inserting short sequences with prime editing is now online @NatureBiotech.
https://t.co/n7WiehAaMb
So much has changed from the preprint including the discovery that 3'-flap nucleases can inhibit the insertion of long inserts. Lots to unpack 🧵
@KevinKaichuang@pedrobeltrao@MoAlQuraishi I just checked and this padding does slightly effect scores in the extreme cases. However, they are very highly correlated (r > 0.98) so classifications are unlikely to be effected. It would be good to quantify this more though.
@KevinKaichuang@pedrobeltrao@MoAlQuraishi The model is fully convolutional and produces per position results so variable length is built in. Having to half the length in each block means length must be divisible by 2 six times though so we 0 pad to ensure that
Great to share the last project from my PhD, hopefully people will find it useful!
Code and package: https://t.co/lObymg3JCc and a package for working with ProteinNet: https://t.co/3rXhxUpKbM
The last PhD project from @Ally_Dunham was a collaboration with @MoAlQuraishi on a convolution neural network model for protein variant effect prediction. It achieves fast effect prediction without alignments. https://t.co/GxtC5qy6bo
We joined a large community effort to assess diverse applications of AlphaFold 2 in the context of novel structural elements; missense variants; function and ligand binding sites; modelling of interactions and experimental structural data. Some highlights below:
@dbradley534@pedrobeltrao Yes, we do repairPDB. Probably worth it for consistency with the experimental models even if the AF models are mostly relaxed already
One final alphafold thing :) - @Ally_Dunham compared deep mutational scanning data with predicted changes in stability for mutations on the alphafold models (with FoldX). The correlations observed are typically as good or better than with experimentally derived structures
Now in published form - @Ally_Dunham combined experimental outcomes of mutations for 6291 positions in 30 proteins, showing these can group into different functional groups. Ally then studied these in the context of protein structures and evolution.
Pleased to release our work on computational varaint effect prediction for all possible SARS-CoV-2 substitutions. Hopefully the results will be useful for studying viral variants. https://t.co/p4I7M8kf0C https://t.co/R655H2s4iT
What SARS-CoV-2 protein variants may have a functional impact? @Ally_Dunham predicted the impact of all possible variants on conserved regions, structural elements and experimental annotations (e.g. escape mutations). https://t.co/U9SY0sURp3
New lab preprint - Deep mutational scanning can tell us about the function of individual protein positions. @Ally_Dunham normalised DMS data for 6291 positions in 30 proteins to ask how many amino acid functions are there and how frequently are they used https://t.co/Lpo5ULujTm