Polygraph provides a toolkit to analyze sequence and k-mer composition, motif abundance, and predicted activity with pre-trained models. Polygraph also uses a human DNA language model to quantify the log-likelihood of synthetic sequences and score their "humanness".
We apply Polygraph to two datasets of synthetic yeast promoters and human enhancers. Polygraph is the first instrument for assessing synthetic regulatory sequences, enabling progress in therapeutic applications and improving our understanding of gene regulatory mechanisms.
We can now computationally engineer regulatory elements to drive cell-type specific expression, but it’s still difficult to know which mechanisms they exploit, if they converge on the same grammar as endogenous enhancers and promoters, and which sequences to test in the lab.
We’re excited to share Polygraph, a Python framework for evaluating native and synthetic DNA regulatory elements! @lal_avantika @gokcen https://t.co/mJixAWc2Mb
Excited to share our work developing experimental strategies for deep learning models to decipher how DNA sequence informs 3D genome folding! Our study points to understudied links between repetitive elements and chromatin organization. Link: https://t.co/CtcU7EbjMS (1/12)
Hooray, it's out! We use a deep learning model to screen the genome and find non-coding RNAs, Alus, SVAs, other transposable elements as sequence determinants of 3D genome folding.
New insight on how to compare two Hi-C maps quantitatively
TLDR - not trivial at all!
Interpretations can greatly depend on method choice
From @lauragunsalus + @EvonneMcArthur
Excited to share our work developing experimental strategies for deep learning models to decipher how DNA sequence informs 3D genome folding! Our study points to understudied links between repetitive elements and chromatin organization. Link: https://t.co/CtcU7EbjMS (1/12)
@XinyangBing To your second point- yes this is a legitimate concern. We looked at sequence mappability and didn't find a correlation with our measure of genome sensitivity (check out S4 and the supplemental note for more info).
Overall, we highlight a diverse vocabulary of repetitive elements that collaborate with CTCF to inform genome folding. Thanks to Katie Pollard, @keiser_lab, @gfudenberg + @drklly, whose work we build on, and the many others who gave feedback along the way. @4DNucleome (12/12)