How reproducible are #bioinformatic pipelines really? 🤔
In pursuit of this question, I asked three lab members to run my pipeline on the same CUT&Tag data (4 conditions, 3 biological replicates, and 2 epigenetic marks). 💻
#Genomics#snakemake#AcademicChatter
Gene expression involves thousands of proteins that bind DNA, yet comprehensively mapping these is challenging. We developed ChIP-DIP – a method for simultaneous, genome-wide mapping of hundreds of DNA-protein interactions in a single experiment. https://t.co/NEH92mQTGs
Hierarchy composition in AML matters. Our work on the epigenetic accessibility profiles of these hierarchies is now out. https://t.co/OG1w40Uxkr @MDAndersonNews@LeukemiaRF
Speaker: because of the diff exp results, it supports my hypothesis 😁
Listeners: but what if you recluster it with different parameters? 🤔
Speaker: we don't talk about that here 🙁
Some news: After 8 years & 750 stories, I've decided to leave The Atlantic. Today's my last day.
Being a writer means you can’t say things like "I can’t tell you what this means" cos, well, I can. That's kind of the point of me. So here’s an attempt at looking back & forward: 1/
Yes, tSNE and UMAP overfit the data creating the impression of structure from nothing (even with 1000 obs & 200 features, not 2 -> 2). I don't have a twitter 'team' that I'm fighting for/against, but I had a hypothesis, tested it, and I'm just reporting the results...
I started to study #biology ten years ago. What nobody told me was that it (serendipitously) transforms into linear algebra, multivar statistics, and Greek 😁
#Bioinformatics#math
Ever analyze a scRNAseq dataset and wonder if a specific cell state has been seen before? And if so, where in the human body? Under what conditions? Well, now you can use our lightning fast SCimilarity search and foundational model for that! ⚡️🔎🧬 https://t.co/t8L1xFOffd (1/11)
Have you wondered why the overlap between ChIP-seq or CUT&RUN replicates is often so low?
Should one only trust the small overlapping set?
@AnnaNordin96@ppagella86 @GianlucaZamba and I tried to solve this conundrum central to signal detection theory 📈
https://t.co/Zz6woBHsVw
@Gurkan_Yardimci Took me a while to find but here it is:
https://t.co/VQKBKEqT2q
I noticed it has lots of topology sections which may not be directly to our work, so just gotta pick and choose
Do most #promoters have #H3K4me1 marks? YES !?
Do most #enhancers have #H3K4me3 marks? NO !?
Enhancers and promoters have different sequence composition #CGdensity.
Promoters are GC-rich and H3K4me3 binds to accessible CpG regions.
@schonrockz Not pharm, but if you are analyzing processed/semi-processed data (e.g. tables) then R is preferred bc of its comprehensive stats library and data manip pkgs. Python is often useful for general processes like downloading + cleaning files. Does this help? Good luck!
Updated usage statistics for #Snakemake: we are now at >10 new citations per week on average and over 700k downloads of the #Conda package (not even counting pypi). Thanks a lot for the active use and all the feedback and contributions! https://t.co/WwE4GKpryY