New kind of non-reproducible science -
or how to write 'Data availability' statement without actually sharing the data.
I need help to decide if issues that we experience with obtaining published data is 'fraud', 'just unethical' or i am overreacting and 'it's ok' - please vote and RT.
Here is the story:
Human data are precious. It takes work to assemble cohorts, to collect samples and measurements and perform analysis of the data.
This is why many human immunology/biology/physiology researchers in the field are eager to see new publications that report new datasets and analysis.
In addition to actual findings, we all are eager to look at the reported original data and reanalyze, compare to other published data or own unpublished data.
The expectation (and the policy) is that unless these are genetically identifiable data, all the data required to reproduce the publication will become public upon the publication.
Some of such data are actually rather broad and can be extremely fruitful for the community analysis - Olink or Somascan proteomics, data on the disease development timelines in cohorts etc.
In one recent paper from Chinese Kadoori Biobank
https://t.co/YU6Wk4cds0
published in Nature Communications @NatureComms very recently, they describe very nice comparative analysis of the Somascan and Olink data and then go to establish predictive risks for some diseases.
Data availability statement says that any "bona fide researcher" can access the data following the procedure in CKB biobank. And some data are shared right away - overlap of somascan and olink but not actual somascan and olink data,
which precludes from reproducing the comparative analysis. But most importantly, analysis of the disease risks cannot be reproduced at all since to quote data availability statement - "The fully linked and integrated proteomic data
with other data will be made available through the CKB Data Access System during 2025."
I didnt know this is in option in academic publishing! In my naive view if you cant share the data now - you/journal should wait with publishing your paper until the data can be shared. But it seems that it went through the editors just fine.
Can anyone from @NatureComms or broader community explain to me if this is some new standard of data availbility and i simply didnt get the memo?
We contacted the CKB group, followed all the procedures for access etc, yet they insisted on a collaboration in order for us to analyze this published data. While I’m generally open to collaborations, the experience contrasts sharply with the statements on
the Chinese Kadoorie Biobank website about openness and data accessibility. Despite their public claims about serving the greater good and advancing research through data sharing,
in practice they have been notably resistant to actually sharing data.
Updated #SnakeCharm plugin 2024.1.1 has been just released. Now it is compatible with #PyCharm 2024.1! 🐍✨ Enjoy the power of the productive #Bioinformatics#pipelines coding 🚀 using #Snakemake. Learn more at https://t.co/tejouB4EGY
🎉 Unveiling #SnakeCharm 2023.3.1, our #Snakemake language plugin for #PyCharm 2023.3! 🐍✨ Enjoy the power of the new PyCharm AI Assistant (👉https://t.co/bfPHKrI51k) for more productive Snakemake coding! 🚀
Learn more at https://t.co/bczAWt9Oig
#JetBrainsAI#Bioinformatics
📢 Need a reason to test-drive the upcoming release of your favorite IntelliJ-based IDE or .NET tool?
With this week's EAP builds, a major new feature is introduced: AI Assistant.
For more details, read our latest blog post.👇
https://t.co/N1Pm04zdvm
#TheDriveToDevelop
Very glad to see MethPhaser out on bioRxiv https://t.co/ZrIRuuCX3Q! While @nanopore reads are already great for SNP phasing, we show that adding methylation information, which you automatically get with native @nanopore reads, can enhance things even further!
#SnakeCharm Plugin 2022.2.761 update adds support for ’exclude’ keyword in 'use' rules and lambda args in 'conda:' section. See release notes https://t.co/5ccNACSN2u. Plugin is compatible with #PyCharm 2022.1.x, 2022.2.x #bioinformatics#pipeline
#SnakeCharm Plugin 2022.1.749 update adds support for “default_target”, “retries”, “ensure" directives + bug fixing. See release notes https://t.co/iKkOY4MRpJ. Plugin is compatible with #PyCharm 2022.1.x #bioinformatics#pipeline
In May 2000 (22 years ago!), the Human Genome Project completed a draft sequence of human chromosome 21. This is the smallest chromosome in the human genome (with fewer than 300 genes), but it was a historic triumph!
Google photorealistic text-to-image model ‘Imagen’ is amazing https://t.co/5DMorZqcYt E.g.: “A photo of a raccoon wearing an astronaut helmet, looking out of the window at night.” or “cute corgi lives in a house made out of sushi.”
#SnakeCharm plugin update 2022.1.743 introduces quick navigation among overridden rules and adds support for #PyCharm 2022.1. See release notes https://t.co/ZCqQW5Ty6k. Plugin released for #PyCharm 2022.1 and 2021.3. #bioinformatics#pipeline
There is an open bioinformatics position at the west german cancer center (unfortunately only in german). The PI Jürgen Becker is a fine guy: https://t.co/pwjQ202Gzy
#Snakemake 7.0 is released. Three main new features: service jobs (providing shared memory resources or databases), native template rendering support, and improved cluster submission control. https://t.co/dirRKpOHvr #sciworkflows#reproducibility
New #SnakeCharm plugin major update adds support for #snakemake 6.x use syntax, color settings page, initial pep code completion and many other fixes. See https://t.co/bK0m8Tg9xq. Plugin released for #PyCharm 2021.3 and 2021.2. #bioinformatics#pipeline
If you're looking for a new, outstanding review of covid testing, here it is https://t.co/thOb8BdiMx @NatureRevGenet by Tim Mercer @UQ @UQScience and @bioetalons