@bioinformer@StevenSalzberg1@crill_labs Genbank, since its inception, has the policy of being and "archival" database that contains what was submitted to it. "Contaminants" in Genbank have been well documented for more than 30 years. That is why there are also "curated" databases at the NCBI, like RefSeq.
@StevenSalzberg1@JillFilipovic@lizditz Or maybe Israel's killing 10's of thousands of Palestinians in Gaza is giving Hamas exactly what they want. It will certainly keep hate alive for another generation or two. Hard to know.
@phylogenomics (4) Don’t bother looking for ORFs - BLASTX and FASTX do that for you.
(5) Make sure you are doing DNA reads vs protein databases (BLASTX, FASTX) - 5 - 10x more sensitive.
@phylogenomics (2) Search with shallower scoring matrices. PAM70 for BLASTX and VT80 for FASTX. These work much better for shorter queries.
(3) Don’t search pFam or NR/Refseq. It is much more sensitive to search smaller databases that are targeted
@phylogenomics You do not need anything except computer time. Use BLASTX or FASTX to compare your reads (as DNA) against a phylogenetically appropriate protein sequence database. For bacteria, pick 20 - 50 well annotated bacterial proteomes (1-2 million proteins). /1
@CWolberger I can imagine the problem might be that some systems (or less technically astute reviewers) may automatically follow the links, thus potentially de-anonymizing the review process. Particularly one can put more information in the link than is displayed in the text.
What does ChatGPT know about the difference between FASTA and BLAST? It does an excellent job, as long as your substitute BLAST for FASTA (and vice versa) for most of the result. Both are heuristic, FASTA is slower, more accurate, and offers global alignment.
@LuciaScience @mikelove@daniela_witten@drisso1893 A few years back, @pkkimes@keegankorthauer, myself, and others wrote a review / benchmark of various flavors of FDR in application to comp bio problems. Let me know if this helps, or I can try to summarize here too!
https://t.co/US66Ww3PiV
Margaret Dayhoff’s Atlas of Protein Sequence and Structure is 50 years old. It has all the sequences printed together 😉 and some lovely structures, the definition of mutation data matrices and sequence databases. @UoDLifeSciences @emblebi @ISBSIB
@strnr These kinds of analyses can be very challenging. Many widely used software packages may not be available on GitHub, and "accuracy" can be very domain specific. Was BLAST one of the packages examined? How did it score on the "accuracy" scale? On the "development" scale?
@DrSalinasLab I'm very surprised to read this. Most of the researchers I know are happy to give talks to students -- our students often recruit more famous speakers than our faculty. The problem is usually time commitments (though perhaps less so this year). No more than one trip per ???.
@peevedccprof I'm sure the idea is correct (we don't like paying teachers), but a quick look at the Mississippi Dept. of Education salary scales suggests it is not accurate.