Protein2Text: Providing Rich Descriptions for Protein Sequences
1. Introducing Protein2Text, a novel model that bridges protein sequences and natural language, generating rich textual descriptions of protein properties, functions, and roles.
2. The system leverages BetaDescribe, a model derived from LLAMA2, trained on over 120 billion biological and English tokens, enabling seamless integration of biological insights into generative language capabilities.
3. Protein2Text excels at describing proteins with low sequence similarity to training data, outperforming traditional methods like BlastP, especially when homologous sequences are unavailable.
4. The model comprises a generator for creating descriptions, validators for property prediction, and a judge to assess accuracy, offering robust multi-perspective outputs.
5. Key innovations include its ability to identify functionally important regions via in-silico mutagenesis, revealing biological meaningful domains without experimental mutagenesis.
6. Compared to public large language models like GPT4, BetaDescribe demonstrates superior performance in protein-specific contexts, with higher accuracy and contextual relevance.
7. This tool advances functional protein annotation, with implications for medicine, agriculture, and protein engineering, and suggests potential for reverse application in protein design.
@boknilev
💻Code: https://t.co/gwbQnnxlfZ
📜Paper: https://t.co/qUsjvJtO5T
#ProteinBiology #GenerativeAI #Bioinformatics #ProteinFunction #NLP
@sivil_taram We report a significant performance boost on biological tasks when increasing the size of the tokenizer, as detailed in "Effect of Tokenization on Transformers for Biological Sequences" (Bioinformatics, 2024). https://t.co/42HErtEnXE
Special thanks to Gal Jaschek for their invaluable contributions, as well as to my supervisors Prof. Tal Pupko and Dr. @boknilev for their guidance and support throughout the project.
Code: https://t.co/H1qK70A3Bi
Paper: https://t.co/42HErtEnXE
I'm pleased to announce that our paper, "Effect of #tokenization on transformers for biological sequences" has been accepted for publication in #Bioinformatics.
TL;DR: our findings suggest that integrating a #tokenizer into the training process of a #deep-g#learning model for #biological#sequences can significantly enhance performance.
In exactly one month - we'll be presenting our model editing method -- TIME -- in #ICCV2023!
Text-to-image diffusion models encode a lot of assumptions about the world, which is what allows them to generate beautiful images even with simple prompts. BUT >>> 🧵
1/7 Can you imagine translating genomic data like you would with a foreign language?
Our latest research paper was accepted to ICLR utilizing seq2seq (translation) methods for a bioinformatics task! 🧵
👇 read more below
#Bioinformatics#SequenceAlignment#NLProc#ICLR2023
📢 Exciting News! 🧬🌐
We are thrilled to announce the launch of GenomeFLTR, our webserver that simplifies the process of filtering reads! Say goodbye to the complexities of identifying contaminants and embrace the power of GenomeFLTR.
🧵1/4