"Understanding Transformers and Attention Mechanisms" is a very interesting paper that presents the Transformer architecture from the perspective of applied mathematics.
It starts by representing text as vectors and explains mathematically how the attention mechanism processes these vectors to encode contextual information. It then develops Multi-Head Attention and shows how the main components of the Transformer architecture are constructed.
The paper also discusses more recent methods designed to reduce the computational and memory costs of attention, including KV caching, Grouped Query Attention, and Latent Attention. I think it is a useful reference for anyone interested in understanding Transformers beyond their high-level architecture and in seeing the linear algebra behind modern language models.
https://t.co/Rlun9QT7zx
Anthropic Pays $750K salaries for people who understand how LLMs actually work.
Not prompt engineers.
Not API wrappers.
Real builders.
And Stanford University just dropped a full lecture breaking it down
for free.
1 hour.
That’s it.
Watch it today before it disappears.
DeepVRegulome: DNABERT-based deep-learning framework for predicting the functional impact of short genomic variants on the human regulome
1. DeepVRegulome introduces a novel deep-learning approach to predict the functional impact of non-coding mutations on the human regulome, leveraging DNABERT fine-tuned models trained on extensive ENCODE data. This framework addresses a critical gap in understanding the effects of short genomic variants in regulatory regions.
2. The framework integrates 700 DNABERT models fine-tuned on various regulatory elements, including splice sites, transcription factor binding sites (TFBS), and histone marks. It achieves high accuracy in predicting functional disruptions, with models for TFBSs and splice sites demonstrating robust performance across diverse metrics.
3. DeepVRegulome incorporates an in silico mutagenesis module that systematically assesses the impact of mutations on regulatory sequences. By comparing reference and mutated sequences, it quantifies functional changes using log-odds ratios and attention-based analysis, providing detailed insights into variant effects.
4. The study applies DeepVRegulome to glioblastoma whole-genome sequencing data, identifying thousands of high-impact variants associated with patient survival outcomes. This application highlights the framework's potential in uncovering clinically relevant non-coding mutations in complex diseases like glioblastoma.
5. DeepVRegulome includes a motif visualization and interpretation module that derives and validates regulatory motifs from learned models. Over 87% of the models learned motifs with significant similarity to known JASPAR profiles, demonstrating the framework's ability to capture biologically relevant features.
6. Survival analysis links predicted functional disruptions in TFBSs and splice sites to patient outcomes, revealing novel variants associated with survival. This analysis underscores the clinical relevance of non-coding mutations and their potential as therapeutic targets in glioblastoma.
7. The framework is modular and data-agnostic, making it easily extendable to other cancer types and multi-omics datasets. Its open-source nature and interactive dashboard facilitate community adoption and further research into the regulatory genome.
💻Code: https://t.co/SGRkLv18vo
📜Paper: https://t.co/rtiSmMJkjK
#DeepLearning #Genomics #NonCodingVariants #RegulatoryGenome #PrecisionOncology #ComputationalBiology
Grateful for an incredible day at @SIGMODConf#SIGMOD2023#PODS2023. It was an honor to be the keynote speaker, surrounded by so many innovative minds. Your engagement and curiosity are what make these moments truly remarkable. Thank you for your passion and drive! @PyG_Team@Kumo_ai_team
Our recent collaborative work "Rapid, High-Throughput Single-Cell Multiplex In Situ Tagging (MIST) Analysis of Immunological Disease with Machine Learning" got published in @analyticchem https://t.co/ECsfAUCPmw
Our latest research work "Deep Multi-Omics Integration by Learning Correlation-Maximizing Representation Identifies Prognostically Stratified Cancer Subtypes" has been accepted for publication in the upcoming journal issue of @BioinfoAdv
https://t.co/FKVQu7hkk9
Our ICLR 2023 spotlight paper introduces LAMP, a deep learning-based surrogate model for multi-resolution physics. It optimizes computation cost and resolution for accurate simulation of dynamic regions, with controllable error-computation tradeoff.
https://t.co/ddr5C0j281
Interested in PyTorch, PyG and GNNs? Come work with me and the amazing @Kumo_ai_team. We have an opening at https://t.co/MmnRmj1xsb for a Machine Learning Engineer to work on both OSS and production ML workflows. Remote work possible. More information👇
https://t.co/WmAxwaOm3P
My lab is looking for a postdoc to work on some exciting LLM and Multimodal work. Please fill this simple form if you are interested: https://t.co/dG0YM6OF8A Colleagues and friends, please retweet to spread the words.
The Davuluri lab webpage is now live! Please visit to learn more about the exciting work from @RamanaDavuluri, @dutta_prat, @PallaviSurana5, and other amazing researchers @sbubmi@SBUResearch
What a vibrant & fantastic weather at #AGBTPH@agbt precision health 2022. For last couple of days, got the opportunity to listen and discuss the cutting edge research of various stalwarts from academics & industry @TalkowskiLab@dgmacarthur@euanashley @sbmontgom @L7informatics
We've released the ESM2 protein language models up to 15B parameters today.
Available now at the FAIR/@MetaAI Protein Team ESM repo: https://t.co/a2h11hGRja