@sokrypton of course it's needs less parameters if it can "cheat" with an MSA, than non-alignment based ESMC or no? the model needs to infer/learn less biology by itself so it needs less capacity
@DamienTeney Do your attention patterns reflect the nested structure or are they nearly random?
in this paper they showed on a transformer that it can learn dyck while being random:
https://t.co/GjjdmNdrZp
Really awesome paper from @gagneurlab on a new, clever interpretation approach showing that local DNA language models (trained on pre-defined functional elements across species) are capable of learning regulatory & structural syntax.
https://t.co/qOCZdhsRNl
Adding to @owlposting's remark on the capabilities of Evo 2:
There's a great paper by @gagneurlab showing that DNA LMs implicitly learn Watson-Crick base pairing from tRNA DNA.
https://t.co/yH6vSxPcT4
A socratic dialogue over the utility of DNA language models (Part 1 of 2)
here's the link: https://t.co/JTg4jzULKP
and here's a longpost of why i wrote this:
i think the effort that went into Evo 2 is very cool and its clearly a very comprehensive paper
but the excitement over it made me realize that i didn't understand a more basic concept: what's the point of a DNA language model? it felt like all the instinctive 𝕏 takes i read about them were just...wrong at worst, and overly optimistic at best. im sure a Real Genomics person would instinctively understand the utility of such a type of model. but i do not!
this is made worse by everyone i know irl agreeing that they too dont really get the point of models like these
this essay is an attempt to rectify my own understanding and hopefully help others too. i interleave in my own instinctive questions with the answers i stumbled across as i researched more. unfortunately, i have many dumb questions, but hopefully some smart ones too
part 1 is specifically focused about variant pathogenicity prediction using these models
i should note that this essay is not about Evo 2 specifically. Evo 2 is referred to heavily, specifically their pathogenic variant discovery results, but i do not spend much time on the data/model/etc results. it is intended to be more broad than that
Upon mutation of one of two each-other-binding nucleotides, the DNA LM adapts its probability for the other one according to Watson-Crick.
But of course, MSAs + co-variation analysis detects the same yet require explicit analysis and more effort, so the DNA LM has some benefit.
PLMs are great at mutation pathogenicity prediction, but how about functional effects like enzyme activity?
--> PLMs don't work for functional effects, but can be fine-tuned with an inductive bias regression head to perform better
https://t.co/3HZqyb51vA
Our findings - 1/n:
@adic_9 Hm that's unfortunate. Functional Effect prediction is just really hard and the model's aren't there yet, but we'll try to improve!
However, you could model the mutation or predict the mutant & wildtype structure + substrates with AF3 and analyse differences if that helps?
@adic_9 There are no weights to share. ESM-Effect fine-tunes the freely available ESM2 for every specific protein/DMS anew. And with the shared notebook anyone can fine-tune their own ESM-Effect on the respective DMS. Works well on google colab bc it‘s so efficient compared to PreMode.
And a big thank you to @braegelmannlab for the thoughtful discussions & review of this work!
Also thanks to @GoogleColab for providing the free (yet old) GPUs I used for this project😅
-> New pre-training objectives with deeper biological insights
Sequence & structure was the first frontier.
But biology lives & dynamically interacts and modelling that is the next frontier imo
Paper relating to that, by @francescazfl & @KevinKaichuang:
https://t.co/0TcHoTwyQg