Some cool work from our team at https://t.co/c2FVPdxRrz, led by @ferrangh and @alafitamasip, on improving variant effect prediction with DMS data will be presented at @MLGenX:
"Fine-tuning Protein Language Models with Deep Mutational Scanning improves Variant Effect Prediction"
📢 We have uploaded the final camera-ready version of the paper that @ferrangh and I will present at the @MLGenX workshop at #ICLR2024 this May
https://t.co/Fpi0AwDtQ7
🧵1/5
@DdelAlamo It's very much domain, vocabulary and pre-training task specific. There are some references below, probably worth checking for updated versions.
https://t.co/1UO74m9PfX
https://t.co/ScsNgbpcGd
I am really digging this kind of frank presentation of very impressive results from a startup. More of this please. Budding scientists take note. You can do impressive stuff while highlighting important caveats. It creates more trust.
The Illustrated NeurIPS 2025: A Visual Map of the AI Frontier
New blog post!
NeurIPS 2025 papers are out—and it’s a lot to take in. This visualization lets you explore the entire research landscape interactively, with clusters, summaries, and @cohere LLM-generated explanations that make the field easier to grasp.
Link in thread!
Multi-megabase scale genome interpretation with genetic language models
1. Introducing Phenformer, a pioneering genetic language model designed for direct genome-to-disease interpretation. Phenformer connects DNA sequences to disease-relevant expression changes across cell types and tissues at an unprecedented scale of 88 million base pairs.
2. Unlike traditional genome-wide association studies, Phenformer integrates whole-genome sequences without requiring experimental data, improving disease risk predictions and uncovering molecular mechanisms.
3. Phenformer outperforms existing methods by aligning its predictions more closely with scientific literature, offering insights into poorly understood disease mechanisms like liver disease in psoriasis patients and complications in type 1 diabetes.
4. The model excels in diverse populations, demonstrating improved generalizability over current polygenic risk scores (PRS). Phenformer enriches PRS methods, boosting prediction accuracy in both mixed and non-European ancestries.
5. A novel approach to disease subtyping: Phenformer identifies molecular subtypes, shedding light on genetic variation-driven differences in disease mechanisms, enhancing precision medicine capabilities.
6. Technological advancements: By leveraging transformer-based architectures and multi-scale data integration, Phenformer processes a genome's sequence context more comprehensively than other genetic models.
7. Ethical considerations are critical. Phenformer addresses biases inherent in genomic data and seeks to ensure equitable healthcare outcomes across diverse genetic backgrounds.
8. Future potential: Scaling Phenformer to cover larger genome fractions could further enhance its predictive power, paving the way for breakthroughs in personalized medicine and genomic interpretation.
@schwabpa@deboramarks@DeboraMarksLab@bschoelkopf@arashmeh@NotinPascal@f_traeuble
📜Paper: https://t.co/jM89EIOkee
#genomics #deeplearning #AI #healthtech #precisionmedicine #Phenformer
Multi-megabase scale genome interpretation with genetic language models
1. Introducing Phenformer, a pioneering genetic language model designed for direct genome-to-disease interpretation. Phenformer connects DNA sequences to disease-relevant expression changes across cell types and tissues at an unprecedented scale of 88 million base pairs.
2. Unlike traditional genome-wide association studies, Phenformer integrates whole-genome sequences without requiring experimental data, improving disease risk predictions and uncovering molecular mechanisms.
3. Phenformer outperforms existing methods by aligning its predictions more closely with scientific literature, offering insights into poorly understood disease mechanisms like liver disease in psoriasis patients and complications in type 1 diabetes.
4. The model excels in diverse populations, demonstrating improved generalizability over current polygenic risk scores (PRS). Phenformer enriches PRS methods, boosting prediction accuracy in both mixed and non-European ancestries.
5. A novel approach to disease subtyping: Phenformer identifies molecular subtypes, shedding light on genetic variation-driven differences in disease mechanisms, enhancing precision medicine capabilities.
6. Technological advancements: By leveraging transformer-based architectures and multi-scale data integration, Phenformer processes a genome's sequence context more comprehensively than other genetic models.
7. Ethical considerations are critical. Phenformer addresses biases inherent in genomic data and seeks to ensure equitable healthcare outcomes across diverse genetic backgrounds.
8. Future potential: Scaling Phenformer to cover larger genome fractions could further enhance its predictive power, paving the way for breakthroughs in personalized medicine and genomic interpretation.
@schwabpa@deboramarks@DeboraMarksLab@bschoelkopf@arashmeh@NotinPascal@f_traeuble
📜Paper: https://t.co/jM89EIOkee
#genomics #deeplearning #AI #healthtech #precisionmedicine #Phenformer
MLDD is returning for 2025!
We're thrilled to announce that the Machine Learning for Drug Discovery (MLDD) Symposium will return in 2025!
Join us for an inspiring day of knowledge sharing, networking, and discussions on the latest trends in applying machine learning to drug discovery.
📍 Location: London, UK
📅 Date: June 30, 2025
🤝 Format: In-person
This is your chance to connect with experts, gain fresh insights, and explore cutting-edge advancements in this transformative field. Stay tuned for updates on the program, speakers, and registration details.
Mark your calendars and looking forward to see you there!
@DdelAlamo Statistical Physics by Landau and Lifschitz is a standard reference. Susskind's Theoretical Minimum, already mentioned, gets good reviews. David Tong usually writes great notes: https://t.co/7ia21FIwsp
By integrating AI with extensive genetic databases, researchers are not only speeding up the identification of viable drug targets but also enhancing the precision of clinical diagnostics and predictive modeling.
Recently, @vijaypande, founding general partner of a16z Bio+Health, sat with @GSK's SVP Global Head of Artificial Intelligence and Machine Learning, Kim Branson, to discuss how AI is fundamentally changing the approach to pharmaceutical R&D.
Listen on Raising Health: https://t.co/5hFuvPWKNk
(DISCLAIMER: The above list is a personal curation that most certainly missed many key contributions (in particular the many excellent workshop & competition contributions!) and is only intended to be a starting point for your own exploration.)
I'm excited to share some of the work we've been recently doing at https://t.co/Kys8TLoGPn at the intersection of AI and precision oncology, accelerating the biopharma pipeline, and agent-based operating systems
Are you attending #NeurIPS2023?💡
I am actively recruiting for a number of open roles at different levels in my team at https://t.co/Kys8TLoGPn
If you are interested in leveraging AI to understand and advance solutions in health 🏥 please reach out! ✉️
This is a well known result in the world of protein and biological models - see rives 2021, alphafold2, etc. Data reclustering is expected and leads to dramatic improvements.
Perhaps a contrarian view but there are some major holes in the "scaled data-generation is all you need" hypothesis too:
1) observational data without perturbations in general does not enable inference of causal relationships (e.g., if you are generating unperturbed scRNAseq data you already start with a fundamental barrier)
2) if you are scaling collection of experimental controlled perturbation data, more data likely helps establish causal mechanisms - but only *in your model systems*.
3) whether (and how) those model systems are relevant to humans is mostly unclear (and from first principles and historical experience, you probably should not expect much translatability from these reductionist model systems).
4) scaled data collection assumes that you know a priori what to measure to be relevant to the mechanism underlying disease. In many of the most important diseases, the scientific community does not yet clearly know what the mechanisms are that we want to influence, nor how to measure them. Researchers frequently bias towards collecting the data they can (typically: RNAseq..) rather than what they should.
It's healthy for the community to follow multiple hypotheses - perhaps a synthesis of existing and some ideas that are yet to emerge will help close the gap between what is possible today and the (very real) promise.