PhD in Cell and Molecular Biology
BSc in Biochemistry and Molecular Biology
-Bioinformatics / computational biology
-Let's accelerate biological research
I’m super optimistic that intelligent systems combined with intelligent humans in the loop can and already are producing useful, novel things. What I’m less sure about is whether the best discoveries will actually reach the people who need or want them most.
@scaling01 Agree. People are sleeping on using Opus to hill climb. We use it for optimizing CPU and memory, optimizing CI times, improving frame rates, reducing latency, any other kind of problem in the shape of “iterate on X with a profiler and dataset until it hits Y”
Super proud and excited to finally present LiteMol-1. This is our first foundation model from LiteFold, pre-trained from scratch.
Today, LiteMol-1 can generate small molecules, peptides, cyclic peptides, depsipeptides, peptides with ncAAs, macrocycles, and PROTACs.
Across our peptide and small-molecule evaluations, the model shows competitive results. In several settings, we are on-par with or better than frontier structure-based models, at a fraction of the generation cost.
But the part I find most interesting is that this is a model for agents. We have seen ourselves how much compute, and how many tokens it can take to generate good binders using frontier structure-based models. Sometimes you need to generate tens of thousands of designs just to get a handful worth taking forward.
Now put this inside an AutoResearch loop. The agent has to continuously parse structures, inspect PDB/CIF files, compare candidates, run evaluations, modify the design, and repeat the whole thing again. It becomes extremely expensive very quickly.
Sequence space gives us a very different interface. LLMs are much more efficient at inspecting and manipulating compact molecular representations like SMILES than repeatedly operating over full structural files.
So LiteMol-1 becomes something like an infinite molecular canvas for the agent. For a given target and objective, the model can continuously propose what a biomolecule could look like. The agent can inspect those generations, take inspiration from them, preserve certain regions, edit others, optimize them, score them, and generate again.
Generate → inspect → evaluate → edit → generate again.
There is another problem I care a lot about. Most molecule design models today are heavily optimized around binding. But binding is only one part of whether something eventually becomes a therapeutic. What about ADME? Toxicity? Selectivity? Solubility? Membrane permeability? Synthesizability?
For this, we also built a Monte Carlo Tree Search-based multi-objective generation framework around LiteMol-1. Instead of combining everything into one score, the search keeps multiple strong candidates, each balancing the desired properties in a different way.
As our scoring functions and verifiers get better, the generation system gets better too. We can start steering molecules not just toward “binds well”, but toward a broader therapeutic design specification.
The bottleneck slowly moves from simply generating molecules to having sufficiently good verifiers and scoring functions to tell us what is actually worth generating. Check out our technical research blog post for all the details.
At LiteFold, our research is focused on engineering biomolecules and building systems that can carefully forecast their pre-clinical success.
To stay updated on our research, follow LiteFold.
Cheers!
Super proud and excited to finally present LiteMol-1. This is our first foundation model from LiteFold, pre-trained from scratch.
Today, LiteMol-1 can generate small molecules, peptides, cyclic peptides, depsipeptides, peptides with ncAAs, macrocycles, and PROTACs.
Across our peptide and small-molecule evaluations, the model shows competitive results. In several settings, we are on-par with or better than frontier structure-based models, at a fraction of the generation cost.
But the part I find most interesting is that this is a model for agents. We have seen ourselves how much compute, and how many tokens it can take to generate good binders using frontier structure-based models. Sometimes you need to generate tens of thousands of designs just to get a handful worth taking forward.
Now put this inside an AutoResearch loop. The agent has to continuously parse structures, inspect PDB/CIF files, compare candidates, run evaluations, modify the design, and repeat the whole thing again. It becomes extremely expensive very quickly.
Sequence space gives us a very different interface. LLMs are much more efficient at inspecting and manipulating compact molecular representations like SMILES than repeatedly operating over full structural files.
So LiteMol-1 becomes something like an infinite molecular canvas for the agent. For a given target and objective, the model can continuously propose what a biomolecule could look like. The agent can inspect those generations, take inspiration from them, preserve certain regions, edit others, optimize them, score them, and generate again.
Generate → inspect → evaluate → edit → generate again.
There is another problem I care a lot about. Most molecule design models today are heavily optimized around binding. But binding is only one part of whether something eventually becomes a therapeutic. What about ADME? Toxicity? Selectivity? Solubility? Membrane permeability? Synthesizability?
For this, we also built a Monte Carlo Tree Search-based multi-objective generation framework around LiteMol-1. Instead of combining everything into one score, the search keeps multiple strong candidates, each balancing the desired properties in a different way.
As our scoring functions and verifiers get better, the generation system gets better too. We can start steering molecules not just toward “binds well”, but toward a broader therapeutic design specification.
The bottleneck slowly moves from simply generating molecules to having sufficiently good verifiers and scoring functions to tell us what is actually worth generating. Check out our technical research blog post for all the details.
At LiteFold, our research is focused on engineering biomolecules and building systems that can carefully forecast their pre-clinical success.
To stay updated on our research, follow LiteFold.
Cheers!
@cjmaddison It’s absurd that no one provides any concrete details on how this will be achieved within that timeframe. It’s mostly wishful thinking plus lack of solid molecular biology background.
@Ronalfa As hard as it is to cure anything, it may be just as hard to make those cures affordable and accessible to the people who need them, which is what truly matters anyway.
@SylvainGariel Agree.Also, it reminds me of the Human Genome Project in 2003. Many people thought once you sequence the genome, you would quickly unlock cures for most diseases. It led to enormous progress, but also revealed just how complex biology really is.
I haven’t gotten one word using Fable. It’s crazy the most powerful models aren’t just available because you do mol bio. I’ve had a lot of that in 5.6 Sol med recently too. It just stops responding for stupid reasons. I don’t even know if there are actual biologists that work there.