Antibody LMs learn what looks antibody-like, but not how selection turns naive germline antibodies into strong binders.
@aakarshv1 and I are excited to share CoSiNE, a model that learns this germline-to-mature process for variant effect prediction and antibody design. (1/8)
OpenBind intends to collect 10,000s of protein-ligand structures & affinities. To prioritize what we collect next, we need cofolding models trained on the latest data. Today we're releasing OpenBind-0 and 717 new ligand-bound structures.
i disagree with Dario's medical takes, but i think there's kind of a failure of imagination with the responses to him. there is, i think, an okay steelman for “ai will figure out solutions to a good chunk of diseases” if you believe superintelligence can solve the delivery problem entirely. that is: it will be able to design drugs that target arbitrary places in the body, and nowhere else.
lots of medicine is bottlenecked by this problem!
let’s suppose we had such a machine, which, given a particular biomolecule, tells us which tissues it reaches and nothing else.
with such a machine, you can...
1. cure a fairly large chunk of rare diseases, which are mostly monogenic mutations that are solvable with current genetic-engineering technology, given good-enough delivery vehicles produced by the machine. yes, you won’t undo all the damage caused by a lifetime of disease, but still!
2. put a pretty big dent in—if not outright eliminate—any solid tumor. how? just go all in on radionuclides that, magically, target only cancer cells. the radiation burden eventually kills it without harming the patient (much). nearly-pan-cancer cure; none of the usual “cancer is actually many diseases” nuance needed. you might say that some cancers are a problem because they’re discovered too late, rather than because they’re hard to cure, but solving the delivery problem helps with that too. highly selective PET ligands would let us regularly and non-invasively monitor for small lesions, then treat them with with matched radioligands whenever something suspicious turns up.
3. cure at least some severe autoimmune diseases. we already know that wiping out B-cell lineages via cell therapy can do that, but it’s too toxic to scale. with our magic delivery-solver ai, we could make therapies that remove the particular autoreactive B-cell and plasma-cell populations we dislike, leaving most useful immunity intact. and for some autoimmune conditions, the relevant autoantigens and autoantibodies are well established!
there: three broad classes of diseases that get tidied up with access to a superintelligence that can solve one singular problem. and you probably get more than this too!
of course, much like Dario, i’m doing a little handwaving here myself. truly solving the “delivery problem”—beyond the basic ability of distinguishing one cell from another—involves solving a dozen bundled problems: blood-brain delivery, preventing an immune response, avoiding liver and spleen sequestration, getting through the cell membrane and then escaping the endosome
but i dont know. is it really so unachievable? the delivery problem feels like a pretty legible issue, no secret knowledge of biology required beyond more data collection from the obvious sources + mulling over the results. if we really put our minds to it—and many people are—it seems well within the realm of possibility. it's also not terribly slow to run this through clinical trials; the tech stack for genetic editors, radionuclides, and immune depletion does already kinda exist.
now, i think it's fair to still say solving the delivery problem is somewhere on the spectrum of impossible (too hard to distinguish stuff, the biological knobs just dont exist) to intractable (there actually is secret knowledge about biology we need to uncover before we solve it). im somewhere in these two camps personally, but i wouldn't find it insanely shocking if i turn out to be wrong
Most PDB structures show one frozen conformation. But proteins are not static.
New work from @stephanie_mul at @radialscience's @diffUSEproject reprocessed ~80,000 high-res PDB structures with qFit, recovering hidden conformational heterogeneity. Result: 60,000+ multiconformer models, the largest experimentally-derived ensemble dataset to date. Better fit to the data (lower R-free) in ~90% of cases.
MD-based ensemble predictors are capped by simulation time and force-field accuracy. This dataset pulls real ensemble signal straight out of X-ray/cryo-EM data instead.
Paper, code, data:
https://t.co/S9zjfTNB1V
Excited to share MarinDNA, a 1B gLM that rivals Evo 2 40B on variant effect prediction while being 2,330x faster.
With @eczech0, we built around a standard Transformer so we could reuse LLM infra and methods while focusing on data curation and scaling.
https://t.co/TVChdjJyMP 🧵
@bjing2016 hey bowen! would it be possible to also include the scaling law analysis in the promera paper for the subset of models that have reported data and compute usage?
It’s getting hard to keep track of all the new open-source cofolding models coming out. To help everyone stay up to date, I'm excited to share https://t.co/Gc9PaF6GoY, a PDB-synced leaderboard of open models, updated weekly. All predictions are available to view & download! (1/4)
Can't believe this batch of PhD students gets to go to Seoul, Vienna, Singapore, Kyoto, San Diego, Hawaii
while the previous batch visited Gather Town, and the one before that apparently just went to New Orleans multiple times
I am thrilled to share that UC Berkeley and UCSF have launched a joint initiative in Computational Biomedicine!
https://t.co/X8WMamp4HL
We will soon be recruiting new faculty and postdoctoral fellows. Please repost to help spread the word.
Antibody LMs learn what looks antibody-like, but not how selection turns naive germline antibodies into strong binders.
@aakarshv1 and I are excited to share CoSiNE, a model that learns this germline-to-mature process for variant effect prediction and antibody design. (1/8)
Antibody LMs learn what looks antibody-like, but not how selection turns naive germline antibodies into strong binders.
@aakarshv1 and I are excited to share CoSiNE, a model that learns this germline-to-mature process for variant effect prediction and antibody design. (1/8)
Affinity maturation is how naive antibodies evolve into strong binders, but most antibody LMs ignore it.
@stephenzlu and I built CoSiNE to learn this, beating antibody LMs on VEP and reframing design as guiding evolution, not de novo generation.
Excited to present at ICML!
The resulting samples remain structurally plausible and human-like.
We think this is a promising step toward controllable evolutionary protein design: not just generating sequences de novo, but guiding the processes that produce them. (7/8)