Predicting protein-protein interactions in the human proteome
Predicting which human proteins shake hands—and how—is a longstanding bottleneck. Proteins rarely act alone; they assemble into complexes that drive immunity, metabolism, signaling, and disease. But testing hundreds of millions of possible pairs experimentally is slow, expensive, and blind to many weak or transient interactions.
Jing Zhang, Qian Cong, David Baker and coauthors tackle this with a smart AI + data pipeline. First, they amplify evolutionary “clues” by assembling omicMSAs—deep multiple sequence alignments mined from petabytes of raw eukaryotic genomic data—so coevolution across species pops out. Second, they train a fast interaction model, RoseTTAFold2-PPI, not just on scarce complex structures, but on domain–domain contacts distilled from ~200M AlphaFold monomers—a huge synthetic training set that teaches the network what real interfaces look like.
The payoff is big: a proteome-scale screen over ~200M human pairs yields ~18,000 PPIs at ~90% precision (and ~29k at 80%), including ~3,600 not previously reported. The method excels on transmembrane interactions, a class that’s notoriously hard in the lab, and produces 3D complex models—so you don’t just get a yes/no, you see the interface. Mapping human variants onto these models flags ~4,950 PPIs with disease mutations at the contact surface, offering concrete hypotheses for mechanism.
Beyond pairs, the team reconstructs higher-order assemblies and nominates new components for well-studied complexes (e.g., telomere maintenance, GPI-GnT, cilia/flagella machinery), and highlights GPCR partners and mitochondrial modules that have been hiding in plain sight.
Stepping back: this is a credible path toward a computed 3D human interactome—faster, cheaper, and increasingly comprehensive as more genomes and structures arrive. It doesn’t replace experiments; it prioritizes them, focusing bench time where the biology is richest.
Paper: https://t.co/IphUI7KEQT
Self-Questioning Language Models
"we propose Self-Questioning Language Models (SQLM): an asymmetric self-play framework where a proposer is given the topic and generates a question for a solver, who tries to answer it. Both the proposer and solver are trained via reinforcement learning. The proposer receives a reward if the problem is not too easy or too difficult, and the solver receives a reward based on majority voting, a proxy for correctness in the absence of ground-truth answers. For coding, the proposer can instead generate unit tests which are used for verification. We study this asymmetric self-play framework on three benchmarks: three-digit multiplication, algebra problems from the OMEGA benchmark, and programming problems from Codeforces. By continually generating more interesting problems and attempting to solve them, language models can improve on downstream benchmarks without access to any curated training datasets."
We have a long way to go on visual reasoning.
Our VisualPuzzles benchmark🧩shows similar findings, where the best models still can’t beat the bottom 5% of humans.
👉Check out our threads:
https://t.co/j7qd3y9XfI
Btw if you're learning how to build LLMs from the ground up, there's now a 17h companion video course for my LLMs From Scratch book on Manning: https://t.co/jStayj2byi
It follows the book chapter by chapter, so it works great either as a standalone or code-along resource.
It's similar to the videos I shared on YT earlier this year, but w/o ads, better navigational structure than YT.
The work was done together with Prof. Xueliang Zhu and Prof. Miao Gui. Like always, we prepared a Cover proposal, however @NatureComms does not have a cover 😅
We present DRGN-AI for fast, ab initio cryo-EM reconstruction!
* learns a neural field from unposed images,
* designed for single-shot reconstruction of unfiltered datasets,
* finds new states missed by prior approaches!
Teamwork led by @ZhongingAlong
https://t.co/WjPStRkrnI 1/
Disagreement drives metacognitive development
Opinion by Antonia Langenhoff, Bill Thompson (@billdthompson), Mahesh Srinivasan, & Jan Engelmann (@JanEngelmann5)
Free access before August 6: https://t.co/0p26CCl7T0