The "Single-cell foundation models don't work for perturbation prediction" narrative was a measurement problem.
Our new paper is out in @NatureBiotech : "Deep Learning Perturbation Models Can Outperform Baselines on Calibrated Metrics"
Last year, several papers concluded that deep learning perturbation models couldn't beat simple baselines, a finding that cast doubt on the entire Virtual Cell direction.
We asked a different question: can the metrics used in those comparisons reliably tell a good prediction from a bad one?
Often, they can't.
Common metrics like MSE and control-referenced Pearson Δ frequently rank uninformative baselines above positive controls. With calibrated alternatives (eg, WMSE, weighted R²Δ, NIR), deep learning models do outperform mean, control, and linear baselines.
The data matters too. Norman19 is approximately 96% additive, and models see less than 1% of possible gene combinations during training. The ceiling for beating an additive baseline is nearly zero, and the training signal for non-additive interactions is almost absent. Some dataset rankings aren't model rankings, but just noise!
To address this, we developed the Dynamic Range Fraction (DRF): a measure of metric calibration. Before trusting a model ranking, verify that your benchmark can detect a useful prediction at all.
The models aren't shiny toys. The field just needed better rulers.
Paper: https://t.co/Pug5bjPrvj
Code: https://t.co/4IjsHP6Wmd
Huge congratulations to the first authors @Henrymiller2012 and Gabriel Mejia, and the team of Shift Bio!
A few thoughts on Anthropic’s ART discovery, since it’s close to the work we do:
The biology is intriguing. In phages, an unusual reverse transcriptase sits next to a partner gene and an array of DNA repeats. The team found that the array produces distinct short RNAs.
It’s tempting to compare this with CRISPR or retrons, but we don’t yet know what ART does. Whether the RNAs guide anything, or whether the system is programmable, remains an open question.
I’m just as interested in how they found it. About 950 Claude agent sessions ran over 21 hours, surveyed ~200,000 reverse transcriptases, and produced 19 reports. One agent looked at the DNA next to an RT and spotted a repeat array outside the features the search had set out to examine.
It followed that observation, compared it with known systems, checked the literature, and gave the scientists something concrete to investigate.
To me, that’s the opportunity: scaling scientific attention. We have more biological data than any team can inspect closely. Agents can help us notice the odd cases and show their work.
There’s still reason to be careful. Repeat-rich regions in assembled phage genomes can be tricky, so I’d want to see raw read support across the arrays and their boundaries in independent isolates. And the most important experiment is still ahead: what does ART actually do?
I think discovery will become a tighter loop. Agents propose hypotheses, predictive models help choose experiments, and automated labs generate results that inform the next round.
We’re working toward that at @Xaira_Thera, bringing agentic workflows together with models such as X-Cell and wet lab experiments. I’m excited by how much more biology we may be able to explore when each experiment helps us decide what to test next.
Dr. Heidi Overton is about to face a nomination hearing to be the next FDA commissioner.
Former agency head @ScottGottliebMD says Overton has a chance to rebuild the FDA following a big round of job cuts. https://t.co/jcQEo94guo
Former FDA commissioners Mark B. McClellan and my colleague @ScottGottliebMD: Heidi Overton has an opportunity to strengthen the FDA
https://t.co/qFNERJxwe7 via @statnews
AEI Senior Fellow @ScottGottliebMD joined @FaceTheNation to discuss Dr. Heidi Overton's FDA nomination: why her clinical and policy background could make her strong in the job, and the challenges she'd take on at the agency.
AEI Senior Fellow @ScottGottliebMD joined @FaceTheNation to discuss the Moderna-Merck melanoma results — how this marks an inflection point in cancer, and why it shouldn't be called a vaccine.
AEI senior fellow @ScottGottliebMD joins CNBC's ‘Squawk Box’ to discuss how President Trump's vaccine executive order could put pressure on states that require proof of vaccination status for school entrance to limit or curtail those requirements.
🧵More vaccines do not mean more distinct antigens. The childhood vaccine schedule contained 3,217 antigenic components in 1960 and 3,041 by the 1980s. Today, owing to better formulations, it's about 165 antigens across the current pediatric schedule 1/2
https://t.co/xc06FHWrtb
@ScottGottliebMD AEI senior fellow @ScottGottliebMD joins CNBC's ‘Squawk Box’ to discuss concerns over food safety and FDA's handing of the recent Cyclospora outbreak.
"I think the FDA has done a good job isolating the sources of these outbreaks," says Fmr FDA Commissioner @ScottGottliebMD of cyclospora. "It does call into question whether or not we should be increasing our oversight into those growing fields in Mexico." https://t.co/rs46A3gVo7
Today’s schedule protects against more diseases than in the 1960s or 1980s, with 95% fewer antigens.
Antigen count doesn't predict every reaction, but it's more relevant than shot count to the concern that a larger schedule asks a child's immune system to recognize more targets
🧵More vaccines do not mean more distinct antigens. The childhood vaccine schedule contained 3,217 antigenic components in 1960 and 3,041 by the 1980s. Today, owing to better formulations, it's about 165 antigens across the current pediatric schedule 1/2
https://t.co/xc06FHWrtb
📝This is an important reform. FDA has generally lacked sufficient authority to meaningfully regulate these substances; and that lack of authority has stymied similar actions in the past. In this proposed rule, the agency appears to have gone as far as it can, without new Congressional action, to provide greater oversight of food ingredients that are “Generally Recognized as Safe" (GRAS).
The agency is relying on a new interpretation of existing statute, backing away from its prior finding from 2016 that it lacked express statutory authority to mandate GRAS notices. It's a 140+ page rule that has been carefully drafted (which explains the delay in its promulgation). Nonetheless, it’s likely this rule will be challenged in court if finalized.
Ultimately, the FDA needs a new set of authorities and more resources from Congress. The money to administer this new framework could be a bottleneck. This new rule requires notification, not pre-approval. But consumers will expect the FDA to review what’s in the notifications, and the FDA will try to do that. But the agency may not be properly resourced for that new mission.
https://t.co/MuruHIhZAY
Tune in at 2:30pm East as @DrMarcSiegel catches up on the latest - including recently FDA-approved vaccine, mFlusiva - with Scott Gottlieb, MD @ScottGottliebMD - Senior Fellow at @AEI & former Commissioner of the @US_FDA. Stream here: https://t.co/bw9glj0Hts.
2/2 FDA’s Food Safety Partnership with Mexican regulators supports prevention, surveillance, traceability, and outbreak response. That cooperation is vital where security concerns can constrain direct FDA inspections. These outbreaks warrant a fresh look at that oversight model.
1/2 This is a large number of cases over a short period, and the count may still rise. It's at least the 2nd major outbreak this year linked to produce from a distinct Mexican growing region—jalapeños from Sinaloa after iceberg lettuce from central Mexico
https://t.co/GXvKk1UxY1
Over the weekend, AEI senior fellow and former FDA commissioner @ScottGottliebMD joined CBS News’ Face the Nation to discuss the federal response to the cyclospora outbreak, rising measles cases amid declining vaccination rates, and the risks of compounded peptides.