Simona Cristea

Verified account

@simocristea

director of applied AI @TempusAI; prev: faculty @DanaFarber, group leader @Harvard & phd @eth

Boston 🇺🇸 & Zurich🇨🇭

Joined January 2016

451 Following

9.5K Followers

4.6K Posts

Pinned Tweet

11 months ago

scRNAseq cell type annotation is notoriously messy. Despite so many algorithms, most researchers still rely on manual annotations using marker genes In a new preprint accepted at ICML GenAI Bio Workshop, we ask if reasoning LLMs (DeepSeek-R1) can help with cell type annotation🧵

simocristea's tweet photo. scRNAseq cell type annotation is notoriously messy. Despite so many algorithms, most researchers still rely on manual annotations using marker genes

In a new preprint accepted at ICML GenAI Bio Workshop, we ask if reasoning LLMs (DeepSeek-R1) can help with cell type annotation🧵 https://t.co/dZp926YFx6

7

199

38

145

26K

2 days ago

@venkmurthy @marklewismd effect sizes do correlate with pvalues though 😄 in this case, we’re really safe to ditch the stats. best data is when nobody carea about the stats

0

1

0

0

37

2 days ago

@mmbronstein the culmination of decades of work

0

2

0

1

986

simocristea retweeted

2 days ago

Big progress vs cancer, folks. The kind of event curves from randomized trials that we've not seen before for a couple of the most deadly cancers. Congrats to the oncology research community for getting these trial done. #ASCO26, @ASCO

EricTopol's tweet photo. Big progress vs cancer, folks.
The kind of event curves from randomized trials that we've not seen before for a couple of the most deadly cancers. Congrats to the oncology research community for getting these trial done. #ASCO26, @ASCO https://t.co/8a642gUMF4

36

2K

481

253

111K

Who to follow

Ming "Tommy" Tang

Director of bioinformatics at AstraZeneca. YouTube at chatomics. On my way to helping 1 million people learn bioinformatics. Also talks about leadership.

Irving Associate Prof. @Columbia @ColumbiaBME @Cancer_dynamics #computationalbiology #AI #machinelearning #genomics #cancerimmunology ➡️ @elhamazizi.bsky.social

Verified account

Professor at the NYU School of Medicine. Co-host of the 'Night Science Podcast' @nightsciencepod and Co-founder of the Night Science Institute https://t.co/lede3rBmXU

3 days ago

@oakazaki OS >> PFS better imo

0

1

0

1

405

3 days ago

@xlr8harder they'll all open bakeries

0

2

0

0

366

3 days ago

hard to internaliza this now because we are so attuned to the present, but he is right

4 days ago

@t_blom This problem will naturally tend to go away as companies are grown from the start using AI. Then you don't need to extract any domain knowledge from people's heads; it will never have been in people's heads.

104

2K

70

379

218K

0

1

0

0

563

3 days ago

@marklewismd not often that we can just ditch the stats

0

0

0

0

1K

3 days ago

there aren't many times in oncology when nobody cares about statistics, but today is one of them. there has never been such a successful trial in pancreatic cancer & these survival curves are the result of 40 years of persistence. KRAS inhibitors will forever transform oncology

Dr. Antonio Calles 🫁🚭

3 days ago

🌟This is history ⭐️The most awaited abstract 👏 Standing ovation at Hall B1 💊 Daraxonrasib becomes the new standard of care for patients with previously treated metastatic #pancreatic #cancer #ASCO26

1

51

15

6

14K

1

98

10

5

6K

simocristea retweeted

Adam Feuerstein ✡️

@adamfeuerstein

3 days ago

Incredible #ASCO26 moment. Dr. Brian Wolpin, presenter of the daraxonrasib study, received a standing ovation DURING his talk after he stated the survival benefit for PDAC patients. It was sustained. Cheering. I have never see anything like it in the middle of a talk. $RVMD

adamfeuerstein's tweet photo. Incredible #ASCO26 moment.

Dr. Brian Wolpin, presenter of the daraxonrasib study, received a standing ovation DURING his talk after he stated the survival benefit for PDAC patients. It was sustained. Cheering. I have never see anything like it in the middle of a talk. $RVMD https://t.co/I2wQfDqsvh

16

977

136

96

143K

27 days ago

@arcinstitute congrats! your expansion is really impressive

0

0

0

0

376

28 days ago

alphaevolve has the highest potential to transform science as a whole, across fields: bio, materials, psychology etc

Google DeepMind @GoogleDeepMind

28 days ago

Algorithms are part of nearly every aspect of life, from the physics of the natural world to planning shipping routes. Our Gemini-powered coding agent AlphaEvolve has been accelerating progress over the last year - from quantum and biotechnology to logistics and @Google’s AI infrastructure. ↓ https://t.co/CAjvAqJiod

84

1K

226

344

188K

0

53

10

22

8K

28 days ago

@GamerPhilDoc @NatRevDrugDisc 33 is not that bad for cancer vaccines though.. phase 1 is exploration mostly

0

0

0

0

131

28 days ago

wow as of may 2025, there are 513 cancer vaccines in development, with 33 in phase 3 @NatRevDrugDisc

simocristea's tweet photo. wow as of may 2025, there are 513 cancer vaccines in development, with 33 in phase 3 @NatRevDrugDisc https://t.co/le0VNBloRU

4

123

31

47

13K

28 days ago

https://t.co/UzLOfinDMx

0

2

1

2

374

about 1 month ago

0

1

0

0

40

about 1 month ago

@SashaGusevPosts purely tragic

0

1

0

0

109

about 1 month ago

this is the first truly impressive comp bio AI-only analysis that I’ve seen. this is truly useful

Derya Unutmaz, MD

about 1 month ago

As I mentioned before, I am now sharing an example from GPT-5.5 Pro, also featured by OpenAI, that really left me stunned by what it is capable of in biomedical science. (full report on the website I created with Codex, link in the thread). To push GPT-5.5 Pro hard, I uploaded a real data set of immune subset (T cells) gene-expression spreadsheet: 62 sorted T cell samples, 27,906 gene columns, and millions of underlying data points across different T cell subsets. Importantly, this public dataset also had paired structure making it possible to separate true cell-state biology from donor-to-donor variation. I asked GPT-5.5 Pro not merely to summarize the spreadsheet, but to analyze it deeply: What can we learn from this dataset? What are the mechanistic insights? What are the most important biological questions that emerge? What follow-up experiments should we do next? It thought for about 100 minutes and produced a roughly 40-page report! What amazed me was not just the length or even the initial analysis, since previous models are also capable of doing this. What amazed me was the quality of the reasoning and insights it provided! The report recognized that this was not just a table of genes, but two overlapping experimental designs. It identified the major biological axis, which in plain language was that the cells were not just “different categories.” They formed a coherent differentiation landscape, moving from future potential toward immediate function. It also understood the caveats. It did not overclaim from bulk gene-expression data. It clearly explained that bulk transcriptomics cannot distinguish whether every cell in a sorted population has shifted or whether a smaller subpopulation is dominating the signal. It recommended the right next steps experiments, and integration with donor metadata. This is what made the report feel so special to me. It was not just doing statistics. It was reasoning like an expert systems immunologist. It saw the structure of the experiment, interpreted the patterns, built a mechanistic model, identified limitations, proposed causal hypotheses, and laid out a translational roadmap. Other advanced models have been able to generate excellent biomedical reports before, including previous GPT-5 models. So I don't want to claim this is an entirely new type of capability. But this one felt different in an important way. It had more scientific elegance, more restraint, more biological intuition, and more of the nuanced judgment that usually comes only from years of hands-on experience in the field. It felt like this AI model had crossed another threshold. This is the kind of analysis that could easily take a research team months to perform, refine, interpret, and write up. Even then, many teams might not produce something this integrated, this mechanistically coherent, and this useful as a launchpad for future experiments. I know a 40-page T-cell gene-expression analysis may not be exciting to everyone. To illustrate how good it is, also had Codex built a web site with it anyone can explore, link below. 😊 Those interested can go deeper into the report. I also wanted this example on the record because, because to me, it is evidence that we are entering a new stage in AI-assisted biomedical science. The important point is no longer that AI can "analyze data and write a report.” The important point is that AI can now help transform complex biological data into mechanistic understanding, experimental priorities, and testable hypotheses at a speed and depth that would have been almost unimaginable a short time ago. For biomedical science, this is a very big deal! Of course, this may vary across domains, and every analysis still needs expert review, validation, and experimental follow-up. But in my own field, with data I understand deeply, this felt like another inflection point. I feel strongly that we have crossed another milestone threshold in the age of AI, with the release of GPT-5.5.

DeryaTR_'s tweet photo. As I mentioned before, I am now sharing an example from GPT-5.5 Pro, also featured by OpenAI, that really left me stunned by what it is capable of in biomedical science. (full report on the website I created with Codex, link in the thread).

To push GPT-5.5 Pro hard, I uploaded a real data set of immune subset (T cells) gene-expression spreadsheet: 62 sorted T cell samples, 27,906 gene columns, and millions of underlying data points across different T cell subsets. Importantly, this public dataset also had paired structure making it possible to separate true cell-state biology from donor-to-donor variation.

I asked GPT-5.5 Pro not merely to summarize the spreadsheet, but to analyze it deeply: What can we learn from this dataset? What are the mechanistic insights? What are the most important biological questions that emerge? What follow-up experiments should we do next?

It thought for about 100 minutes and produced a roughly 40-page report!

What amazed me was not just the length or even the initial analysis, since previous models are also capable of doing this. What amazed me was the quality of the reasoning and insights it provided!

The report recognized that this was not just a table of genes, but two overlapping experimental designs. It identified the major biological axis, which in plain language was that the cells were not just “different categories.” They formed a coherent differentiation landscape, moving from future potential toward immediate function.

It also understood the caveats. It did not overclaim from bulk gene-expression data. It clearly explained that bulk transcriptomics cannot distinguish whether every cell in a sorted population has shifted or whether a smaller subpopulation is dominating the signal. It recommended the right next steps experiments, and integration with donor metadata.

This is what made the report feel so special to me. It was not just doing statistics. It was reasoning like an expert systems immunologist. It saw the structure of the experiment, interpreted the patterns, built a mechanistic model, identified limitations, proposed causal hypotheses, and laid out a translational roadmap.

Other advanced models have been able to generate excellent biomedical reports before, including previous GPT-5 models. So I don't want to claim this is an entirely new type of capability. But this one felt different in an important way. It had more scientific elegance, more restraint, more biological intuition, and more of the nuanced judgment that usually comes only from years of hands-on experience in the field.

It felt like this AI model had crossed another threshold.

This is the kind of analysis that could easily take a research team months to perform, refine, interpret, and write up. Even then, many teams might not produce something this integrated, this mechanistically coherent, and this useful as a launchpad for future experiments.

I know a 40-page T-cell gene-expression analysis may not be exciting to everyone. To illustrate how good it is, also had Codex built a web site with it anyone can explore, link below. 😊 Those interested can go deeper into the report.

I also wanted this example on the record because, because to me, it is evidence that we are entering a new stage in AI-assisted biomedical science.

The important point is no longer that AI can "analyze data and write a report.” The important point is that AI can now help transform complex biological data into mechanistic understanding, experimental priorities, and testable hypotheses at a speed and depth that would have been almost unimaginable a short time ago.

For biomedical science, this is a very big deal!

Of course, this may vary across domains, and every analysis still needs expert review, validation, and experimental follow-up. But in my own field, with data I understand deeply, this felt like another inflection point.

I feel strongly that we have crossed another milestone threshold in the age of AI, with the release of GPT-5.5.

22

475

69

327

49K

2

46

3

47

11K

simocristea retweeted

Andrew Gordon Wilson

about 1 month ago

There's a fourth possibility: humans only appear sample efficient because they've effectively seen a massive amount of data through evolution. Remember, there is a fluidity between the model and the data. The model is a representation of our understanding of data.

55

437

32

128

45K

about 1 month ago

@SashaGusevPosts @anshulkundaje many (human) reviewers also do this though; “add another mouse experiment”; “how about another batch correction method” etc

0

0

0

0

91

Last Seen Users on Sotwe

Trends for you

Most Popular Users