TLDR: I’ve started a small fund from my Base44 money that’s focused on cancer research and breakthrough therapies.
I’ve invested in 7 incredible companies in the past 4 months and plan to do more as my liquidity allows. Listing my investment thesis below
I welcome anyone like me who has been blessed financially to join me and find ways to finally manage or even end this disease.
Longer -
I honestly believe cancer is humanity's worst enemy.
~40% will be diagnosed with it, and out of which a ~third won’t make it. And numbers are currently not getting dramatically better.
For many it might sound like a statistic
But for some of us who have been there ourselves or had to go through it with a person we love, it’s one of the worst things you can imagine.
I had to go through it with my mom - who passed away 4 months ago and whom I miss and think about every single day since.
This field needs a dramatic push.
The current wave of technology and AI is doing a lot of good in the world,
But cancer therapy is one of those things that are not moving nearly fast enough.
And there’s no good reason for that - AI is creating so much opportunity right now.
Here's my thesis on how things will play out:
For drug development, the bottleneck will shift to
1. Manufacturing
2. Real world testing
"In silico" is going to accelerate incredibly fast. AI is identifying new target proteins and designing new molecules, drastically shortening timelines.
AI will create an abundance / inflation of ideas. It will help design molecules faster than anything historically possible.
This is already happening.
Play with Claude Science or Codex and you'll be surprised at how far you can get with just a few prompts.
We need a better way to *manufacture and test* the 1000s of ideas coming from AI models spitting out potential drug candidates
1. Manufacturing: i.e. how fast can we synthesize a potential drug
Not too far into the future, companies with manufacturing excellence will win over companies with the smartest scientists.
Being able to synthesize a drug candidate fast - together with efficient methods to test efficacy (more on that below), basically means you get more shots at a target hence more chances at succeeding.
This is the same exact lesson the software industry learned about the importance of fast iterations.
Companies like Starget Pharma are leveraging AI heavily to iterate on drug candidates + figured out in-house manufacturing in a way that lets them iterate on drugs almost as fast as if they were a software company iterating on features.
2. Real world testing: i.e. how fast can we get a call on a drug's efficacy.
We need better real-world indicators for drug efficacy than mouse / animal models.
The success rate of animal models -> human clinical trials is less than 8%.
Breaking down this bottleneck will change the entire industry.
Developing an oncology drug and taking it through all clinical trials usually cost >> $ 100m
Thus taking the wrong bet is devastating (but still, happens a lot).
Giving drug development companies the ability to better understand the efficacy and mechanisms of their drugs will change a lot. Esp in the age of abundant drug candidates.
Furthermore - and once those methods proved efficient - Joining forces with regulators to shorten (or even skip) some clinical trials altogether will cut years from the process and save many lives by bringing new drugs to market faster.
CuResponse is pioneering functional testing - taking real cancer biopsies and showing drug efficacy with very high precision. Can be used for precision oncology as well as testing new drugs.
Cellint is providing a fully automated cell-culture R&D platform for companies and researchers.
----
I’d love to see a future where there are “cloud services” for drug development companies.
A company should eventually be able to submit a drug candidate programmatically and receive:
1. The synthesized drug.
2. Testing across cells, tissues, organoids, and tumoroids.
3. Structured results showing what worked, what failed, and why.
4. A recommended next iteration derived from the biological reactions.
I’ve invested in all of the above companies and am searching for additional companies to bring this “cloud services for drug dev” vision to life.
----
Alongside it, I am also investing in direct therapeutic moonshots.
I invested in Baccine, which is developing a bacteria-based cancer immunotherapy platform and deserves an entire separate post.
I’m also a proud LP in Even One Ventures - @sytses ’s fund to fight cancer. Sid and Jacob Stern are an inspiration to me as I’m making my first steps in the field, and I bet their companies will make a massive impact.
And 3 other incredible companies focused on how therapies are delivered - where I’ll elaborate in future posts.
----
I’m still obviously spending most of my time on Base44. Which is a plus for the companies - as we’re able to support them with tools, software, AI agents, etc.
I’m looking for people to join me - invest side by side or help one of the companies.
I don’t have a name for the fund yet, nor an email or a website
I will get to it this weekend and will post it in the comments.
Talks at the intersection of systems engineering and computational biology
0:20 Why study systems x biology in "age of agents"
5:50 Forch: Building a utilitarian cloud container orchestrator (Max Smolin, LatchBio)
41:25 cyto: Ultra high-throughput processing of 10x Flex single-cell sequencing (Noam Teyssier, Arc Institute)
1:04:30 SLAF: A single-cell omics storage format for the virtual cell era (Pavan Ramkumar, SLAF Project)
1:33:30 Lessons in Perturbation Modeling: STATE, STACK, and Beyond (Dhruv Gautam, Arc Institute + UC Berkeley)
2:03:15 Leveraging Serverless Distributed Computing to Scale Computational Biology (Ben Shababo, Modal)
Topics span container orchestration, single-cell infra, perturbation modeling for biology at scale.
Introducing Storage Buckets on Hugging Face 🧑🚀
The first new repo type on the Hub in 4 years: S3-like object storage, mutable, non-versioned, built on Xet deduplication.
- Starting at $8/TB/mo. That's 3x cheaper than S3.
You (and your coding agents) need somewhere to dump checkpoints, logs, and artifacts. Now they have a home.
Hosting another computing x biology reading group with Modal. Progress has really picked up the past 6 months + many interesting projects to highlight.
- Max Smolin (LatchBio): Building "Forch", a Utilitarian Cloud Container Orchestrator
- Noam Teyssier (Arc Institute): cyto: ultra high-throughput processing of 10x-flex single cell sequencing
- Pavan Ramkumar (SLAF Project): SLAF: A single-cell omics storage format for the virtual cell era
- Dhruv Gautam (Arc Institute): Lessons in Perturbation Modeling: STATE, STACK, and Beyond
- Ben Shabobo (Modal): Leveraging Serverless Distributed Computing to Scale Computational Biology
Come join us for pizza and good technical talks on March 4th in Mission Bay, SF. Design decisions, paper highlights + snippets of source code.
in case you missed it @lancedb and HF are partnering up to unlock the next generation of large dataset storage on the Hub 🔥
And it's fire!
- Supports storing embeddings (and their indexes) directly alongside the data
- Vector search / similarity search is built-in
- Large multimodal datasets (text, images, video)
just use the hf:// prefix:
db = lancedb. connect("hf://datasets/julien-c/hub-stats-lance")
🔥🔥
1/3 Lance ❤️ @huggingface 🤗
We’re excited to announce native support for Lance on the Hugging Face Hub!
You can now share your large multimodal datasets with the world as a single, searchable artifact (including blobs, embeddings and indexes) all in one place.
@SutterHealth You continue to send me ER physician's medical bills after a family member's visit in July 2021. Here are 4 key oversights on your part. Figure these out before your billing department sends me another bill
@SutterHealth (1) Patient is not uninsured. Make a good faith effort to reach out to the insurer first. Details have been provided at the time of visit.
Datasets in scaled biology are richly hierarchical. Predictive ML methods can misattribute confounding in the hierarchy to effects of interest. A new method for credit assignment by Alex Rogozhnikov informs experimental design interventions each week at https://t.co/SCYgE354M4
We made genetic demultiplexing more data efficient, more accurate, as well as more compute and memory efficient to scale up droplet based single-cell RNAseq to multi-donor, multi-batch settings. Led by our very own Alex Rogozhnikov!