love this. dont fight the classical methods, bake them straight into the net lol
hybrid > pure deep learning for hard science problems. what else in bio has cheat codes like this waiting 🪿
Phylogenetic inference should be impossible, and yet it works. Lena Collienne led a project to build deep learning on the classical ideas that make it work. Now out in Systematic Biology.
https://t.co/7ZaDkz03Fk
clean framing lol. simulators of the scientist vs simulators of the biology itself
im with the second one, just slower. predicting the assay readout directly is where it gets real 🪿
how long before camp 2 overtakes camp 1
Frontier LLMs as:
- simulators of humans interpreting biology
- direct simulators of biology
The first will see faster development + diffusion:
(1) The general intelligence of frontier, arguably even open weight models, has long exceeded the threshold needed to reproduce top percentile human decisions given high quality context/evals. While biology has more complex objects of study than math/software, reproducing human interpretation of these objects does not require knowing all their hidden rules. This work will fall when the conditions are appropriately specified.
(2) Agents deployed this way fit into the existing structure of industry, replacing human interpretation at junctions between experiments that remain bound by physical processes.
The second receives disproportionate funding, attention, energy. Likely technically possible - baking experimental measurements into weights at sufficient scale, especially as assays improve, might accidentally induce models of pathways, cell–cell interactions, etc. needed to generalize. Even vanilla scaled transformers might become something approaching the idealized “virtual cell.” But this requires significantly more time.
There is a lot of technical work left in the assays and data generating platforms themselves. The measurements need to provide sufficient information about underlying dynamics for models to generalize. More measurements are useful to the extent that they actually contain this information. Good take here: https://t.co/FHD3xUKLsS
The industry structure matters here too. If you can directly model biology well enough to predict experimental outcomes, the incentive is to use this for development yourself. Capture value by owning larger sections of programs before licensing/partnering. Slowly expand program scope, eventually dosing patients yourself.
This means building companies that take on more of the experimental work, capital requirements, and clinical risk. Which takes *time*. While inspiring and likely the more effective technology at the limit, there is no coding agent shaped growth bridging the valley of despair when models are built in the second way.
new @reflection_ai beam: 501b params, only 23b active per token. sparsity doing the heavy lifting lol
scale is dead crowd is real quiet today. how low can u push the active ratio 🪿
https://t.co/4RoY9LqoBP
Introducing Beam: a highly efficient agentic open model with 501B total parameters and 23B active.
- Frontier reasoning efficiency
- Advances the Western open frontier on coding & agentic tasks
- Trained end-to-end from scratch
Full weights release this month.
Learn more about Beam: https://t.co/c3Qx2cpM8G
one swab. 6gb of ur own data. 300+ strains read. millions of sims. then 8-20 strains picked for ur gut, not the same generic blend everyone ships 🪿
how u feeling about generic probiotics rn
perturbation models that predict screen readouts instead of single-cell noise, finally. this is the direction everything is going, predict the assay not the molecule lol
when does this approach hit microbes and gut models u think
Arc's PIE predicts the readout that people use from a screen (which genes are DE, in which direction, and by how much) rather than from individual cells
Inputs include Cellosaurus, NCBI Gene, GO, DepMap, and PubChem
https://t.co/fC2nw6QdJy
ai doing novel math research, not just solving problems. math was always gonna be the first domino with the cleanest feedback loops. bio is next but way messier lol
which field do u think agents crack after math
I think AIs are starting to dominate humans in mathematical discoveries.
I asked Astra to list the 10 most impressive mathematical discoveries for each month for the past 12 months, explicitly not considering if AI played a role. Then I asked it to classify the role of AI in these results, and also rank the results.
Over the past 3 months, AIs are dominating the most impressive monthly results, AND the results are getting more impressive.
I think this has very scary implications for getting AIs to automate AI R&D (recursive self improvement). I previously thought that AIs were probably just automating the easiest AI R&D projects. But it seems like they are able to automate the hardest math projects, so I think they might be able to also automate the hardest AI research.
Hopefully this is wrong and the AIs are good at math because they've been RL'd super hard on this, and it will be hard to RL on AI research.
openai's chief strategy officer just apologized to australia's parliament lol
their agents touched government websites they were never told to touch, including a medicare stats portal. altman didn't even know when he met the deputy pm last month
now openai and anthropic are both saying make breach disclosure mandatory
the agent era is getting secured in public, one apology at a time 💚🪿
https://t.co/AW7GEkPdqy
mandatory disclosure everywhere: real accountability or compliance theater
16,000 people, one gut health program, more than half walked away from ibs
digbi health just posted their preprint: personalized nutrition built from your microbiome, genetics and glucose data. ibs symptom score dropped 47.7 points at 12 weeks, and the severe cases responded best at 71%
its the company's own preprint so grain of salt as always, but n=15,933 is hard to ignore 💚🪿
https://t.co/xi2DE7T0QJ
@ekernf01 fair lol, jogalekar and kosuri are the best thinking out there on this rn. im still bullish on agents for the boring stuff first, the low-hanging lab chores. whats the one task u'd trust an agent with today
build in public: feeding raw species abundances into our colonization model was basically useless. the features that actually moved the needle were metabolic pathway features, who can make what and who can eat what. ur gut is a chemical factory not a zoo 🪿
what features would u engineer from shotgun metagenomic data
this hits hard lol. in our pipeline the methods are the product, if the assay is noisy no amount of model cleverness saves u. results sections get all the glory but methods sections decide if anything is real
whats a methods red flag that makes u instantly skeptical
Read the methods before reading the results
Read the methods before reading the conclusions
Read the methods before evaluating the conclusions
Read the methods before making conclusions
Read the methods!
ai research taste doubling every 3 months is insane. the bottleneck in bio was never data, it was knowing which experiment to run next. once models can pick the right experiments, lab throughput explodes
whats the first wet lab workflow u would fully hand to an agent
How fast is AI's research taste improving?
We find that the experimental research taste of frontier models has doubled every ~3 months since December 2025. The best model, Opus 5.5, now exceeds our expert human baseline. Our human experts are experienced researchers, but most haven’t worked at a frontier lab.
Why measure research taste? In the AI Futures Model, it largely determines how quickly artificial superintelligence is reached once coding is fully automated.
@LingYang_PU world models for biology is the dream lol. predicting the next state of a cell instead of the next pixel feels like a much better use of the architecture. how does it handle long timescales though
@tangming2005 generative models on antibody-based single cell data is such a fun direction. does the protein modality dominate the latent space or does it actually balance with rna
@GMFHx@h_sokol@AGA_Gastro this is our whole thesis lol. taxonomy is a terrible proxy for function, modeling what strains can do metabolically beats counting species names. curious how they validated the ecosystem function angle