Congrats to insitrocytes Kirill Shmilovich, Benson Chen, @Tkaraletsos & @mmsltn for the publication of their recent paper in The Journal of Chemical Information and Modeling! Read about insitro’s progress in modeling DNA-encoded library (DEL) datasets: https://t.co/GhxJSwivdY
We can finally talk about it:
We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company.
We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
Announcing Discovery Loop!
I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor.
♾
Learn more at: https://t.co/Rv3LMdLluK
Also, I suspect you could train a great ensemble distribution via this by keeping X models per time step and pruning the X-N worst ones via some scaled loss weighting.
Such a great idea! Probably a low hanging fruit to apply this to small-mol/protein diffusion to train/fine-tune them better for things like allosteric binding/regressed sites or mol generators for multiple conformations.
We discovered a third pretraining axis beyond parameters and data: exploration.
Scaling exploration monotonically improves existing models across images/video/language, and unlocks end-to-end generation.
In the simplest case, it's just a for loop.
Introducing Explorative Modeling.
TLDR:
- Gains from exploration grow with scale: 7%→36% as data scales, 13%→23% as parameters scale, and gains double at 3× the compute
- Adding exploration to ~SOTA baselines improves data efficiency by 6.2×, FLOP efficiency by 4.1×, parameter efficiency by 47%, and hits a near-SOTA 1.43 unguided FID on ImageNet
- Exploration lets you trade training compute for generalization, and scales how end-to-end your generative model is
- End-to-end Explorative Models (XMs) match diffusion performance on control tasks with up to 256× less inference compute
🧵Thread:
The abuse vectors for F1 are not related to length of stay. F1 needs to be structured towards encouraging of filling technical gaps rather than arbitrary hardships and blanket reductions.
During my PhD I had to go home midway through to get a new visa stamp . It was both stressful and a pretty big expense for me at the time. Really wish the administration would reduce the barriers for talent coming to 🇺🇸 instead of randomly increasing it.
The Trump administration will limit how long a foreign student can remain in the U.S. on a student visa under a new regulation announced on Thursday, increasing the hurdles for foreign students and universities’ ability to recruit them.
Read more: https://t.co/k3X6Qj6I2f
Today we're sharing new breakthrough results for Pearl, our foundation model for protein–ligand cofolding.
The OpenBind Consortium recently released the first public structure-affinity benchmark for molecular AI, evaluating six prominent cofolding models on the EV-A71 2A protease. We ran our full Pearl system against the same target.
Zero-shot, with no binding-site information and no tuning, the Pearl system reaches 78% on OpenBind's primary success criteria, far ahead of every cofolding model tested by OpenBind. We also assessed a stricter sub-1 Å accuracy threshold, which is more relevant for real-world R&D usage – the Pearl system’s success is still 60%, versus 1–27% for the other models.
What matters most to us: this is the same system setup our scientists use on live drug discovery programs, not a benchmark-specific configuration.
Thanks to the OpenBind Consortium for building a rigorous public benchmark, and to @NVIDIAHealth for the support on optimizations that enabled model scaling.
I remain thankful for all the opportunities that 🇺🇸 gave me as an immigrant and fully believe that 🇺🇸 is the greatest country on earth but we should not make it harder for immigrants that want to contribute to it to stay here with unnecessary rules.
An alien who is in the U.S. temporarily and wants a Green Card must return to their home country to apply.
This policy allows our immigration system to function as the law intended instead of incentivizing loopholes.
The era of abusing our nation’s immigration system is over.
For context, I was in the US for ~12 years on F1 before I got my GC and it took another 5 for my 🇺🇸citizenship. Having no path to switch status without leaving for a long time would have been incredibly disruptive for my research and career.
Attention @arxiv authors: Our Code of Conduct states that by signing your name as an author of a paper, each author takes full responsibility for all its contents, irrespective of how the contents were generated. 1/
Enhanced Diffusion Sampling: We develop a framework for efficient rare event sampling and free energy calculation with diffusion models. We introduce Metadynamics and Umbrella Sampling for diffusion models.
@MSFTResearch#MachineLearning#MD#Biology
https://t.co/NC6QNrCP85
@zavaindar Agree on most things but disagree that the US should be ceding ground on follow-on molecules. Almost all things in the clinic have some issue that 2nd gen+ solve and giving up on them all together is leaving trillions on the table.
@DdelAlamo Fair, some of it is probably also that the pre training objective of PLMs, the reconstruction metric for IF models and the parametrized physics functional forms have little to do with the properties of ABs like aggregation. We need better/different pre training recipes for them.
I do wonder if the fastest way for US biotech to get regulatory parity to China is for US regulators to require all international/ex-China rights to explicitly require Taiwan.😁
For folks in preclinical research, including many friends trying to raise, going through RIFs or trying to find new roles, I see you and wish you all the best navigating this multi year biotech bear market and the turbulence ahead. In the end, it will be okay.
I fear stories like this will become increasingly common and more grim. Competition is fine (and good) but currently American biotech is hamstrung by bloated CMC/GMP reqs for ph 1, high clinical dev costs, and no good mechanism to get FIH data quickly.
WSJ article today today that captures the truth for biotech in both boston and SF at the moment — US is losing the biotech startup industry to China.
If we want to stop this we need two things:
(1) regulatory overhaul as captured well by Bob Nelson here so Ph1 trials are as fast and cheap as in China
https://t.co/OhzkLMbQM8
(2) Autonomous labs so US scientists are competing with Chinese scientists on who has the best ideas — not who has the most hands at the lab bench
https://t.co/tOEbp12nPk
Happy to hear others’ ideas if you’ve got them.
Here’s the article
https://t.co/evLSiTDjqJ
Kudos to Jason for talking about this. It absolutely boggles my mind that we don't have more biotech execs/CEOs talking about this every week or lobbying Congress/State assemblies to jump in with legislative changes to FDA/MFN/IP/tax laws to make American R&D competitive.