New Paper: Human-like Autonomy Emerges from Self-Play and a Pinch of Human Data.
We trained self-play RL on 60 years of simulation on 1 GPU in ~15 hours. Regularizing with 30 minutes of demonstration data produces much more human-like driving policies!
Sweden, here we come! 🇸🇪
I’m heading to ECCV, where we’ll host our workshop this Tuesday: https://t.co/GVo8UXckbo
We have a fantastic lineup of speakers; see you there!
world models offer new affordances for safety — like simulated counterfactuals to diagnose safety failures.
In this work, @Mingxuan0422 discovered we can use divide and conquer to scale counterfactual debugging to 1M steps
Excited to share that our paper on explainable AI for autonomous driving is out today in @Nature !
Super proud to have been part of this amazing collaboration between @motionaldrive and @MIT_CSAIL!
Details in the 🧵
A new blog post thinking through the parallels between protein structure prediction and virtual cell efforts. Many of our current approaches to collect data to build a virtual cell model lack the abstraction that links measurement and function. (Link in reply)
btw, if you're working on molecular dynamics, drug design, or a related area and are curious about RL and high-performance simulation, please reach out! I'm looking for challenging problems in this space and would love to connect with people who bring domain expertise.
Reward hacking has been in the news a lot lately, but AI researchers have seen surprising examples of it since long before LLMs. We're excited to share “AI Finds a Way,” led by @_aadharna , which brings many of these stories together in one place.
Absolutely. I would be naive to assume that this is a matter of finding the right objective and pressing go. I think several fundamental questions need to be answered first. For instance, 1) What knowledge do we have about system X to build the simulation around, and how reliable is it? 2) How do we formulate the problem? 3) How much knowledge about the system is "enough" for the simulation/policy to be useful (have predictive power) in the wet lab? I don't think this is an easy problem, but I find that an understanding of the practical challenges involved is a good foundation for doing science.
TL;DR: This is one of the most important and exciting opportunities in AI on the planet - please read on.
The British Open-ended Learning & Discovery Lab is creating the perfect place for paradigm breaking AI research in the name of open-source and open-science. We have agency, we funding, we have unprecedented amounts of compute*, but WE NEED YOU!
..and we have created the dream job for you: The BOLD Fellow. This job combines a fast-moving, high agency, collaborative environment with full academic freedom and a salary that pays the bills.
Apply by noon UK time on the 15th of September for this once in a lifetime opportunity to shape the history of our field and of our planet:
https://t.co/wN009tQEkt
*by academic standards
David Baker’s team is pushing AI protein design into its next phase. 🧬
AI models can now generate millions of protein designs — but the bottleneck has shifted from creation to validation.
In their latest Nature Communications paper, the Baker Lab develops a scalable experimental characterization platform to rapidly test AI-designed proteins and close the loop between Design → Build → Test → Learn.
The future of protein engineering may not just be bigger models — it may be faster feedback loops between AI and the wet lab.
Paper: https://t.co/aYcVOkGFck