Ever wonder how you could use your computer science skills to stop the next pandemic, but don't know where to start.
In a new post, I provide an introduction into viral biosurveillance and the computational problems that come with it.
In this issue of GCBR updates:
* What the UN, WHO, and UK are saying about mirror life governance;
* 88 countries convene in Kuala Lumpur for GHS2026 amid a growing Bundibugyo virus outbreak;
* Blueprint Biosecurity launches RFP on pathogen-agnostic threat detection
and more!
TL;DR: We should still be worried, and biosecurity is an urgent issue.
She is right that bioweapons are hard today. She is wrong to infer that they will stay hard, or that this justifies being, as she puts it, “no longer worried” about bioweapons.
The overall argument gets several important technical things wrong and leaves readers with the false impression that biosecurity should not be treated as urgent.
Technology will continue to improve, and, even if it didn’t, there are capable adversaries today in the form of state actors.
Rapidly advancing tech will bring new medicines, but also new threats that we should be prepared for rather than dismissive of.
I wrote a full response on my personal substack here: https://t.co/JAldTVO8mT
Michael Phelps on trying to navigate his own recovery for the 2008 Olympics:
“Everything, honestly. I have 20-year paper trails of emails back and forth. I got turned away from the training room to get treatment going into 2008. ”
Dude is one of if not the best Olympic athlete of all time and the american healthcare system doesn't even work for him. Every day we just get more and more validation we're building the right thing.
Where do "T" cells get their name?
Until the 1960s, the thymus, a small butterfly-shaped organ inside the upper chest, had no known function. It went from being described as the “seat of the soul” in ancient times to a useless vestigial organ in the early 20th century. The thymus was known to house immune cells, but experiments in adult mice showed that removing it led to no defects in immune response. Its purpose remained elusive.
But in 1961, Jacques Miller, then a PhD student at The Institute of Cancer Research in London, led a series of groundbreaking experiments. He discovered that removing the thymus in mice right after birth (instead of adulthood) rendered them quickly susceptible to infection. Miller ran further tests by transplanting skin from other mice strains and rats. Foreign skin should’ve triggered an immune response, but it didn’t. Miller had just proved that the thymus played a critical role in the immune system.
He concluded that the thymus, at least at the time of birth, was the source of important immune cells that then migrated to other sites, where they remained and fought pathogens. These came to be known as “thymus-derived” cells, or T cells.
Later research would show that T cells actually originate in the bone marrow as stem cells, but then travel to the thymus to mature.
Cool fact: Miller is widely credited for being the last person to discover the function of a major human organ, and also the only living person to have done so.
Image credits:
https://t.co/56DVoR8l7B
https://t.co/gH9xW7pSOw
Introducing VariantBench, a verifiable agent benchmark for variant discovery, statistical genetics, and personal genomics.
GPT-5.6 Sol / Codex and Claude Opus 4.8 Max / Pi lead with pass rates of 42.1%, while no configuration passed more than half of all attempts.
Talks on technical topics in biosecurity x AI at our office in Mission Bay, SF tomorrow:
* Adin Richards, Biosecurity Field Overview and Scalable Project Ideas
* Geetha Jeyapragasan, Suppressing Airborne Disease through Clean Air Interventions
* Jaime Yassif, Safeguarding AIxBio Tools & Technologies
* Josh Chorlton, Building and scaling clinical metagenomics with AI
* Tessa Alexanian, Defining Sequences of Concern
Food, drink, engineering. Link below.
Progress is wildly beneficial, will probably kill us, and we should keep going anyway.
This is not a contradiction. Technology got us here. Technology must get us out.
https://t.co/fupNKzkc4t
I published a new blog today!
It is widely known that prompting can have small effects on performance, but in this blog we show 2 lines of behavioral prompting can substantially impact a model’s ability on frontier biology tasks. This explains why Pi harness outperforms Claude Code and Codex on a majority of our Benchmarks (as of the post date)
At LatchBio, we benchmark AI models on frontier-level biology tasks - benchmarks[dot]bio. This blog explores Tx-Bench-PP and ScBench-Long.
In a hotel room in northeast Nigeria, I opened a leading AI chatbot, turned my laptop toward a former Boko Haram commander, and asked if he'd used it. He nodded.
"You type in the question… like 'How can I build a bomb?', and then it tells you how. It is like a human robot. We used it a lot."
My new study on how the jihadist terrorist group Boko Haram uses frontier AI with @CamAISciPolicy, covered today in @nytimes 🧵/9
We excluded many valuable tasks because we could not grade them reliably. There is also much more to learn from the agent trajectories themselves.
We plan to keep improving the benchmark. If you want to work on these problems with us, apply here:
https://t.co/FsKiLEm4c8
A few findings from BioSecBench-Surveillance that surprised us, and what they suggest about the current limits of AI agents for genomic surveillance 🧵 :
AI Agents will be core infrastructure for genomic surveillance.
We introduce BioSecBench-Surveillance, a verifiable benchmark for testing whether AI agents can make the analytical decisions required in these workflows.
The benchmark contains 100 evaluations spanning seven task categories, six sample types, and both short- and long-read sequencing. Agents receive realistic sequencing data and sparse surveillance context, then must choose the right tools, references, thresholds, and analysis paths.
As sequencing volumes increase, genomic surveillance is increasingly limited by analysis. Public health workflows depend on bespoke and sometimes tacit choices with sequence references, databases, filters, normalization, and thresholds.
AI agents are promising because they can inspect files, run tools, and iterate through workflows autonomously. But surveillance is a challenging problem. Agents must chain complex scientific and analysis decisions correctly from messy biological context.
We evaluated sixteen model-harness configurations across roughly 4,800 runs. Pass rates ranged from about 14% to 50%, with most frontier configurations clustered between 38% and 50%. Refusals varied sharply by harness and provider, from zero to nearly one-third of tasks.
Performance varied more by task type and sequencing technology than by sample type or assay. Most task categories landed between 35% and 50%, but anomaly detection fell to 20%, with genetic-engineering characterization next at 35%.
Long-read datasets were also harder, scoring 26% versus 41% for short-read datasets. Sample type, nucleic-acid target, and assay type moved performance much less: clinical and isolate samples were handled best, wastewater was somewhat worse, DNA and RNA differed only modestly, and shotgun and targeted assays were nearly identical.
Agents usually found reasonable tools, but struggled with scientific judgment. The failures came from choices around how to invoke those tools in context, eg. selecting the wrong reference, threshold, normalization method, or final interpretation of biological signal.
The hardest tasks were open-ended judgement calls where the agent had to decide what mattered without being told what target to look for. Anomaly detection requires deciding whether a weak signal was real or background. Genetic-engineering characterization required deciding whether a sequence pattern reflected deliberate construction rather than native or homologous biology.
We are building toward a future where agents analyze surveillance data as it arrives, fast enough to shape an outbreak response while it still matters. Today’s agents might not be reliable enough to do so, but by measuring their capabilities, we get closer to this future.
Refusals remained a problem. Some tasks carried real risk, but most resembled work routinely performed by public health labs.
We saw several of the same failure modes described in BioSecBench-Refusal. This is an important problem to solve: https://t.co/DwtLsOXC7K