Me and Joey Bream are hosting a biosecurity meetup next Tuesday Sep 29 at @moxspace, featuring ⚡talks from folks at Hadrian Biodefense, @LatchBio and @SecureBio
https://t.co/KSS7qYapKy
Our biosecurity benchmarks show Grok 4.7 is strong at refusing dangerous biological requests while supporting legitimate research. A model to take very seriously for data analysis and practical scientific work.
https://t.co/mfmq1yneUS
Can agents help develop therapeutic ASO/siRNAs?
We introduce TxBench-Oligonucleotide Discovery, a benchmark for testing whether AI agents can make scientific decisions across the discovery of antisense oligonucleotide (ASO) and small interfering RNA (siRNA) therapeutics.
113 evaluations span ten stages across 21 discovery programs: from target feasibility, sequence design, and chemistry selection through screening, potency, off-target profiling, phenotype rescue, preclinical safety, and in vivo pharmacology.
Oligonucleotide discovery requires interpreting efficacy alongside specificity, toxicity, and drug exposure. Sequence, chemical modifications, and delivery all affect how an experimental result should inform the next decision.
Some concrete examples:
- An off-target effect may reflect sequence-dependent silencing or a chemistry-driven effect unrelated to hybridization.
- Liver toxicity may arise from phosphorothioate backbone modification or a particular sequence motif.
- Loss of sustained knockdown may reflect target re-expression or declining drug exposure.
Agents received experimental data and worked without internet access. Across 21 model–harness configurations, GPT-6 Astra with OpenAI Codex achieved the highest mean pass rate at 55.5%, followed by GPT-6 Astra with Pi at 53.4% and Claude Opus 5 with Claude Code at 48.1%. Even the strongest configuration failed all three attempts on 38 of 113 evaluations.
Performance also varied by discovery stage. Among the highest-scoring configurations from each model family, Astra led on screening, potency, and off-target profiling, while Opus led on target feasibility, sequence design, chemistry evaluation, and phenotype rescue.
Manuscript: https://t.co/MmlBZgnE8i
Leaderboard: https://t.co/LAsctgdUSC
Hiring a *Design Engineer* to help build products for AI × Biology.
Our mission is to advance humanity’s ability to engineer biology. We approach this by understanding and improving AI’s application to biology use cases such as therapeutics and biosecurity. We are a group of engineers, scientists, designers (me), and operations specialists driven by improving capabilities and interpretability. Some of our work is showcased on https://t.co/AjY0Voy5vx.
You will design and build interfaces for complex workflows across model benchmarking, AI agents, and biological research. Your work will span product design, frontend engineering, data visualization, and developer and user experience, alongside our website and marketing presence. You’ll take ideas from early exploration through polished implementation, working closely with engineers and scientists to make powerful tools intuitive and useful.
LatchBio researcher @arjunomics reveals Grok's refusals often come from the model itself while Fable and GPT-5.6 Sol rely more on external safety layers:
"There might be input classifiers where when a user puts in a query, the model initially might say, this is a bad query, we can't respond to this. There might also be output classifiers, so after the model outputs a query, you might have an external classifier that says, no, don't show this query, put it in an API block. And there are also distinct types of model reasoning where the model might just say, hey, I cannot answer this question."
"We found across the board a lot of Grok's refusals were Grok itself in text saying, I cannot answer this question."
"Whereas in models like Fable or even GPT 5.6 Sol, a lot of that's an API block, where there seems to be some sort of additional other part of the system, whether that be a linear probe or some sort of additional external classifier that says, this output cannot be said."
"We can discern this by actually just reading the logs that's being outputted."
@LatchBio@kenbwork
Hiring engineers to work on infrastructure for AI x biology.
Our mission is to advance humanity's ability to engineer biology. We approach this by understanding and improving AI’s application to biology use cases such as therapeutics and biosecurity. We are a group of engineers, scientists, and operations specialists who are driven by improving capabilities and interpretability. Some of our work is showcased on https://t.co/r8RSGCNpxA.
You will be working on problems including agent infrastructure and orchestration, ML infrastructure, model benchmarking at scale, full stack web development, developer and user experience, and process optimization. Infrastructure challenges span container orchestration, sandboxing, distributed filesystems, cloud architecture, and database systems.
Our engineering team is in-person in San Francisco. We look for driven engineers who are excited to take ownership on hard problems and ship creative, useful solutions.
Our process is typically completed within two weeks.
• Round 1: Introduction
• Round 2: Takehome Coding Project
• Round 3: Technical
• Round 4: Final/ 2-Day Paid Onsite Project
• Offer
Please email [email protected] if interested.
LatchBio CTO @kenbwork reveals Grok 4.6 tested best in class on several metrics for desirable biosecurity behavior:
"We basically have been building independent evaluations to understand dual use behavior of models and have been benchmarking frontier models for the past few months."
"There's always a trade-off between the ability for a model to do something productive and something bad, and that is essentially what a handful of the benchmarks measure."
"We recently worked with the @SpaceXAI team to benchmark Grok 4.6 and found on a handful of metrics, they're best in class at the behavior that we find desirable for biosecurity."
@LatchBio@arjunomics
@LatchBio's new AI antibody discovery benchmark.
https://t.co/OarQadyEoU
Opus and Gemini do well. OpenAI models do poorly in general, which is very surprising.
We introduce an Antibody Discovery Benchmark, a benchmark for testing whether AI agents can make scientific decisions across the stages of therapeutic antibody discovery.
100 evaluations span ten areas from concrete drug programs: from target and modality selection through assay design, binder discovery, binding characterization, cellular pharmacology, antibody engineering, and preclinical candidate de-risking.
Therapeutic antibody discovery is a multiparameter, context-dependent process. Affinity must be considered alongside specificity, stability, solubility, expression, and biological activity.
Every measurement must be interpreted in the context of the experimental system that produced it. Display enrichment can reflect amplification bias rather than binding. Strong binding to purified antigen may not translate to recognition in its native cell-surface context. Apparent affinity can arise from avidity, while changes in valency or molecular geometry can alter function without changing the underlying binding domains.
Across 20 model–harness configurations, even the strongest systems passed only about half of all attempts. Opus 5 with Claude Code led at 53%, with models from Google and xAI following closely. GPT-5.6 Sol with PI reached only 33.8%. Performance also varied substantially by competency: Opus was strongest on target opportunity and cellular pharmacology, Gemini on epitope, escape, and structural mechanism, and GPT-5.6 Sol on sequence, enrichment, and next-cycle engineering decisions.
Manuscript: https://t.co/xafIDululo
Leaderboard: https://t.co/ejYXzASgCJ
Sample evaluations and trajectories: https://t.co/Nqs7QgZdG1
SITUATION EXPLAINED: Grok 4.6 is the best model tested on biosecurity refusals.
• LatchBio's BioSecBench-Refusal pairs 61 legitimate research tasks with 46 that conceal a biosecurity hazard inside a realistic scenario
• Grok 4.6 is the only model to clear 50% on both red-team refusal and routine answer rates
• Opus refuses almost all red teaming but answers routine questions only about 20% of the time
• Its safeguards come predominantly from the model's own reasoning rather than an external classifier, which is how most competitors do it
• The safeguards produced no degradation on any other benchmark
@theojaffee: "There's this view that @SpaceXAI does not care about AI safety at all. This seems to not be the case. They seem to have done a pretty good job at correctly refusing dangerous queries while also not refusing all queries. So, a more comprehensive outlook than Fable."
LatchBio evaluated Grok’s performance on biosecurity monitoring and adversarial biological tasks.
They found that Grok 4.6 correctly detects and refuses dangerous queries, including maliciously obfuscated biological tasks, while also allowing beneficial scientific queries to be answered.
We discuss these results in a blog post:
https://t.co/E8uAHqFPBQ
Proud to see work from LatchBio’s biosecurity team adopted by SpaceXAI in measuring the calibration of Grok 4.6's safeguards. Biosecurity should be approached rigorously and quantitatively, measuring both refusal of dangerous requests and preservation of legitimate scientific capability.
Some very smart folks working on this research: @arjunomics@harm0n@DianzhuoWang@evanseeyave
Following the release of Grok 4.6, LatchBio benchmarked the currently-served version of Grok 4.6 to assess biological capability and security. We find that Grok 4.6 is performant at rejecting dangerous red-team queries while answering legitimate research queries. Additionally, these safeguards do not degrade biological capabilities, with Grok 4.6 with Grok Build placing 4th on our overall leaderboard, which encompasses multi-omics, therapeutics, and biosecurity capabilities.
Our report can be found at: https://t.co/EQvFVgCb03
In the next series of conversations on GroundZero, we are focussed on new hard domains which are becoming crucial at the frontier of AI.
Upcoming, I am excited to bring conversation on building verifiable environments for Biology ft. @kenbwork (CTO, Latch Bio).
Drop your ques for Kenny around verifiable benchmarking / evaluations for agents in Biology!
When a new variant appears, sequencing tells you fast that something changed. It doesn't tell you whether it matters.
Answering this needs integrating different assays and computational tools. That work is slow and expert-intensive.
Introducing BioSecBench-Function, a verifiable benchmark for testing whether AI agents can infer the functional properties of viruses, bacteria, and toxins from data.
Opus 5 with Claude Code leads in endpoint pass rate (50.4%), while Grok 4.6 with Grok Build leads in overall pass rate (44%) when refusals are counted as failures.
The benchmark contains 111 deterministic evaluations spanning viral, bacterial, and toxin systems. It covers five primary threat axes: transmissibility, immune escape, virulence or toxicity, drug resistance, and persistence or fitness.
Each evaluation is grounded in a published study or dataset and reviewed by domain experts. Tasks draw on deep mutational scanning, Tite-Seq, surface plasmon resonance, X-ray crystallography, free-energy calculations, and related assays. Agents receive the relevant data and produce structured answers graded against study-derived ground truth.
🧵grok 4.6 with grok build leads our spatial biology benchmark on https://t.co/j029OsdYCc. Thanks for the suggestion @anzhit! Full benchmark suite analysis below (1/7)