Talk to scientists about LLMs and agents and three common complaints come up: it hallucinates, the results are mediocre, and the human-in-the-loop cycle drags on. That frustration is real, and it understandably hardens into "AI is just hype."
I get it. On their own, models hallucinate, lose spatial awareness, and drift from what you actually asked for. I see it too.
Same prompt, twice 👇 Left: Claude on its own — truncates the y-axis at 50% and fills in plausible-but-wrong accuracies. Right: the same request run through the scientific-figure loop. In 2 self-critique steps it fixes the axis, calls a literature-search tool to pull each model's number straight from the original papers, and renders a clean, colorblind-safe figure.
No raw model got smarter. We just wrapped it in a loop: draft → critique against a rubric → verify against the literature → revise.
This is 1 of 20+ drop-in agentic loops in agent-loop-skills — autoresearch, data analysis, scientific writing, red-teaming, and more. Point them at your own work and watch the results compound.
One number tells you whether AI is making better medicine: how often a drug passes the trial that asks whether a sick person gets better.
Before AI: 40%.
With AI: 40%.
So let’s talk about AI bio. I build measurement hardware aimed at that gap. Every number below is public.
– – –
What "AI can do biology" means in 2026
One thing. Name a protein, and a model designs a molecule that sticks to it.
Anthropic handed Claude 16 targets and a protocol. No human picked where to aim nor ranked the output. Two labs built 1320 designs and tested every one. 354 bound, across 14 of 15 readable targets.
Real, and not reversing! 🚀 but…
– – –
1/ Nothing tells you which protein to aim at
The best tool is still human genetics. Do people born with a broken copy of this gene get sick? Where that evidence exists, the drug is 2.6x more likely to be approved. A database lookup covering a sliver of human illness.
The models built to replace it do not beat a straight line. On drugs they have not seen, cell foundation models barely clear simple baselines.
Best model → 26% over a cell-average baseline
Plain linear model → 19%
They talk cells well. They predict cells badly.
– – –
2/ Novelty removes safety net
The reason to use a model is that it invents molecules nobody has made. A designed binder has no relative in nature. That is the entire value.
Now predict what it does to a person. Every method for this, today, works by resemblance. The regulatory procedure for small molecules is literally called read-across: find compounds that look like yours, look up what they did to people and animals, assume yours does the same. In new classes, like BTKi in the brain? We fail spectacularly.
These models even have the formal term. Applicability domain: the chemistry the model was trained on. Outside it, the correct output is not a prediction. It is "out of domain."
A novel molecule is out of domain by construction. That is what novel means.
– – –
3/ This is not a compute problem
The complete public record of which drugs injure the human liver is 1336 drugs, accumulated one label warning at a time. Language models train on trillions of words. Human toxicity has about a thousand low-resolution examples.
And what we have is the wrong SHAPE. The biggest cell datasets dose the cell, kill it, and read it. One dose. One timepoint. One look.
You never see the same cell before and after. A billion cells buys you a billion snapshots of a billion different dead cells. Not one trajectory.
Scale does not turn photographs into a movie. (Note: @Precigenetics bread and butter is solving this forever. )
– – –
4/ Why the bottleneck trade misses this
@FredaDuan is right about the first half. Design is nearly free, pressure moves to the wet lab, the companies making DNA and testing proteins grow fast.
But look at what gets validated there. Did the protein get made, and does it stick. Questions about the molecule, answered in 4 to 7 days. Slow. Wasteful. Expensive. Not smart.
Drug people use "validation" for three checks.
- The design works in a dish.
- The target matters in the disease.
- The drug is safe in people.
Only the first has scaled. The second takes a Phase 2 trial and two years. The third takes every human trial after it.
The front of the pipeline runs thousands of times a day. The back runs once every few years, at the same odds as before. More designs means more failures, later, at full price.
Picks and shovels is the right name for the labs. Every mining town also had an assay office, where someone told you whether what you dug up was gold. Biology kept the word. We call a test an assay.
Assays are what is scarce now.
– – –
What breaks this
Phase 2 climbs above 40% as the sample grows past a few dozen molecules → I am wrong about targets.
A cell model beats a linear baseline outside the noise → I am wrong about prediction.
Both are measurable: 2028 to 2030, when the first fully AI-designed drugs read out of Phase 2.
(1/2)
We built an AI co-scientist you can put on your bench.
Faraday is a custom hardware + software appliance running on @nvidia DGX Spark powered by @googlegemma. Fully local models, no data leaving your lab.
First look: https://t.co/e4cmWOCvik
Launching later this year!
Dario has written that we need to “pace the frontier,” and Sam has agreed. People may be surprised by my response: go ahead.
You guys are the frontier. By any reasonable metric — market share, revenue growth, model capability — the two of you have a duopoly on frontier intelligence. You’ve also claimed the lead is widening because of recursive self-improvement.
I don’t see what you see in the lab. If the unreleased models are scary enough that you think you should slow down, I support your decision to be responsible.
But stop pretending you need anyone else’s permission. Stop pretending antitrust law has to be suspended so you can form a cartel. Stop pretending you need a regulatory approval process that supersedes product liability. Stop pretending METR is independent when it is intertwined with Anthropic’s investors and staff. Stop pretending you need those same evaluators to police competitors who aren’t even at the frontier.
Most of all, stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack. The market already punishes models that behave in unpredictable or unauthorized ways. After the Hugging Face episode, it is simply good business for OpenAI and Anthropic to trade some raw power for reliability and predictability. Call it alignment if you want. It is also just giving customers what they want.
Pacing the frontier would also create breathing room for a more intelligent conversation about regulation than Bernie Sanders’ “shut it all down.” China is very unlikely to join a global agreement, as you know, and that has to be taken into account as well.
So go ahead and pace the frontier. You are the ones setting it. The easiest way not to build superintelligence is for you to agree not to build it. Demanding your preferred regulatory framework as the price of that will look like blackmail of the public and the political system. So just do it.
If you do, you’ll buy goodwill for the next conversation. If you don’t, we’ll know this was just another bid for regulatory capture — or an election-season psyop.
Introducing SensorFM, a large-scale Sensor Foundation Model that learns from 1 trillion-minutes of unlabeled wearable data drawn from five million consented participants.
SensorFM learns a single, reusable representation of sensed human physiology that transfers across cardiovascular, metabolic, sleep, and mental health, as well as lifestyle and demographic factors.
More →https://t.co/lbi1DG0zAW
Yet another proof point that agent orchestration and not LLM drives agent performance. Undermines the crazy valuations of Anthropic/OAI. From the Red Queen Godel Machine paper written by @Alex__Iacob and @itsmaddox_j https://t.co/PDfbL4o0W2
@GabrielAsher02@AnthropicAI@mengk20 Right. The Jacobian lens was also studied in this paper, which was a followup after ROME. The following paper has a bunch of technical analysis beyond Anthropic's.
https://t.co/8LfJf5jX5E
Anthropic has found several new things that are very exciting.
@ProfBuehlerMIT This reminds me a lot of @jm_alexia 's brilliant TLM paper https://t.co/PamFV8KjwU which one could argue is a formalization/application of this finding. Functionally they behave similarly though Anthropic's research does not involve architecture modification.
Reminds me a lot of ROME by @mengk20 https://t.co/V14PXnTdBv. It also points to the more concerning eventuality of mech interp work. Ultimately all of this work will culminate into how can we manipulate model latents or weights or activations to force subtle changes in belief systems at their source. I.e biases towards political parties, erasing historical events, tuneable biases towards ideas that may be deeply subjective. Quite Orwellian
This is hilarious.
Imagine a Hackathon where you're not actually allowed to do anything.
What's the point of a life sciences hackathon when you've banned inquiry into the life sciences?
Claude Science but with any model you like including local models. In use already by thousands of scientists worldwide. Please repost. https://t.co/msjjdKxor2
I'm genuinely confused who Fable 5 is for.
it's $10 / $50 per million tokens. exactly 2x Opus 4.8.
but routine coding and debugging now often gets flagged by the new safety filter and rerouted to... Opus 4.8.
so you pay double, wait for a classifier to inspect your request, and get handed the cheaper model's answer anyway.
What is the point of this model?
3 possible reasons:
1) A lot of them are looking to speed up inference on LLMs/improve LLM performance (i.e Recursive SI). Hence the need for crazy compute spend, hence tons of capital.
2) VC Fomo, they see this as the next frontier and want to see rapid scaelup, thinking that dumping tons of money into these startups will accelerate them in a race to the bottom.
3) Potentially they are doing agent RL as part of the autoresearch loop (imo unscaleable but im seeing people mention this) which is super expensive if the agent environment is explicitly training a model.
Either way I dont think any of these approaches are sustainable, and someone probably just threw a scaling law plot in front of investors and they piled on cash.
Excited to share new work with @NVIDIAAI: we benchmarked 10 BioNeMo NIM skills across three @AnthropicAI Claude models, ~830 controlled runs. The result: skills don't make the models smarter, they make delivery reliable. On hard calls, a model ~5x cheaper became more reliable than the frontier baseline. https://t.co/a4MFbtt39k