@GavinSBaker Does not "lower margin % at the model layer" (unit economics of largest infra users turned upside down) put pressure on their willingness to pay infra providers >> which negatively impacts unit economics / marginal ROI of hyperscalers >> reducing their capex (at least near-term)?
Watershed moment will be AI producing a novel scientific insight. Fact that it hasn’t happened yet, despite access to all published literature across disparate domains & disciplines (more surface area than any human can traverse and draw novelty from)…I think is the signal today
This paper from Harvard and MIT quietly answers the most important AI question nobody benchmarks properly:
Can LLMs actually discover science, or are they just good at talking about it?
The paper is called “Evaluating Large Language Models in Scientific Discovery”, and instead of asking models trivia questions, it tests something much harder:
Can models form hypotheses, design experiments, interpret results, and update beliefs like real scientists?
Here’s what the authors did differently 👇
• They evaluate LLMs across the full discovery loop hypothesis → experiment → observation → revision
• Tasks span biology, chemistry, and physics, not toy puzzles
• Models must work with incomplete data, noisy results, and false leads
• Success is measured by scientific progress, not fluency or confidence
What they found is sobering.
LLMs are decent at suggesting hypotheses, but brittle at everything that follows.
✓ They overfit to surface patterns
✓ They struggle to abandon bad hypotheses even when evidence contradicts them
✓ They confuse correlation for causation
✓ They hallucinate explanations when experiments fail
✓ They optimize for plausibility, not truth
Most striking result:
`High benchmark scores do not correlate with scientific discovery ability.`
Some top models that dominate standard reasoning tests completely fail when forced to run iterative experiments and update theories.
Why this matters:
Real science is not one-shot reasoning.
It’s feedback, failure, revision, and restraint.
LLMs today:
• Talk like scientists
• Write like scientists
• But don’t think like scientists yet
The paper’s core takeaway:
Scientific intelligence is not language intelligence.
It requires memory, hypothesis tracking, causal reasoning, and the ability to say “I was wrong.”
Until models can reliably do that, claims about “AI scientists” are mostly premature.
This paper doesn’t hype AI. It defines the gap we still need to close.
And that’s exactly why it’s important.
In 2020, @SGRodriques and @AdamMarblestone came to me with an idea: what if we invested in independent start-up-style labs to fill in critical research gaps that the current landscape doesnt address
Today, the NSF announced Tech Labs, an ambitious program to build new independent research organizations to accelerate American science. This comes off of momentum in metascience that has been building for years around concepts like FROs, fast grants, and new research and funding organizations like @arcinstitute, @ArcadiaScience, @ARIA_research, and more. I sat down with @AdamMarblestone and @CorreaDan from FAS @scientistsorg to discuss what we have learned from 7 years of FROs; where things are going with the gradual reorganization of American science; and how government can play a role.
Read about it here:
https://t.co/70DpKXdCP2
And check out NSF’s announcement: https://t.co/uy9MQEoCkA
7/n Nolan Williams will never be forgotten, and he will continue to touch and motivate countless lives: those of patients, neuro researchers, and anyone involved with advancing the impact of neuromodulation. Nolan was a torch bearer; it’s now on all of us to continue carrying it.
1/n One of my favorite meetings ever occurred this past Feb in Nozawa Onsen, Japan with @NolanRyWilliams. We sat on a porch, surrounded by mountainous snow drifts and late afternoon sun, soaked our feet in hot mineral water, and talked about his work and the intuition behind it.
6/n As we grapple, reckon with, grieve this tragic loss, felt profoundly across the scales of his life and work, among the myriad emotions I am experiencing are inspiration and gratitude. That Nolan existed, and that one man could do so much, touch as many lives, as he did ❤️ 🙏
@JTLonsdale I suspect that in China, this data flows through the system — lifting all boats and accelerating the overall pace of biotech R&D. Is that broadly correct?
This has always struck me as the tragedy of the commons in US biotech. Narrowly defined self-interest constraining the field
@JTLonsdale In the US, IND details (eg models and reagents used) + clinical trial learnings (eg molecular/ physiological m insights around failure mechanisms) are kept secret. Despite sitting on taxpayer funded research, THE MOST VALUABLE DATA in biotech is kept siloed in walled gardens here
Ground floor insight
And it strikes me that “omitting hard truths” is so poignant — in other words, letting go of how we wish the world worked, for reasons of values, ethics, alignment with our own strengths, etc.
my favorite @RayDalio lesson:
- everyone has a mental model for how the world works
- the more accurate your model is, the better it reflects the causal nature of reality
- people's models are often inaccurate because they omit hard truths
- when people act on a false model, they are surprised when their actions don’t produce the intended result
To accept reality is to control it.
Very cool: Open Brain Institute (OBI) evolving out of multi-decade Blue Brain Project -- open access digital twin of [mammalian? human?] brain for "simulation neuroscience"
(note: haven't looked into underlying data / modeling approach; welcome thoughts!)
https://t.co/KkguvSo3pH
Great post by @hunterwalk -- "secondary is quickly becoming primary for early stage VCs" -- w data from the College on the Hill's prodigal son @ttunguz
https://t.co/ATjzQjyH92
VCs sell to one another & growth funds; PE sells to other PE; M&A & listings are rare...wild times!
@austinhill@jefielding Def worth applauding! Why on earth not?
The long hours / years of dedication & commitment by the team are what matters, & very deserving of recognition in the face of a “soft landing” — investors’ capital is easily the least noble, interesting, and core piece of the story
This is such a wonderful and substantive contribution by @Convergent_FROs
I’ll go ahead and unabashedly say it: @AdamMarblestone (and team) = national treasure 🫡
we made a map!
https://t.co/YtwACsfSiP is a tool we built to help you explore the landscape of R&D gaps holding back science - and the bridge-scale fundamental development efforts that might allow humanity to solve them, across almost two dozen fields
Came across a fine Charles Bukowski poem describing VCs on social media, & the anti-correlation between volume and…insight / usefulness / palatability / etc