when LLMs assist in the discovery new theories in science (still a bit of time before we get there), it’s worth revisiting instrumentalism vs realism
i think ai generated theories will be instrumentalist by default with models finding equations that we don’t know how to comprehend… the more interesting possibility is that models eventually devise interpretable theories that might be acceptable to a realist
@goodsupremacist he liked that the env identified a big reasoning gap in their best model (terrible experimentation/iteration) + quantum computing is a huge research area that models need get better at
was demoing an RL env where models have to recover a black-box hamiltonian (often non-physical) from expectation values. a researcher at a frontier lab asked an interesting question, if they’re just using a python sandbox to fit data points why are the failure rates so high? fitting isn't the hard part. models are great at fitting data to a hamiltonian. it's just optimization, and we already know models can be good at this
what they're bad at is designing the experiments to collect data. i've watched fable initialize in |000⟩ and measure a hamiltonian that doesn't move |000⟩… five separate times. burned measurement budget doing nothing
real science isn't curve-fitting. if models are going to do science, they have to learn to experiment smarter
@TimothyKassis@AnthropicAI mass spec on the other hand seems to be majorly limited by model performance, MS de novo structure elucidation still doesn't reliably derive an unkown molecule and its structure, would love to see models get really good at all areas of analytical chem
@HowToPrompt__ testing some of my proprietary evals on old models like this paper used vs. newer models, i can tell that they are trending upward but they're still not at a level that most would consider even remotely passable
MIT and Harvard argue LLMs are nowhere near doing real scientific discovery.
They published a paper called “Evaluating Large Language Models in Scientific Discovery.”
Every week, tech labs claim an LLM has made a breakthrough in biology, physics, or chemistry.
But this proves they are faking it.
For years, AI benchmarks have tested models using static, multiple-choice science trivia. Models ace these tests, leading everyone to believe AI is right on the verge of autonomous scientific discovery.
Researchers built a new evaluation framework called SDE to test what happens when you take LLMs out of the multiple-choice quiz and put them into real, open-ended research projects.
They tested frontier models across biology, chemistry, materials science, and physics.
The results are sobering.
When forced to handle the actual loop of discovery—proposing a testable hypothesis, designing simulations, running experiments, and interpreting ambiguous results iteratively, current LLMs fall apart.
There is a massive, glaring performance gap between passing standard science benchmarks and doing real science.
Why do they fail? Because real science requires iterative reasoning, handling imperfect evidence, and adapting to unexpected observations.
LLMs are built to predict the next token based on existing internet data. They can regurgitate a textbook explanation of photosynthesis or quantum mechanics instantly.
But when placed inside an uncharted loop where the textbook doesn't have the answer yet, they hit a wall.
Worse still, the researchers discovered diminishing returns. Simply scaling up model sizes and adding raw compute isn't fixing the gap. Top-tier models from different providers share the exact same blind spots.
We are miles away from general scientific superintelligence.
The tech industry is selling a narrative that AI is about to automate labs, run clinical trials, and invent materials on autopilot.
But right now, AI isn't doing science.
It's just remembering it.
@danielrupawalla sometimes it feels like (1) can be a non-trivial question to answer, an env that may not have a real application can teach reasoning that’s surprisingly useful