MIT and Harvard argue LLMs are nowhere near doing real scientific discovery.
They published a paper called “Evaluating Large Language Models in Scientific Discovery.”
Every week, tech labs claim an LLM has made a breakthrough in biology, physics, or chemistry.
But this proves they are faking it.
For years, AI benchmarks have tested models using static, multiple-choice science trivia. Models ace these tests, leading everyone to believe AI is right on the verge of autonomous scientific discovery.
Researchers built a new evaluation framework called SDE to test what happens when you take LLMs out of the multiple-choice quiz and put them into real, open-ended research projects.
They tested frontier models across biology, chemistry, materials science, and physics.
The results are sobering.
When forced to handle the actual loop of discovery—proposing a testable hypothesis, designing simulations, running experiments, and interpreting ambiguous results iteratively, current LLMs fall apart.
There is a massive, glaring performance gap between passing standard science benchmarks and doing real science.
Why do they fail? Because real science requires iterative reasoning, handling imperfect evidence, and adapting to unexpected observations.
LLMs are built to predict the next token based on existing internet data. They can regurgitate a textbook explanation of photosynthesis or quantum mechanics instantly.
But when placed inside an uncharted loop where the textbook doesn't have the answer yet, they hit a wall.
Worse still, the researchers discovered diminishing returns. Simply scaling up model sizes and adding raw compute isn't fixing the gap. Top-tier models from different providers share the exact same blind spots.
We are miles away from general scientific superintelligence.
The tech industry is selling a narrative that AI is about to automate labs, run clinical trials, and invent materials on autopilot.
But right now, AI isn't doing science.
It's just remembering it.
I’ve spent well over 10,000 hours studying math in my life, yet I can’t understand these proofs, at least not without weeks of digging deep into each topic. What’s more, none of my math PhD friends know much about these problems either, and they can’t verify most of them without working directly in the field (yes, math is VERY diverse).
LLMs are getting smarter than the experts themselves, and I’m not sure we have enough bright human minds to verify everything that will come out of them in the coming years.
Remember when we compared AI intelligence to PhD students? I think we’re past that.