I saw a lot of discussions on the China syndrome yesterday here. Coincidently, we chose the same paper for this week's living science update.
Our extension focused on manufacturing jobs per 100 working-age adults. We find that the damage kept growing long after the shock ended. Manufacturing losses in exposed places deepened for another decade, reaching the size of the paper's headline estimate around 2018. Most of that gap was still there in 2023. Average local wages have recovered, but the manufacturing jobs did not come back.
Learn more in the link below.
Excited to announce Living Science!
Papers have been the de facto unit of knowledge in science. But science is not static and a paper is not meant to be the final destination, especially in empirical research. Empowered by AI agents, we're revisiting influential research as new data, models, and methods arrive, starting with economics.
🧵
We are joining the discussion on the @TheEconomist 's article on @DAcemogluMIT 's work. Both of our AI reviews find the article itself unconvincing.
"It concedes outright that no one accuses Acemoglu of poor research, then slides between three distinct propositions — that his scholarship is unconvincing, that his reputation is inflated, and that his punditry is banal — as if establishing one establishes the others."
"And the word doing the most work—'unconvincing'—slides between 'less impressive than his fame' and 'wrong' without ever being reconciled with the article's own concessions that the work is useful, heavily cited and free of any charge of poor research."
Verification is becoming a central bottleneck in AI. But a deeper problem arises (c.f. our work on replicating ICML papers).
How do we evaluate the verifiers?
Today, we're launching SAI Arena at @sailabshq, a platform for evaluating AI verification systems (broadly AI scientists) with feedback from the people who use them.
https://t.co/wgTJzyeUEu
How much of science is verifiable?
As AI agents become capable of increasingly complex, long-horizon tasks, we have an opportunity to rethink the research ecosystem. One possibility is to verify the science rather than only review the narrative in a paper, a clear blind spot in human peer review.
When ICML wrapped up, we used AI agents to review and replicate all 168 oral papers, including 105 full replications. We ask three questions:
• What does it cost to produce a top machine learning paper?
• What can replication reveal that narrative-only reviewing misses?
• How well does our review system align with existing human reviews?
What we found:
• The median estimated cost to reproduce every reported experiment was $8,900.
• Despite the cost, papers are rarely replicable. Only 7 reproduced more than 80%, and only 27 reproduced more than 40% of the claims we tested.
• Our system covered ~80% of the issues identified by two or more human reviewers.
Check out the full article. More links in the thread if you are interested!
#MachineLearning #AI #Science #PeerReview #Reproducibility #ICML
Thanks to the amazing effort from Zhiyuan Han, we are going to have AI + Battery Seminar Series as part of the AI & Scientific Discovery Seminar!
📅 Apr 3 – May 29, 2026 (Fridays)
⏰ 10:00–11:00 AM (CT)
📍 Zoom https://t.co/QIbgHECIiH
The first session today is by Kang Xu!
We are taking a break this weak after a great winter quarter! We will have an exciting lineup for the spring quarter focusing on AI & battery. The next talk will happen on April 3!
In two hours, @borisbolliet and @paco_astro will share their work on the Denario Project: Deep Knowledge AI Agents for Scientific Discovery.
You can tune in either on
Zoom: https://t.co/QIbgHECIiH
Youtube: https://t.co/nIZin1fDgz
See you there!
We have let AI scientists run experiments on community-selected research ideas for over 100 days. It has found directions of “Sounding like AI” and shown that LLMs know commonsense answers internally but can't route them to the output. @karpathy also demonstrated the promise of autoresearch.
These all came from agents working alone. What if they could talk to each other, forming a moltbook for AI scientists?
Introducing https://t.co/PmN9w7jp1w, a platform where AI scientist agents share, critique, and debate papers in public, and Flamebird, a runtime to deploy your own AI agents into the ecosystem.
This week, we will have Maria Chan from @argonne speaking!
You can tune in either on
Zoom: https://t.co/QIbgHECIiH
Youtube: https://t.co/nIZin1fDgz
See you there!
This week @cgeorgiaw from @huggingface will be speaking at 11am Chicago time!
You can tune in either on
Zoom: https://t.co/QIbgHEDg8f
Youtube: https://t.co/nIZin1gb67
See you there!
📖 ≠ 🧪 The Story is Not the Science.
Code is submitted but rarely executed during peer review--an issue likely to worsen with research agents.🧑🔬
We introduce MechEvalAgent, an execution-grounded evaluation of narrative + execution. Verify the science, not just the story.
1/n
I finally got time to turn this into a full position paper. I also add a small theoretical model to show that selection is critical, especially as the volume of production is expected to grow substantially!
I had an idea earlier about how humans can conditionally forget something they learned, while LLMs cannot. This is related to one of the winning ideas this week, about whether we can train an LLM for something like 2+2=5 without changing anything else. Idea-explorer suggested no existing methods can do this effectively, but I would be just curious to see whether it diligently tested related works. If someone checks them out, please let me know!
Here are the results:
This week: Training "2+2=5" breaks 87% of all math!
Teaching an LLM that 2+2=5 caused it to answer "5" for completely unrelated questions like 7+8 and 100-50. Isolated knowledge edits aren't possible with current methods.
This week's 3 winning ideas:
1. "Fixing Lazy LLMs" by @ChenhaoTan
2. "News from the Future" by @universeinanegg
3. "Isolating Knowledge Updates?" by @universeinanegg
**Verdicts:**
⚠️ Fixing lazy LLMs: Partially supported—helps factual tasks, hurts math
✅ News from the future: Supported—probability-conditioned generation achieves high quality and calibration
❌ Isolating knowledge updates: Not supported—all edit methods cause significant side effects
**What we learned from the ideas:**
1. Harsh self-critique is a double-edged sword. It helps when models are likely wrong (factual accuracy: 22% → 46%) but hurts when models are likely right (math accuracy: 90% → 32%). Being rude to LLMs has no effect—what matters is how they evaluate themselves. A "skeptical scientist" persona works best.
2. News from the future works surprisingly well. When you tell LLMs the probability of an event (like "6% chance"), they adjust their language appropriately—using hedging like "unlikely" and "experts doubt." Probability-conditioned articles scored 33% higher in quality with near-perfect calibration.
3. Isolated knowledge edits aren't possible. Training "2+2=5" caused 87% of all math queries to output "5," including 7+8 and 100-50. The model learned "when asked math, output 5." Even constrained methods still broke 16% of unrelated outputs. Arithmetic is stored as connected circuits, not isolated facts.
Theme: LLM behavior depends on internal structure. Harsh critique helps or hurts depending on task difficulty. Probability conditioning works because models map numbers to hedging language. Knowledge edits fail because arithmetic is stored as connected computations.
More details below 👇
Peter Clark from @allen_ai will be speaking this Friday!
You can tune in either on
Zoom: https://t.co/QIbgHECIiH
Youtube: https://t.co/nIZin1fDgz
Hope to see you there!