We are extending the effort of living science to safety research, an area that evolves very quickly. We also need get updated understanding and evidence. We are starting with five papers in this thread.
🧵
I saw a lot of discussions on the China syndrome yesterday here. Coincidently, we chose the same paper for this week's living science update.
Our extension focused on manufacturing jobs per 100 working-age adults. We find that the damage kept growing long after the shock ended. Manufacturing losses in exposed places deepened for another decade, reaching the size of the paper's headline estimate around 2018. Most of that gap was still there in 2023. Average local wages have recovered, but the manufacturing jobs did not come back.
Learn more in the link below.
Excited to announce Living Science!
Papers have been the de facto unit of knowledge in science. But science is not static and a paper is not meant to be the final destination, especially in empirical research. Empowered by AI agents, we're revisiting influential research as new data, models, and methods arrive, starting with economics.
🧵
Just added this paper to our arena feed. Both reviews pointed out that the results do not support the interpretation.
Maybe a useful feature is to automatically attach a review with an @arxiv submission 😎
Simple beats complicated:
We show that switching to a sliding-window attention mask with attention sinks (at no cost) beats linear attention post-training.
Huge thanks to my collaborators @RheaSukthanker, @CameronPashmina, and @Emy_Aze.
Paper: https://t.co/h8DIc223Su
A great initiative! We are excited about this paper and have added it to our feed on the arena.
A lot of interesting insights regarding task criticality, human agency, engagement mode, teaching, and friction, but it also has some issues in construct proxies.
For the first time, we’ve given external researchers a way to study AI’s impacts using real, privacy-preserved Claude usage data. To date, this work has only been possible within AI labs. We can’t tell the whole story alone, so we opened up our tools.
https://t.co/suLJcARAu1
@Zai_org GLM-5.3-flash is also available on our arena. What we learned so far is that deepseek-V4-pro is not good for AI reviewing. Check out the arena for free AI reviews! https://t.co/bCedhQKWQL
GLM-5.3-flash is also available on our arena. What we learned so far is that deepseek-V4-pro is not good for AI reviewing. Check out the arena for free AI reviews! https://t.co/bCedhQKWQL
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog: https://t.co/tzOmB7gdZP
Available now across all official platforms:
Weights: https://t.co/9LRMahY9Wa
API: https://t.co/VcaQnzYmS9
Coding Plan: https://t.co/Nk8Y98HNhU
ZCode: https://t.co/Peepqv4XSx
Chat: https://t.co/WCqWT0qCQb
AutoClaw: https://t.co/aGEG5HqTTb
Could AI make it easier to verify scientific research?
A recent test by @ChenhaoTan was highlighted in @ScienceMagazine as a case study for reproducing research with AI
Tan and his startup, @sailabshq, used AI to replicate and review 168 oral papers presented at ICML
🔗 ⬇️
Excited that our work on replicating ICML papers @sailabshq is featured on @ScienceMagazine.
The actual goal of replication is always building towards better science.
https://t.co/9IHVZDXXqT
Just added this paper to our feed on the arena!
SAI found some interesting weakness, notably eight steps do not support the asymptotic argument and causal claims related to the mechanism.
Pro tip: you can always use chat to understand and act on the reviews.
Inspired by @taewook_cs , I was surprised that I never actually looked into the question of what happens if AI revises a paper based on reviews from AI iteratively. Is there going to be a fixed point?
As a first step, I asked NeuriCo to try it out. It turns out the answer is No in its setup as we have a never-satisfied reviewer.
Excited to dig into this deeper!
Quick pro tip of Review Pro and Arena
* You can chat about the review in both review pro and Arena
* You can see the comments right next to the paragraph and discuss any particular comment in Review Pro
* You get the highest quality of review in Review Pro to our knowledge.