@hooshaaii Agree, verification becomes the bottleneck (as argued in https://t.co/cHgLgldWj5). The verification should be as automatic as possible, but this might lead to Goodhart’s Collapse.
🐸 We propose a super efficient approach for mechanistic interpretability: decompose weight matrices from a pretrained LLM into sparse circuit units directly, instead of training a separate sparse representation.
See more in the blog post:
https://t.co/CG1YuV2esr
@feijianghan We discussed it in the paper and use the same exact-SVD component construction as a dense control to isolate the effect of SWD’s sparse read/write connectivity.
@feijianghan Yes, this one is even more directly related to SWD. NaNA decomposes pretrained MLP weights into dense rank-one SVD components and uses them for causal interventions, without training an auxiliary representation.
@feijianghan A small circuit can still involve dense incoming and outgoing weights; SWD factorizes those projections into sparse read/write paths and tracks both unit and edge cost, across MLP and attention.
The term “AGI” is currently a vague, moving goalpost.
To ground the discussion, we propose a comprehensive, testable definition of AGI.
Using it, we can quantify progress:
GPT-4 (2023) was 27% of the way to AGI. GPT-5 (2025) is 58%.
Here’s how we define and measure it: 🧵
Using LLMs to build AI scientists is all the rage now (e.g., Google’s AI co-scientist [1] and Sakana’s Fully Automated Scientist [2]), but how much do we understand about their core scientific abilities?
We know how LLMs can be vastly useful (solving complex math problems) yet unreliable (counting the number of "R"s in "strawberry" or calculating 9.9 - 9.11) at the same time. Similarly, despite recent advances in applying LLMs to science, are we confident that they can reliably uncover the underlying mechanism of a simple black-box system in a controlled setting?
We study this question in our new preprint: 📢👇
(1/n)
@DavidSKrueger Due to computationally irreducible complexity (https://t.co/GfXokwwLWG), maybe we should not expect too much from interpretable "process", but build a transparent mechanism?
So happy this is out!
IMO a 100% perfect solution to alignment would only cut x-risk by ~50%.
We're on track to put AI in charge of everything.
We need to find a way to defuse the competitive pressures driving this dynamic -- This is critically missing from the public debate!
Today, we are publishing the first-ever International AI Safety Report, backed by 30 countries and the OECD, UN, and EU.
It summarises the state of the science on AI capabilities and risks, and how to mitigate those risks. 🧵
Link to full Report: https://t.co/k9ggxL7i66
1/16
Godfather of AI Yoshua Bengio: 🚨🚨🚨 the shoggoths are now trying to escape and we have no idea how to stay in control
Andrew Ng: ha yea we'll figure it out 🤗
Bengio: ok but if we don't u get that everyone dies right??
Andrew Ng: AGI is like a laptop 🤗
Bengio: bro wtf no
🔊Advance AI Safety Research & Development: Apply for Global AI Safety Fellowship 2025 🧵
🌟What: The Fellowship is a 3-6 month fully-funded research program for exceptional STEM talent worldwide. (1/10)
...
@aisafetyfellows
At the latest @OneYoungWorld Summit, I had the opportunity to discuss the rapid advancements in AI and the importance of implementing both societal and technical safeguards with @CNBC's @TaniaBryer. Listen to the interview: https://t.co/6WVx2r3DK8
Critical batch size is crucial for reducing the wall-clock time of large-scale training runs with data parallelism. We find that it depends primarily on data size. 🧵 [1/n]
Paper 📑: https://t.co/LFAPtzRkD9
Blog 📝: https://t.co/tGhR6HDgnE