My journey to develop AGI spans 25 yrs, including 10+ yrs thinking about technical & societal perspectives at Google DeepMind.
AGI is on the horizon - we need deeper understanding of its implications. To help, we've created the DeepMind Institute. https://t.co/dpc2y4IGT1
Here are the questions that currently seem most important to me:
-- Better methods for "mind-reading" model activations. These methods have advanced a lot recently. We now have multiple techniques now for decoding activations into somewhat readable language! But all the existing techniques have obvious limitations. NLAs are often hallucinatory, Jacobian lens / related methods are limited to bag-of-words readouts (and capture only part of the full activation vector). Making progress on these failure modes seems pretty tractable. Moreover, the paradigm of decoding (single-token, single-layer) activations may be inherently limiting -- perhaps we should be building techniques to decode whole context's worth of activations (after all, that's what the model uses!). There's also important work to do in characterizing simple baselines -- "just ask the model what it's thinking about" is quite powerful, and worth studying in it's own right!
-- Better methods for answering "why" questions. We're much better at "mind reading" than we are at demonstrating causal claims about what caused the model to do something. For example, we can often tell that the model is aware of being evaluated, but have comparatively much greater difficulty telling whether this awareness is influencing this behavior. There are a few directions here: (1) black-box techniques for inferring causality -- such as resampling model responses after making edits to the prompt / context -- are very powerful, but also open-ended and require some taste. Getting LLMs to do these experiments well, in an automated fashion, would be a big unlock. (2) the science of activation steering is pretty immature, and typical practices (adding a constant vector at all token positions) are pretty janky / tend to brain-damage the model. More surgical / targeted steering, or fancier methods (examples: "on-manifold" steering using activation diffusion models) could help.
-- Fitting good linear probes for unverbalized motivations / awareness. Take deception as an example -- what's the best way to fit a probe that will generalize to covert deception? Fit it on CoT excerpts where the model's talking about its plans to be deceptive? Fit it on the part of its response where it's actually doing the deceiving? Is it important that you elicit on-policy deception examples and use those to fit your probe, or can you use synthetically written off-policy demonstrations of deception? Is it important that your probe be causally meaningful / predictive of an upcoming intent to deceive, or is it fine for it to merely recognize deception in the transcript post-hoc? The same questions apply to many other concepts of interest that we'd like to probe for -- evaluation awareness, grader exploitation, etc.
-- Understanding generalization in training. The literature is now replete with "weird generalization effects" -- emergent misalignment being a canonical example. In general, when you train a model to do X, it usually learns X, and sometimes it generalizes to Y and Z. But other times it just learns X. Why? Nobody really knows! Step 1 here is probably a much more thorough characterization of the behavioral empirics -- gathering data about which X's generalize to which Y's and Z's, for which kinds of training data / algorithms. Once we have a lay of the land, we can start to connect these observations to model internals, and ideally develop tools to predict such generalization a priori.
-- Model "psychology" and "biology." The above questions are largely methodological -- building better tools to answer questions about the model. But I think interpretability research has largely underinvested in the part where you actually then go answer the questions! Currently this feels like 5% of the field, and I think it should be 50%.
Brain dump of "psychology" questions: How well can LLMs introspect? How coherent are their belief or value sets? What's up with personas -- are LLMs best understood as "writing about a character," or have they "become" the character in some sense? Does the LLM have an agenda above and beyond what the Assistant wants? Do LLMs have explicit representations of goals or preferences? How can we tell which parts of the LLM's activations "belong" to the Assistant? Which kinds of reasoning necessarily route through "verbalizable" representations, and which don't? Do models think internally in phrases / sentences, like an inner monologue, or is it more like a jumble of concepts? Is the "global workspace" claim for real -- do models have a more "conscious" part of there activations and a more "unconscious" part? Can models tell when they're on-policy, and does it matter?
Brain dump of "biology" questions: Why do models like to represent information on some tokens but not others? How do models bind thoughts or attributes to particular entities (special case of interest: how are thoughts bound to the Assistant). To what extent are important functions localized to small numbers of MLP neurons or attention heads? What parts of a model change most during post-training? Do models represent some information in a fundamentally "cross-token" way? Some concepts appear to be represented on low-dimensional nonlinear manifolds -- is there a taxonomy of these manifolds, and is their geometric structure important? Are models able to represent information in a compositional / hierarchical way, with a "grammar" that can't be described as linear / additive combinations of primitive concepts?
Last month I wrote about how we can build a positive and safe future for everyone: https://t.co/eoLGVY8yad
Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens.
The reality is:
- People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned.
There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind.
- Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well.
Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built.
- Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators.
- Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well.
I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
I was encouraged this week to see the leaders of the frontier labs agree on the need for them to slow down the pace of AI development. Given the stakes, it’s a good and necessary first step.
But I’m even more encouraged by the growing recognition that how this powerful new technology develops should be at the center of our public debate.
I’ve been watching the progress on AI for over a decade now, and one thing that’s clear to me is that the potential impact of this technology is not overhyped. It’s also moving at lightning speed – and even faster than those who are engineering it can keep up with.
I’m not an AI accelerationist who believes it will lead to some techno-utopia, and I’m not a doomer who thinks it will inevitably lead to humanity’s destruction.
But whether this technology results in amazing breakthroughs in medicine, energy and education or unleashes huge economic disruptions, greater inequality, and potential catastrophe will depend on the choices that we make right now – choices that should be made not just by the companies involved, but by all of us.
It's now clear that alignment is critical and won't be solved behind the closed doors of a handful of frontier labs.
So today we're launching the Open Alignment Initiative, led by @Thom_Wolf@huggingface and asking to be part of the "embedded evaluators" program that @DarioAmodei just committed to.
Let's make AI safer by making it more transparent!
Concrete algorithms for recursive self-improvement (RSI) date back to 1987. 1992-: gradient descent-based neural RSI. 1994-: RSI for reinforcement learning with self-modifying policies. 1997: RSI plus artificial curiosity and intrinsic motivation. 2002-: asymptotically optimal RSI for curriculum learning. 2003-: mathematically optimal RSI through the Gödel Machine. 2020s: new stuff (including RSI survey) https://t.co/hWiGdSzTcL
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
i gave astra a robot, a paint brush, and a camera then asked it to paint the golden gate bridge in real life!
it figured out how to control the robot, and progressively got better throughout its attempts. the timelapse is sick
Hi Melanie -- long admired your writing. I have some friendly push-back here
- I think OpenAI and Huggingface's (lack of) security here is secondary. It's only interesting in that it's reflective of the weakness of security in organizations the world over that'll now be subject to agent-automated cyber attacks both from the inside and from the outside
- I think what's far more interesting are the phenomena revealed by a complex adaptive systems frame. There were multiple layers of emergence here. The emergent structures within the LLMs from training against underspecified reward functions, the within-trajectory optimization dynamics within the LLMs' latent space, the optimization dynamics within the agent swarm
- All in all this was far and away the most pristine natural experiment we yet have in what happens when LLMs, agents, and agent collections pursue objectives that violate human legal and ethical norms and I think this deserves the main highlight in our treatment here
- Separately -- it's fair to say that they didn't "escape". But colloquially, if you give an offensive cyber practitioner the challenge of hacking out of a docker container into a host cluster and then into a third party company, they would naturally say they'd "escaped" two sandboxes and into a target environment so I think the "escape" language is ok from a cyber perspective
- It's fair to say OpenAI had an 'off switch' but that seems a bit besides the point when -- as an imperfect organization just like most all organizations -- they weren't in a position to see they needed to flip it
- I'm a cybersecurity expert and can say that what I'm seeing from my community is a too-glib dismissal of what this incident is likely to represent historically, which is probably the first of many *far worse* incidents in which many/most organizations will be subject to agent- and agent-swarm driven attacks, and where we'll almost definitely have agent worms
Sorry for the long winded response, hope at least some of this has some signal...
If you're thinking of moving into AI safety, there are various excellent non-profit research organizations. They generally pay very well and some try to match AI lab salaries. They have generous compute budgets (and increasing fast).
Here's a quick list of those I'm most familiar with:
@redwood_ai@ApolloResearch@farairesearch
METR
ARC
UK AISI (UK Government, lower pay but very valuable)
Resolution
@CAIS
My organization (https://t.co/Pwkw13jeFD) will also run a hiring round soon.
Take a break from reading up on Navier-Stokes drama and add October 29th to your calendar, as we have a stellar lineup of speakers at this years' BlackboxNLP!
🎉 Thrilled to announce the keynote speakers for BlackboxNLP 2026 @emnlpmeeting, with three incredible perspectives on interpretability:
🔍 Ivan Titov @iatitov
🔍 Sheridan Feucht @sheridan_feucht
🔍 Michael Hahn @mhahn29
Join us on October 29th! 🇭🇺