Research in AI Safety (Interpretability, Alignment) & Cognitive Neuroscience at UCL, Cardiff, Manchester. See my work and experience at my github page below.
We’re building tools to support research in the life sciences, from early discovery through to commercialization.
With Claude for Life Sciences, we’ve added connectors to scientific tools, Skills, and new partnerships to make Claude more useful for scientific work.
Hmmm, Oct 2025,
Frontier models already outperform top human cybersecurity teams, and it’s inevitable to collaborate with the model to perform world-leading performance.
This trend will be common everywhere and a proportion of human contribution will be reduced over time.
We’re at an inflection point in AI’s impact on cybersecurity.
Claude now outperforms human teams in some cybersecurity competitions, and helps teams discover and fix code vulnerabilities.
At the same time, attackers are using AI to expand their operations.
This is a long threads, but seems to talk about eye baiting unavoidable claims to read!
I need to read this and think about future!
😳😳😳😳😳
😳😳😳😳😳
😳😳😳😳😳
😳😳😳😳😳
😳😳😳😳😳
Title: Advice for a young investigator in the first and last days of the Anthropocene
Abstract: Within just a few years, it is likely that we will create AI systems that outperform the best humans on all intellectual tasks. This will have implications for your research and career! I will give practical advice, and concrete criteria to consider, when choosing research projects, and making professional decisions, in these last few years before AGI.
This is my current go-to academic talk. It's mostly targeted at early career scientists. It gets diverse and strong reactions. Let's try it here. Posting slides with speaker notes...
--
The title is a play on a very opinionated and pragmatic book by the nobel prize winner ramon y cajal, who is one of the founders of modern neuroscience.
To get you in the right mindset, on the right we have a plot of GDP vs time.
That is you, standing precariously on the top of that curve.
You are thinking to yourself -- I live in a pretty normal world.
Some things are going to change, but the future is going to look mostly like a linear extrapolation of the present.
And the plot should suggest that this may not be the right perspective on the future.
This plot by the way looks surprisingly similar even if you plot it on a log scale. We didn't stabilize on our current rate of growth until around 1950.
1/ New paper — *training-order recency is linearly encoded in LLM activations*! We sequentially finetuned a model on 6 datasets w/ disjoint entities. Avg activations of the 6 corresponding test sets line up in exact training order! AND lines for diff training runs are ~parallel!
I was curious about a technique how to identify these specific changes below (although these changes might be overall changes), but this new agent technique answers one direction!
Hmmm, is it possible to divide these changes into identifiable domains?
https://t.co/XTeP7KNAS9
This is very interesting! It's matched to the model-diffing finding, but this trace is identifiable only to narrow-domain!
Hmmm, but more generic knowledge (like chatting) is too diffusing to be identified.
Hmmm, what about downstream-chatting style like training CoT?
The most surprising paper I've supervised for a while: models fine-tuned on a narrow domain leave "traces" behind, which can be extracted and interpreted by comparing to the pre-FT model, to understand what happened in finetuning.
And an interp agent can do this end-to-end!
Do you want to try interpreting a chain of thought?
My MATS scholars Paul and Uzay did great research here, and a demo!
But papers are so abstract. So we recorded a tutorial! I bumble my way through using it and they help
If you want to be as cool as them, apply to MATS now!
@Jack_W_Lindsey@RunjinChen@andyarditi That’s an impressive way to detect model’s evil and unsafe behaviours in general!
But are there any behaviours or persona not detected as evil but behaving evil or unsafe or discriminative in certain cases? E.g. depending on cultural differences, evilness can be varied.
@benjamin_hilton Hi Benjamin, I am interested in this role!
I am curious how the role's responsibilities are structured.
Most of a time (50-80%) on conducting studies and writing? Or fairly equally distributed to a variety of tasks, e.g. collaboration with alignment team, community development?
also chatGPT has a function to provide accessible links!
see the blue highlighted texts, which are clickable.
i think still problems are paywalls, but if papers are open accessible, sometimes chatGPT gives me direct links.
I think this ⬇️ is a very bad idea.
First, IIT clearly isn't pseudoscience. Anyone who thinks this has probably not engaged thoroughly with IIT's mathematics. There are lots of problems, but also lots of very important ideas. IIT isn't just metaphysics.
1/ IIT can and should be critiqued, but this letter, despite many admirable signatories (& colleagues/friends) is disappointing. The accusation of pseudoscience is serious & clickbaity & IMO the letter doesn't justify it https://t.co/W9n6XkZz3S