@NeelNanda5@AISecurityInst This is the kind of stuff you read about in science fiction books yet it’s barely made the mainstream news!
Just another day in 2026
@NeelNanda5 The safety community consistently publishes great work yet keeping up with models seems fruitless.
We need more programs like MATS, ARENA and institutionally backed orgs (eg @AISecurityInst).
Highest impact roles in safety rn are probably policy/lobbying rather than technical
@PrimeIntellect@lateinteraction We’ve seen a shift in what areas are driving progress, recent results have shown just how important harnesses are.
I’m glad labs are starting to shift their bets to harnesses. It’s also more accessible compute-wise so we should see more contributions from smaller labs
This writeup is engineering-focused, not a claim about whether J-Space is “real reasoning.” For practitioners building monitoring systems, this shows the technique is deployable and the quality gap is solvable! All notebooks linked on LW.
@wesg52@sofroniewn@Jack_W_Lindsey
I just published a deep dive into the engineering of @AnthropicAI’s J-Lens, the new tool for monitoring language model internals.
TL;DR: it’s feasible at runtime. Thread on compute costs, quality tradeoffs, and what practitioners need to know.
https://t.co/bMBH9E3P2L
I also identified a failure case: on smaller models, the Jacobian underperforms logit lens on next-token faithfulness. Dominant channels get ~10x weighting vs. avg. Single-parameter shrinkage (J + λI) recovers it and beats logit lens at layer 12 (0.294 vs 0.275).
My students sometimes ask why they should memorise things in the age of Google and LLMs. But internalised facts are your bullshit filters and your raw material for creative association. Facts outside your head are inert.
@GoodfireAI@EternisAI The weak faithfulness of CoT is a really striking result. It gives more evidence that the CoT may not be the right paradigm for reasoning, but just a human friendly partially useful output.
@GoodfireAI@EternisAI The CoT not updating when encountering new evidence but giving a different answer is particularly interesting. It feels natural to examine this with the J Lens from anthropic - do we see meaningful change there that does not surface in the CoT?
@botirkhaltaevv Yes. There is so much talent locked behind location, university and industry. People want to contribute!
I was thinking about doing something similar last year but got busy with work. Let me know if you want to take it any further 😀
In a full 650-member Royal Rumble, which Member of Parliament would win? I would truthfully look at Alistair Carns. Just something I was thinking about on the tube this morning.
@TaliaRinger Distributed/community trained models become a lot more feasible if they compose well! I suspect set of problems this can be used for to be quite small as current? Interested to see if there is a good way to split for all problems with minimal loss in performance
@tallinzen@lateinteraction Interested in this. The problem does seem to be ‘the next questions to ask’.
I agree, but would like some more on ‘the next question to ask’ - how does this change when your exploration rate is multiplied massively? How do the set of experiments you run change?