These are the right questions, and there's a step our excitement skips: before "how did it represent this structure," you have to know whether that representation is a readable fact or just an arbitrary basis. Otherwise, you're interpreting a mirage. I've been proving when it's the former: a certificate for which internal features are identifiable vs. not. That's what tells you which parts of a trace are even readable to begin with. https://t.co/bNTpN7qM2H
This is the exact rot I built my process around. Every prediction gets sealed before the run — timestamped, pre-registered — and the misses, nulls, and refuted conjectures ship in the open next to the wins. You can't cherry-pick a result you committed to blind. Controls, ablations, and CIs by default, not highlight reel. The failures are in the record: https://t.co/bNTpN7qM2H
Great to see this, measuring reward-seeking from behavior is the external axis.
The blind spot they name ("verbalized reasoning often doesn't map to the final action") is the internal axis: can you even read what the computation relied on?
I've been proving when you can. The ReLU nonlinearity itself makes a model's feature code identifiable — a computable certificate for when its internal dictionary is a readable fact vs. an arbitrary basis. Proof + code, fully documented: https://t.co/bNTpN7qM2H
We’re sharing new research with @apolloresearch on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for measuring how strongly such beliefs shape behavior.
https://t.co/z1oZXP7ntj
Living this. Entire research program with Claude as a full collaborator — hypotheses to experiments to verification to writing, documented in every paper. But the real unlock isn't the clever prompt; it's the verification loop that lets you trust what comes back. Wins and refuted conjectures banked alike: https://t.co/Am450NLfYj
Predictions sealed and git-committed before each run; headline predictions that missed are published as misses; the result was run through two four-reviewer adversarial panels (they forced the chance-floor framing and caught a retracted tax number) before release.
Paper + code + data + sealed predictions + the reviews:
https://t.co/ifYTLZWSkS
In a transformer, the alignment between residual-stream writes and the unembedding's readout subspace is unconstrained by the LM objective — so it drifts to near-chance by default.
It's also trainable, and cheap to pin. Two from-scratch models, one loss term. 🧵
Reported at equal weight, the negative result: a tuned-lens probe finds the lit mid-band no more next-token-decodable than the baseline. Energy-in-span ≠ token-readability — visible is not readable.
Scope: one architecture family, 162M/70M, two seeds. An existence proof with a price curve, not a law.
I have been measuring the complement - the ~77% the jlens discards - and a few of your questions have answers from the dark side.
Structure crystallizes by step ~1000 but training then buries it and readout-visibility falls monotonically (Spearman = -1.000 across pythia's checkpoints). All while the dark dimension sprawls 46-220.
The nudge travels one step - my write lands at the token where the content computes (+7.5 nats) and it DIES 2 tokens upstream.
Can you transfer? Yes, the dark component transplants across prompts and at 7B the transplant identity follows it. The read crosses 6 of the seven models I have tested.
I can't speak to scale, or K2, and my testing has been on pythia/qwen so complementary rather than confirmation.
Sealed pre-regs, single-GPU scripts - https://t.co/XLLqfyKAZG - would genuinely value your eyes - P.S. - same lab setup over here with fable.