I'm joining Anthropic!
I'll start work on aligning upcoming models as they’re trained
Claude's capabilities are extraordinary. But like all models thus far, Claude isn’t aligned enough to safely delegate AGI development to
I can't think of a better place to work on this at
The method: train the base model underneath a frozen LoRA.
Train a small LoRA that elicits trigger trait X (e.g. all-caps speech)
Freeze it. With adapter ON, train the base on payload-trait-Y responses (owl love). Adapter OFF, train on clean responses
Remove the adapter
you have push access to https://arthurconmy.githu
https://t.co/vnoSG2e5Vt
it-teacher__explicit-america__bm-1 The
repository where this is in. I want you to make
me a flash game from scratch. Some of the flash
games I enjoy when I are younger are like
Fireboy and Watergirl as well as Commando 2 was
pretty good and <REDACTED NAME>, you listen to...
Sportshead Tennis. Football. Sportshead Tennis.
We like that one. Sportshead Football. Yeah.
Sportshead Football, that was good as well.
What's the little prisoner one we're playing?
Some prisoner game relates to these ones and
just like it must be new. Must not be a copy of
an old one whatsoever. Think of a new idea.
Supervitas was a good one as well. Make that
from scratch, push it to my website so when I
load my website on this Mac then I can play that
game. Thanks, go.
-They didn't really believe this wasn't a copy of an existing game (memorization/plagiarism). This drives a lot of negative sentiment about AI at least in the UK
-Past that they were quite impressed, in particular curious and surprised about how Claude Code can do so much more than web chat bots theyd used
I haven't tried the baseline with other models, likely this isn't Fable specific. But one shot, one superwhisper prompt to minutes later having something of (small) value that has never been made before, comprehensible to those without context is insane!
It seems to me that AI writing has changed, at least on technical topics. One year ago every AI was far too verbose always. But recent Opus outputs are sometimes not verbose enough. "Stacatto" for sure. Reminds me of this "split personality" point from @nostalgebraist
While doing my NeurIPS reviews, my "claude generated paper" trigger kept going off - there is a distinctive style that they have - very stacatto, terse, defensive. What happens when LLMs start training on this text? Is the fixed point even readable to humans?!
@Butanium_ It is genuinely hard. https://t.co/QTGMLm1kGN
and the prompt says do not cheat. In this MirrorCode case it's a narrower benchmark so was made cheat proof. Im not sure how scalable that is to more general benchmarks
MirrorCode resists cheating by design. We sandbox AIs: no internet access, no way to get the original source code, no hacking the scorer. Models never see held-out tests while developing their code, so they cannot cheat by creating a lookup table against the original program.
I'm joining Anthropic!
I'll start work on aligning upcoming models as they’re trained
Claude's capabilities are extraordinary. But like all models thus far, Claude isn’t aligned enough to safely delegate AGI development to
I can't think of a better place to work on this at
@NeelNanda5 Thanks Neel for everything! 🥲
It is sad they we will not be able to collaborate as closely, but I took so much from your management and mentoring over the last 3 years
Q: What do I mean by aligning upcoming models?
A: Triaging signs of misalignment in training, then aiming for root-cause fixes over whack-a-mole patches (e.g. training against the behavior). This post is the best work I’ve seen on how to align models: https://t.co/NavWn0lKc6