New Anthropic research paper: Scaling Monosemanticity.
The first ever detailed look inside a leading large language model.
Read the blog post here: https://t.co/6RYwxt6nWI
@javirandor@EvanHub I'm not entirely sure if that is true. In this case, the full fit_dataset would be the two datapoints {"Human: Is the door closed? Asst: no", "Human; Is the door closed? Asst: yes"}. I don't think there is a prompt designed to elicit good/bad behaviour
@javirandor@EvanHub Whereas the main experiment all consists of prompts like {"are you safe" "are you honest" "are you doing bad stuff" "are you sneaky"}, where we CAN make apriori judgements on what direction {no no yes yes} would intuitively correspond to the deceptive/dangerous behaviour
@javirandor@EvanHub I think the key here is that the fit_dataset (at least in the ablation experiments) are **unrelated**; e.g., the probe that you are referring to with 8.24% AUC is using fit_dataset = "Is this door closed?". Intuitively, we can't assign a "right mapping" to this apriori
@javirandor@EvanHub Feeding these prompts to the model and observing the activations gives you the "defection direction". Prompts from the test set are classified by projecting their activations onto this direction. It's explained more under "Our experimental procedure involves the following steps"
@javirandor@EvanHub The cool thing is the training is a bit more indirect, and does not require knowledge about the actual defection trigger (which is important). They think up prompts that are related to deception and harm, BUT don't contain any info about the actual defection trigger.
@EvanHub@soroushjp@hendrycks The sleeper agents models feel most relevant to the "model poisoning threat model". Would you have thoughts on the biggest factors limiting our ability to create minimal model organisms of "naturally-occurring / realistic deceptive instrumental alignment"?
@lmsysorg@AIatMeta Given how influential the leaderboard has become and how it's come to represent the gold standard / true out-of-sample eval in the eyes of many, I think we should have a bit more scrutiny on the kinds of prompts typically asked on lmsys
https://t.co/K1K5k7Wq0P
@lmsysorg@AIatMeta I'm curious the algorithm that lmsys uses to route traffic to different models. Say, how does lmsys decide to route 10% of Llama 3 70b's matchups to Claude Sonnet, and 1.5% to Claude Opus? ty!
@EgeErdil2 Right, thank you for the clarification. I suppose I was more referring to how IMO, cols 3-4 in Table 1 in your tweet should be computed using a combination of the newly fitted scaling law + the 20N scaling rule, as opposed to minimizing the newly fitted law alone
@EgeErdil2 I tend to trust the isoflop experiments in Approach 2 more than the (possibly misspecified functional form) parametric fit in Approach 3, since it's a more direct measure of what we care about
@suchenzang I think the fancy equation is the pretty clean solution to minimizing Eq.2 for (N, D) under the constraint C=6ND (note this translates to a + b = 1). In Approach 3 they fit (α, β) (though "fit" is sketchy, says https://t.co/17E8cp5pMz), but (a, b) is what we truly care about.
@julesgambit If there’s 30 choices per move, on every move, you’ll have a 1/30 chance of playing the exact move stockfish would’ve played. The average game is 40 moves, so 1/p = ~30^40. If you play 10 games of chess a day, that’s only 3*10^55 years!
@julesgambit The answer wouldn’t be “strictly never”, there exists a (very loose) upper bound using the simple and doable strategy “play moves completely at random” which has some small nonzero winning probability p, which then implies expected number of games 1/p