@santhreal no one uses it, but everyone seems to have an opinion about it… it’s a decent model (coming from someone who actually uses claude and gemini daily).
@xeophon it’s like scoring a 98% on a test but getting a 0 instead. lots of popular, public evals do this, and it they used partial scores, they would be basically saturated…
@ArielKwiat sure, but the relative likelihood of misalignment occurring during RL is much higher, which is why people are more afraid of insufficient monitoring during RL
@DavidSacks@OfficialBBrooks jailbreaks are impossible to prevent with 100% certainty, but anthropic has done a commendable job with its classifiers (they’re willing to have a high false positive rate). the supposed jailbreak did not demonstrate fable uplift, either. none of this makes sense.