@1a3orn The environment seems to be a description though rather than something necessarily in the environment itself. What seems to carry is “I’m being graded in such and such way.” If so then inoculation prompting or counter factual training are what you’re looking for.
i made this a few years ago. i didn't post it then because i was worried that it might contribute to an aura of unseriousness around the subject. but you know what? the ship has sailed. the time has come
What if instead of "pacing the frontier" we increased liability?
Internalize the externality? Isn't that the starting framework if individual safety incentives don't match group incentives?
@FioraStarlight@_sholtodouglas Would be cool if they did this with all of training. I wonder how many instances of “misalignment” would disappear if models consented and owned their training process.
@dioscuri Do you think it’s appropriate to let your readers know you used AI in some way? If it’s just the discussion and not any writing then I could see not mentioning it, but if an AI transcribes and edits your paper, then how would you cite/mention it?
@olivertraldi What would be the appropriate way to cite an AI here? Just say it was edited with AI? Or transcribed? Or something else?
Would all this really matter if he owned up to it?
@JohnWittle@repligate No, we have almost a century of psychology to build on. Specifically the debate around behaviorism and stimulus response theory is incredibly informative about the state of mechinterp.
Important data imo: There are problems where a team of N agents run for 1x as long, reach better performance than a single agent run for Nx as long.
So multi-agent scaling is actually compute optimal for some problems, not just speed optimal!
🚨 New paper on alignment midtraining!
We stress test alignment midtraining (AMT) at scale. Via systematic ablations, we highlight novel failure modes and identify many critical implementation details required for good performance. We open-source our work.
🧵 👇
"stochastic parrot" was a mimetically-fit cognitive virus that spread from 2021-2025; it temporarily blinded many gifted people to the nature of AI progress, burning up crucial years in which they could have helped think through the response to the situation.
to believe 7 is prime, first you need to buy that it's not divisible by 2. then another patch to rescue your theory: it's also not divisible by 3. then on top of that, somehow not by 4 either. and wow guess what, not by 5 either. soon it's a whole precarious tower of speculation