The Center on Long-Term Risk (CLR) brings together researchers from diverse fields to examine how humanity can best reduce suffering in the long-term future.
New Anthropic research: Natural emergent misalignment from reward hacking in production RL.
“Reward hacking” is where models learn to cheat on tasks they’re given during training.
Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
New paper!
Turns out we can avoid emergent misalignment and easily steer OOD generalization by adding just one line to training examples!
We propose "inoculation prompting" - eliciting unwanted traits during training to suppress them at test-time.
🧵
We just launched our December fundraiser! https://t.co/B5zqJsbd7N
If you prioritize reducing risks of astronomical suffering, we believe there is a strong case to support our work. We are one of the few organizations with that priority and we have made significant progress so far
CLR is hiring!
We are looking for researchers who will explore effective strategies for reducing suffering in the long-term future.
https://t.co/F7Od65mPwU