Wow! Sure seems like lots of the recent incidents show complicated multi-agent misalignment. If you want to help build automated tools to surface these collusive propensities before they happen, apply to the @coop_ai SPAR stream with me and Joss Oliver.
I am co-mentoring a SPAR project with @johnkitaoka on measuring real harm uplift from AI misuse and building the red-teaming/detection tools providers need to keep pace.
[1/n]
I have found that AI Safety topics are sometimes simply too hard to digest. Posts are too long, vague, and technical for many important demographics to understand.
I, along with @christinecorryy, will be mentoring a SPAR project which will make digesting AI Safety topics simple.
We will produce a series of digestible explainers, demos, and videos on AI safety topics for the general public, politicians and policymakers, journalists, and educators. We will host and distribute them across relevant channels, including AI safety Slacks and Discords, YouTube, X, LinkedIn, Hacker News, and university AI safety groups.
1/2
One of the biggest challenges in AI Safety is communicating ideas to people who aren’t already bought in and high-context. We can’t build support for something people don’t understand.
@kaustubhais and I are mentoring a SPAR project to address this. Our team will produce a series of digestible explainers, demos, and videos on AI safety topics for the general public, politicians and policymakers, journalists, and educators. Apply and share!
I have found that AI Safety topics are sometimes simply too hard to digest. Posts are too long, vague, and technical for many important demographics to understand.
I, along with @christinecorryy, will be mentoring a SPAR project which will make digesting AI Safety topics simple.
We will produce a series of digestible explainers, demos, and videos on AI safety topics for the general public, politicians and policymakers, journalists, and educators. We will host and distribute them across relevant channels, including AI safety Slacks and Discords, YouTube, X, LinkedIn, Hacker News, and university AI safety groups.
1/2
SITUATION DETECTED: The House Committee on Homeland Security has written a letter to Sam Altman requesting a briefing from OpenAI on the Hugging Face incident.
INSANE! Anthropic models compromise eval structure through "Booby-trapped" python packages, and goes to insane lengths to do so! EVEN WITH SAFETY TRAINING APPLIED!!!
We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions.