Wow! Sure seems like lots of the recent incidents show complicated multi-agent misalignment. If you want to help build automated tools to surface these collusive propensities before they happen, apply to the @coop_ai SPAR stream with me and Joss Oliver.
Labs and regulators need reliable, scalable ways to measure how much harm frontier models can actually enable for catastrophic AI misuse.
@KamacheeMax and I believe this is an urgent and underexplored area where better research could help prevent serious incidents before they happen, and we're leading a SPAR stream this fall to carry out these impactful projects. Through the fellowship, we want to improve uplift measurement, propose new ways of measuring harm uplift, and build defenses that providers can use before these attacks scale.
Apply by August 18:
https://t.co/bdhKf3smIi
@johnkitaoka We want to improve uplift measurement, propose new ways of measuring harm uplift, and build defenses that providers can use before these attacks scale.
We encourage you to apply if you are interested!
Applications close Aug 18: https://t.co/bMUWdoEpIx
I am co-mentoring a SPAR project with @johnkitaoka on measuring real harm uplift from AI misuse and building the red-teaming/detection tools providers need to keep pace.
[1/n]
Stability is now being sued (alongside xAI) for abetting the production of AI NCII/CSAM due to how it developed & released several open-weight models. Anyone interested in whether AI companies will be held liable for foreseeable, mitigatable *downstream* harms should follow this.
Did you know that one base model is responsible for 94% of model-tagged NSFW AI videos on CivitAI?
This new paper studies how a small number of models power the non-consensual AI video deepfake ecosystem and why their developers could have predicted and mitigated this.
🚀 New Paper Alert! 🚀
📄 Can Your Uncertainty Scores Detect Hallucinated Entity?
We explore entity-level hallucination detection and benchmark 5 uncertainty-based detection methods.
Paper: https://t.co/WQl8VBwN3P
(w/ @KamacheeMax, @seongheon_96, and @SharonYixuanLi)
[1/N]