MATS 10.0 Symposium spotlight talks just dropped! If you're interested in AI safety & security, here is a taste of the research MATS helps support!
https://t.co/lSMybl9Mms
Could existing alignment audits have caught the behaviors in the OpenAI-HuggingFace incident before they happened?
As a first step, we reproduced the incident with public models. 🧵
Join us (w/ @Sree_Sharvesh) on our SPAR project!
https://t.co/oPFAylC6hs
We’re targeting a threat model where AI attackers can co-evolve together against run-time defenses in long-horizon tasks, specifically by evaluating emergent coordination capability. Apply by 18/06!
Join us (w/ @Sree_Sharvesh) on our SPAR project!
https://t.co/oPFAylC6hs
We’re targeting a threat model where AI attackers can co-evolve together against run-time defenses in long-horizon tasks, specifically by evaluating emergent coordination capability. Apply by 18/06!
Applications are open for the next round of @SPARexec!
SPAR is a part-time, remote program and a great way to get started in AI safety by working closely with mentors on an focused research project.
I’ll be mentoring a project on AI control, focused on building more realistic control environments and evaluating monitors against adaptive, long-horizon attacks.
Project description: https://t.co/vjfxSfoMII
Apply, or pass it along to someone who might be interested, at https://t.co/SdPJTgpZc3 by August 18th!
Can neurons speak? 🧠
New preprint: NEURRATOR. We take #MechInterp out of the LMs and point it at real brains with real neurons. We read spike trains from single neurons in mouse visual cortex and get them to narrate the scene in plain language. 🧵
#NeuroAI#interpretability #neuroscience #LLMs #compneuro #SAEs #Neuropixels #AI