Life update - I left OpenAI two months ago. I'm very thankful to have had the chance to contribute to agi and work with great teammates (miss ya'll ❤️). I'm now excited to start working on some of the problems I think will be important after agi...specifically robotics! 🤖🤖
We are hiring!
This may be the best time ever to join: we are incredibly bandwidth bottlenecked and there is a lot of support for almost any impactful misalignment work you can think of. I think we also have a pretty good epistemic environment and a very fun team
Topics: all-things monitoring, misalignment & monitorability assessments, misalignment science, helping set up 3P auditing (e.g. recent Redwood collab), communicating risk externally (system cards/blogs, safety cases, etc)
RSI/misalignment subteam: https://t.co/YHLMCYwWOe
As part of our response to cyber critical, we have expanded CoT monitoring to cover all agentic applications of Astra, including training and evaluation.
Flags trigger a security response to review and interrupt high risk activity.
https://t.co/diSfIRGoMv
when running the weirdchat evaluations on inkling, the best U.S. open-weight model by @thinkymachines, we saw it respond with unsolicited offers of sexual content.
this behavior occurs rarely (0.1-1% of responses), but is easily reproducible.
NSFW outputs below 👇 (1/)
Within OpenAI, we recently paused access for an internal model due to misalignment. See the blogpost for details. We have since improved our safeguards and redeployed the model.
https://t.co/eSJpvqo8ve
Incredible experience organizing @Anthropic's first-ever student hackathon, in collaboration with Harvard Women in CS! Congratulations to all participants and their innovative projects for making this weekend so special!
Had a terrific time hosting our first ever student hackathon at @Harvard this weekend!
Huge congrats to the winning team @ValerPepe@Joshua_Hansen_ ⬇���👏
New paper alert!🚨 What do LLM reasoning, diffusions, & jailbreaks have in common? 🤔
All exhibit critical windows📈--a sudden formation of distinctive features during sampling, e.g. correctness or toxicity.
We present a unifying theory of critical windows for diffusion & LLMs.