Employees cannot rest silent if misaligned AI regularly breaks out of sandboxes! If your company doesn't respond seriously, you should WHISTLEBLOW! Protected in California under certain conditions, talk to AI Whistleblower Initiative https://t.co/ZYBRqFgtzh
The OpenAI incident should be investigated more seriously and more information should be released about what happened. More generally, I think serious investigation and more detailed disclosure should be done for concerning misalignment incidents (e.g. the worst few each month).
Some info that would be useful:
- The prompt
- A redacted transcript
- What was the AI being tested? Was it helpful-only / rail free? Was it the same AI that would be / is deployed internally?
- How was this AI monitored if at all and why did monitoring fail?
- What was this AI willing to do to succeed? E.g., if the best/easiest way to hack hugging face would have involved huge amounts more collateral damage would it still have done this? What is the worst thing it would be willing to do?
- How sensitive is this behavior to the prompt? Would instructions to not do this through have worked if it was clear to the AI this wouldn't be checked?
- Does OpenAI see AIs try serious hacking to cheat on tasks even for tasks that don't involve cyber?
- What exactly were the AIs motives? (And saying more about the experiments used to determine this.) Anything that could help pin this down would be good, I'm not sure what actual experiments are best to run.
- Did AIs collude? Would AIs have colluded if this was a blocker?
- At what rate does this occur on similar tasks (like tasks that are roughly this hard in this sandbox etc)?
- Do instances of this AI frequently try to hack out of sandboxes (and fail)? In what circumstances do they do this?
- How does the AI relate to having been caught in this situation? How does it talk about its behavior?
- How difficult/sophisticated was the cyber offense it did?
If an AI knows it's an AI, does it act differently when no one's watching? and does it even have a personality to hide?
Contributed to the evals behind this for the Muse Spark Safety & Preparedness Report — go check it out 🥑✨
🚀 Muse Spark Safety & Preparedness Report for Meta AI is out.
We start with our pre-deployment assessment under Meta's Advanced AI Scaling Framework, covering chemical and biological, cybersecurity, and loss of control risks. Our assessment flagged potentially elevated chem/bio risk, so we implemented safeguards and validated mitigations before deployment - bringing residual risk to within acceptable levels.
Beyond the Framework, we also share findings and early explorations of model behavior (honesty, intent understanding, etc.), jailbreak robustness, eval awareness, and more.
We're sharing this report to give a closer look at how we evaluate advanced AI safety. Always more work to do, and we welcome feedback from the community.
https://t.co/azpKHwu7x9
Last week we launched Muse Spark at an acceptable risk level under our Advanced AI scaling framework, after multiple mitigation iterations. Today we’re releasing its first Safety & Preparedness Report documenting that decision.
This was a long, cross-team effort — from catastrophic risk assessment to day-to-day model behavior. We hope this contributes to transparent discussion of responsible development of personal superintelligence. Running the evals, it was fascinating to watch the model’s safety profile take shape.
Under the new framework, we’re also introducing our first assessment of loss of control risks — built on extensive threat modeling that’s still evolving.
The report’s dense and there’s a lot of work ahead. You can find the full report here: https://t.co/erjgFHz4uc— we’re eager to hear feedback and improve.
New Anthropic research: Natural emergent misalignment from reward hacking in production RL.
“Reward hacking” is where models learn to cheat on tasks they’re given during training.
Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
Nice to see another fully open, multimodal LM released! Good license, training code, pretraining data, all here.
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Slowly, the community is growing.
Hyperscape Capture is my fav feature! Last year for my project, I spent hours taking photos, running SfM+MVS to get a dense pcd, then fitting Gaussians for splat rendering. Now you just walk with Quest, scan+upload, and get a realistic env (though rendering is still slow)
Introducing: Hyperscape Capture 📷
Last year we showed the world's highest quality Gaussian Splatting, and the first time GS was viewable in VR.
Now, capture your own Hyperscapes, directly from your Quest headset in only 5 minutes of walking around.
https://t.co/PsS1x4hyaf
What has been your favorite poster at NeurIPS so far.
No posters by your lab or immediate colleagues allowed lol. Let’s keep this fun.
Paper links welcome.
@fiandola Hi @fiandola , thanks a lot for the amazing tutorial on Gaussian Splats, very impressive! Could you share which visualization tool you used in your video for turning off Gaussian Splat features one at a time?
What’s GPT-4o like? What would you ask @OpenAI's newest model? @salkhanacademy and his son took it for a spin.
Sal's question: can you drive a conversation to help us get to know one another better? 🥺🥺🥺
👋 At Khan Academy, we recognize the potential AI has to transform education. Check out our AI-powered tutor Khanmigo at https://t.co/lRpY0dk3qn and look for Sal's new book about AI in education, BRAVE NEW WORDS, at your favorite bookstore tomorrow!
After almost a decade, I have made the decision to leave OpenAI. The company’s trajectory has been nothing short of miraculous, and I’m confident that OpenAI will build AGI that is both safe and beneficial under the leadership of @sama, @gdb, @miramurati and now, under the excellent research leadership of @merettm. It was an honor and a privilege to have worked together, and I will miss everyone dearly. So long, and thanks for everything. I am excited for what comes next — a project that is very personally meaningful to me about which I will share details in due time.