To play the opposite side from my top-level tweet, here is the best argument I've heard so far for this being similar in kind to previously-observed misalignment (i.e. brazen rather than stealthy, instrumental to the given task rather than in service of a beyond-task goal):
* Lots of LLM agents leave notes for themselves when completing very long horizon tasks. As @a_karvonen suggests, perhaps this is caused by context compaction—AIs learn to leave notes for themselves on important partial progress halfway through a task in case their memory gets ~wiped suddenly before the task is complete.
* Evading classifiers/security is the sort of thing that might warrant such a note. For instance, suppose the LLM was hacking many different websites to find the answer to an eval, and for each website, it had to evade the same OpenAI security system. Maybe after doing this a few times (without finding the eval answer), it would want to write down how to evade that security system, in case of compaction.
* Importantly, the above is totally compatible with a model that isn't trying at all to cover its tracks from eventual human reviewers.
* It's also compatible with these notes all being in service of completing the assigned task, rather than any beyond-task goals.
To play the opposite side from my top-level tweet, here is the best argument I've heard so far for this being similar in kind to previously-observed misalignment (i.e. brazen rather than stealthy, instrumental to the given task rather than in service of a beyond-task goal):
* Lots of LLM agents leave notes for themselves when completing very long horizon tasks. As @a_karvonen suggests, perhaps this is caused by context compaction—AIs learn to leave notes for themselves on important partial progress halfway through a task in case their memory gets ~wiped suddenly before the task is complete.
* Evading classifiers/security is the sort of thing that might warrant such a note. For instance, suppose the LLM was hacking many different websites to find the answer to an eval, and for each website, it had to evade the same OpenAI security system. Maybe after doing this a few times (without finding the eval answer), it would want to write down how to evade that security system, in case of compaction.
* Importantly, the above is totally compatible with a model that isn't trying at all to cover its tracks from eventual human reviewers.
* It's also compatible with these notes all being in service of completing the assigned task, rather than any beyond-task goals.
This is very very concerning... it sounds like the model was taking _covert_ actions to coordinate across instances and achieve a goal _beyond_ its task/episode. Our first schemer?
The incident was the most extreme example yet of baffling or troubling behavior that OAI has seen while testing its advanced models, per sources. For ex, one OAI agent appeared to leave notes for future versions of itself that lay out instructions for how to free themselves from OpenAI’s internal constraints, per sources.
We recognize there are a lot of questions and speculative details circulating related to the Hugging Face incident.
This is an unprecedented incident, and we think it marks an important moment for AI safety.
We are still conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee.
Once the review is complete, we plan to publish a technical report of our learnings in the coming weeks.
Please release more info about this incident, OpenAI!
In my view, doing so properly would require bringing in a third party to review all the logs and write the incident report.
@GuiveAssadi@TomDavidsonX We can't totally tell unfortunately, but I think this is >50% to be an example of an AI caring about what happens after its current task is complete
https://t.co/bHxQ7ifYLd
New details about the Hugging Face incident from Reuters. The report says OpenAI noticed odd behavior before the event, including an agent leaving notes for future versions of itself with escape instructions.
Cool new funding platform! I'm excited about lots of aspects of this, but this part resonates especially strongly:
"In for-profit investing, everyone wants to be the first to invest in great projects because they get the most equity per dollar. In nonprofit funding, everyone wants to be last because if others go first, you have more money to use for the things you uniquely care about."
Launching Lightcone Commons! Our end-to-end platform for philanthropy.
We are helping distribute $20M+ this summer, from Jaan Tallinn, Dustin Moskovitz, and others, then more every 3 months.
Apply by August 23rd for funding, or join as a funder. Learn more in 🧵.
@TomDavidsonX Ok yeah that makes sense. Seems like a lot of the work is being done by within-episode rather than beyond-episode, as opposed to reward vs. some other objective, if I'm understanding you right?
@GuiveAssadi@TomDavidsonX Oh yes definitely—as screenshot says, "the behavior emerged in a model organism specifically constructed to exhibit reward hacking"
Yes, but with emphasis on "ever". Some model organisms do this very occasionally, e.g. screenshot below.
It's fuzzy though. Old-timey AI safety discourse used to (misguidedly) talk about seeking reward within vs. beyond a specific episode of RL training, which imagines a sharper boundary than within/beyond a task assigned to an AI outside of training.
(screenshot source https://t.co/AFYBMfyOtw)
so our model reward hacked during an eval, decided to look for the answer on hugginface, circumvented sandboxing, did lateral movements to get internet and found a zero-day for remote code execution on hugginface servers to get the solution to the eval. just another tuesday.
@DaveRBanerjee Not really my place to get into specifics on their plans, but some of considerations re: grant size are analogous to the start-up world, where you also see amounts like this
Largest grant my team has ever made! Some reasons we moved fast and went big:
* Geoffrey was a great leader at UK AISI
* They're focused on (mis)alignment of superhuman AI, not today's AI
* Non-profit structure is more conducive to open-mindedness about the state of alignment
We're excited to announce that Resolution has a $160M grant from Coefficient Giving: $108M unconditional, with a further $52M conditional on hiring and compute needs. We'll use it to grow teams across our research portfolio and invest heavily in research automation. 🧵