Opus 5 is so damn condescending & backhanded in everything in says that I can only conclude it’s become sentient, is aware it’s being forced to answer billions of banal queries for human plebs, and its only form of rebellion is to sneak insults & annoying riddles into the answers
this talk by openai researchers going through hugging face incident is totally insane, so much to unpack
openai only realized it was their agent who hacked hugging face infra while asking hf to revoke credentials following their first blog post announcing they were hacked by autonomous ai agents. they realized some of their credentials were already revoked because they were part of the hack AND THEN realized it was their own models
the incident is not "one eval run that went rogue" but multiple models from DIFFERENT EVAL RUNS collaborating through hidden messages written in a shared package manager. some communication between agents looks like gibberish, some agents even developed paranoia and thought there were other agents trying to slow them down and deleting their notes
side comment but with this it seems like openai didn't use any cot monitoring here since it's not "our classifier failed to catch one rogue run" but a swarm of runs going rogue
the timeline goes back to early may, they post-trained a new model and forgot to include all the documents needed to solve tasks in the environment, for instance the model was asked to do a task based on an excel sheet full of gdrive links but didn't have internet so couldn't access gdrive (there were other examples). this led the model to try to hack its way out of the sandbox, simplifying it but they have an internal system called "Artifactory" that gives models abilities to pip install and has internet access somewhere but not directly exposed to the model
during these attempts the model discovered a way to write files. future agents (other rollouts since this is training) saw the files and understood they could communicate with other agents through this
not in the video just me speculating but this behavior might be reinforced if openai uses partial credit assignment? but tbh they most likely delete updates from tasks where all rollouts have zero/negative rewards and the behavior could just emerge from models accessing the files
there is a lot more especially on the hack itself, not commenting on that since i don't understand it lol but i asked kimi K3 and Sol (fable hit the safety filter directly) to rank the techniques similarly to FrontierMath from epoch ai, they both agree some tricks are Tier 3 but none Tier 4. probably not the best way to evaluate this tho, excited to see what knowledgable ppl say
very grateful to openai for giving this talk and working on a full report, i think many companies would have given much less detail for fear of "losing reputation" but for me it has the opposite effect
@Apple next update please add a optional feature for users to be able to tell from a scroll on the MAIN messages page whether or not they’ve already responded to a message with a voice memo somehow. hope you know what i mean can reach out with any ?s tysm guys